MMLU
2020-09-07 — 2025-02-24 · lived 4 years, 5 months
Saturated — everyone knew the answers
Fifty-seven subjects, from elementary mathematics to professional law, and in 2020 the best GPT-3 managed 43.9 — nineteen points over guessing. By January 2025 OpenAI's o1 sat at 91.8, and errors in 6.49% of questions put the ceiling near 93.5. Anthropic's next flagship printed no MMLU score.
How the frontier caught up
At launch, then at the date we bury it. The ceiling is what the number could not get past, and it is not the same measurement on every grave — it says which.
A benchmark gets no deprecation notice. This death date is our editorial judgment: we bury a benchmark when frontier scores stop telling top models apart, or when major labs stop reporting it. The sources below are the evidence — judge it yourself.