models.rip · benchmark

MMLU

2020–2025 · 43.9% → 91.8% frontier score

Fifty-seven subjects, from elementary mathematics to professional law, and in 2020 the best GPT-3 managed 43.9 — nineteen points over guessing. By January 2025 OpenAI's o1 sat at 91.8, and errors in 6.49% of questions put the ceiling near 93.5. Anthropic's next flagship printed no MMLU score.