# MMLU

> Fifty-seven subjects, from elementary mathematics to professional law, and in 2020 the best GPT-3 managed 43.9 — nineteen points over guessing. By January 2025 OpenAI's o1 sat at 91.8, and errors in 6.49% of questions put the ceiling near 93.5. Anthropic's next flagship printed no MMLU score.

2020-09-07 — 2025-02-24

A benchmark gets no deprecation notice, so this date is this archive's editorial judgment rather than an announcement.

## The record

- Cause of death: Saturated
- Frontier at launch: 43.9%
- Frontier at death: 91.8%
- Practical ceiling: 93.5% (question errors)

## Sources

- https://arxiv.org/abs/2009.03300
- https://blog.google/technology/ai/google-gemini-ai/
- https://arxiv.org/abs/2501.12948
- https://arxiv.org/abs/2406.04127
- https://www.anthropic.com/news/claude-3-7-sonnet

## Elsewhere

- [Stand at this grave](https://models.rip/mmlu)
- [The benchmarks plot](https://models.rip/benchmarks)
- [The whole cemetery](https://models.rip/)
