HumanEval
2021-07-07 — 2025-02-24 · lived 3 years, 7 months
Saturated — everyone knew the answers
164 hand-written Python problems, of which Codex solved 28.8% first try. By late 2024 GPT-4o, Claude 3.5 Sonnet and a 32-billion-parameter open model sat at 92.1, 92.1 and 92.7 — six tenths of a point across three labs. Nothing was capping the score. It had simply stopped separating anyone.
How the frontier caught up
At launch, then at the date we bury it. The ceiling is what the number could not get past, and it is not the same measurement on every grave — it says which.
A benchmark gets no deprecation notice. This death date is our editorial judgment: we bury a benchmark when frontier scores stop telling top models apart, or when major labs stop reporting it. The sources below are the evidence — judge it yourself.