models.rip · benchmark
HumanEval
2021–2025 · 28.8% → 92.7% frontier score
164 hand-written Python problems, of which Codex solved 28.8% first try. By late 2024 GPT-4o, Claude 3.5 Sonnet and a 32-billion-parameter open model sat at 92.1, 92.1 and 92.7 — six tenths of a point across three labs. Nothing was capping the score. It had simply stopped separating anyone.