# HumanEval

> 164 hand-written Python problems, of which Codex solved 28.8% first try. By late 2024 GPT-4o, Claude 3.5 Sonnet and a 32-billion-parameter open model sat at 92.1, 92.1 and 92.7 — six tenths of a point across three labs. Nothing was capping the score. It had simply stopped separating anyone.

2021-07-07 — 2025-02-24

A benchmark gets no deprecation notice, so this date is this archive's editorial judgment rather than an announcement.

## The record

- Cause of death: Saturated
- Frontier at launch: 28.8%
- Frontier at death: 92.7%

## Sources

- https://arxiv.org/abs/2107.03374
- https://arxiv.org/abs/2305.01210
- https://arxiv.org/abs/2409.12186
- https://www.anthropic.com/news/claude-3-7-sonnet

## Elsewhere

- [Stand at this grave](https://models.rip/humaneval)
- [The benchmarks plot](https://models.rip/benchmarks)
- [The whole cemetery](https://models.rip/)
