# HumanEval

> 164 hand-written Python problems, of which Codex solved 28.8% first try. By late 2024 GPT-4o, Claude 3.5 Sonnet and a 32-billion-parameter open model sat at 92.1, 92.1 and 92.7 — six tenths of a point across three labs. Nothing was capping the score. It had simply stopped separating anyone.

2021-07-07 — 2025-02-24

This URL is the guided walk — the 3D cemetery, standing at this grave. The written record is at [HumanEval's epitaph](https://models.rip/humaneval/epitaph).

## Elsewhere

- [The whole cemetery](https://models.rip/)
- [The dataset](https://models.rip/graveyard.json)
