← the graveyard

HellaSwag

2019-05-192023-03-14 · lived 3 years, 9 months

Saturated — everyone knew the answers

Wrong sentence endings, filtered by machine until they were ridiculous to people and irresistible to models: humans scored 95.6, the best model 47.3. GPT-4 scored 95.3. In under four years the gap it was built to open had closed to three tenths of a point.

How the frontier caught up

At launch, then at the date we bury it. The ceiling is what the number could not get past, and it is not the same measurement on every grave — it says which.

47.3%95.3%ceiling 95.6% (human accuracy)

A benchmark gets no deprecation notice. This death date is our editorial judgment: we bury a benchmark when frontier scores stop telling top models apart, or when major labs stop reporting it. The sources below are the evidence — judge it yourself.

Sources