# HellaSwag

> Wrong sentence endings, filtered by machine until they were ridiculous to people and irresistible to models: humans scored 95.6, the best model 47.3. GPT-4 scored 95.3. In under four years the gap it was built to open had closed to three tenths of a point.

2019-05-19 — 2023-03-14

A benchmark gets no deprecation notice, so this date is this archive's editorial judgment rather than an announcement.

## The record

- Cause of death: Saturated
- Frontier at launch: 47.3%
- Frontier at death: 95.3%
- Practical ceiling: 95.6% (human accuracy)

## Sources

- https://arxiv.org/abs/1905.07830
- https://arxiv.org/abs/2303.08774

## Elsewhere

- [Stand at this grave](https://models.rip/hellaswag)
- [The benchmarks plot](https://models.rip/benchmarks)
- [The whole cemetery](https://models.rip/)
