models.rip · benchmark

HellaSwag

2019–2023 · 47.3% → 95.3% frontier score

Wrong sentence endings, filtered by machine until they were ridiculous to people and irresistible to models: humans scored 95.6, the best model 47.3. GPT-4 scored 95.3. In under four years the gap it was built to open had closed to three tenths of a point.