models.rip · benchmark
GLUE
2018–2019 · 70% → 88.4% frontier score
Nine tasks collapsed into one number, and the number was 70.0 the day it opened — low enough that its own authors called that a problem. Fifteen months later the frontier was 88.4, past the human baseline of 87.1, and those same authors published SuperGLUE because nothing was left here to measure.