Benchmarks
Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
Sort:
KnowledgeEnriching…
2025Humanity's Last Exam
A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.
Pending curation
KnowledgeSaturated
2020MMLU / MMLU-Pro
Broad academic knowledge across 57 subjects — the standard knowledge benchmark.
14,042 tasks
93.0%