Humanity's Last Exam
A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.
This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.
What It Tests
A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.
Discovery notes
Notes the discovery agent wrote when proposing this benchmark.
Addresses benchmark saturation: frontier models score <10% vs >90% on saturated MMLU. Published in Nature (January 2026). Features 2,500 public + private holdout questions. ~10% multimodal. Poor model calibration revealed.