benchmark.darvinyi.com
← Back to Benchmarks
Knowledge

Humanity's Last Exam

A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.

Year2025

This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.

What It Tests

A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.

Discovery notes

Notes the discovery agent wrote when proposing this benchmark.

Addresses benchmark saturation: frontier models score <10% vs >90% on saturated MMLU. Published in Nature (January 2026). Features 2,500 public + private holdout questions. ~10% multimodal. Poor model calibration revealed.

Links