Benchmarks
Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
Sort:
MathActive
2024AIME
Annual olympiad-level math competition used as a fresh, contamination-proof AI benchmark.
30 tasks
100% (Heavy)
MathEnriching…
2026EternalMath
A living, dynamically evolving benchmark that automatically transforms recent peer-reviewed mathematics papers into executable reasoning tasks with deterministic verification, updating continuously as new mathematics is published.
Pending curation
MathSaturated
2021GSM8K
Grade-school math word problems requiring 2-8 step arithmetic reasoning.
8,500 tasks
99.7%
MathNearing Saturation
2021MATH Benchmark
Competition-level math problems across 7 subjects, from AMC to AIME difficulty.
12,500 tasks
~97–98%