Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
LLM agents across 8 interactive environments: OS, databases, web, games, and more.
Annual olympiad-level math competition used as a fresh, contamination-proof AI benchmark.
An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.
Grade-school science questions that simple retrieval systems can't answer.
23 hard reasoning tasks where chain-of-thought is required to exceed human performance.
A benchmark specifically evaluating code agents' ability to ask clarifying questions and seek information when faced with ambiguous programming instructions, not just task completion.
A living, dynamically evolving benchmark that automatically transforms recent peer-reviewed mathematics papers into executable reasoning tasks with deterministic verification, updating continuously as new mathematics is published.
Multi-step real-world tasks that are conceptually simple for humans but require tool-using agents.
PhD-level science questions so hard that even experts with Google still struggle.
Grade-school math word problems requiring 2-8 step arithmetic reasoning.
Commonsense reasoning — pick the most plausible continuation of an everyday activity.
Python function completion from docstrings, evaluated by test execution.
A 3,000-question benchmark of expert-vetted academic questions across 100+ subjects (STEM, humanities, sciences) designed to be resistant to internet lookup and requiring genuine understanding, created by Center for AI Safety and Scale AI.
Contamination-resistant benchmark refreshed monthly from recent sources with no LLM judge.
Contamination-resistant coding benchmark using freshly released competition problems.
Crowdsourced human preference Elo ratings from millions of real user comparisons.
Competition-level math problems across 7 subjects, from AMC to AIME difficulty.
Broad academic knowledge across 57 subjects — the standard knowledge benchmark.
Benchmark for automatically converting GitHub repositories into autonomous, interoperable software agents for the 'Agentic Web'.
Can AI resolve real GitHub issues on production codebases?
Real Upwork freelance software tasks mapped to $1M in economic value.
A simulated software company with 16 AI colleagues testing real office work tasks.
A 1,375-question benchmark with 13,311 manually curated reasoning steps, specifically evaluating thinking efficiency and Chain-of-Thought quality in Large Reasoning Models across math, physics, and chemistry.
Can AI avoid repeating common myths and falsehoods that pervade its training data?
A comprehensive multi-task, multi-modal time series reasoning benchmark covering 4 dimensions (Perception, Reasoning, Prediction, Decision-Making) across 15 tasks and 14 distinct domains in finance, healthcare, industrial systems.
A 2,000-question multimodal benchmark for evaluating visual-textual search capabilities in MLLMs, requiring actual image analysis beyond text cues for complex fact-finding tasks with minimal dependence on general knowledge.
A CVPR 2026 Oral benchmark testing whether vision-language models can perceive, reason strategically (theory-of-mind), and make good decisions inside 10 vision-grounded multi-agent games.
Autonomous browser agents completing realistic tasks on functional sandboxed websites.
A 2,000-question multimodal reasoning benchmark built to maximize visual diversity, not just task diversity, across seven real-world visual domains.
AI customer service agents that must follow policy while solving real customer problems.