Benchmarks
Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
Sort:
CodingEnriching…
2026ClarEval
A benchmark specifically evaluating code agents' ability to ask clarifying questions and seek information when faced with ambiguous programming instructions, not just task completion.
Pending curation
CodingSaturated
2021HumanEval / HumanEval+
Python function completion from docstrings, evaluated by test execution.
164 tasks
97.6%
CodingActive
2024LiveCodeBench
Contamination-resistant coding benchmark using freshly released competition problems.
1,055 tasks
91.7%
CodingContaminated
2023SWE-bench
Can AI resolve real GitHub issues on production codebases?
2,294 tasks
80.9%
CodingActive
2025SWE-Lancer
Real Upwork freelance software tasks mapped to $1M in economic value.
1,488 tasks
66.3%