benchmark.darvinyi.com
Updated 2026

The AI Benchmark
Explorer

Deep-dives into every major LLM benchmark — what they test, how they work, real task examples, and where the models actually stand.

Benchmarks

30

Agent Evals

9

Models Tracked

45

Recorded Scores

262

Across all benchmarks + agent evals

The Real Work Gap

Models that ace structured benchmarks often fail dramatically on real end-to-end professional work. The gap is larger than most people expect.

LiveCodeBench

91.7%

Fresh competition coding problems

Contamination-resistant coding benchmark

GAIA (Top Agent)

67.0%

Real-world multi-step tasks

With tool access; human baseline is 92%

RLI Automation Rate

3.75%

Real Upwork freelance projects

$143,991 of actual paid work • 240 projects

The same models scoring well on structured benchmarks complete under 4% of real freelance projects to client-acceptable quality.

Browse by Category

30 benchmarks across 11 categories.

Benchmark Status

Which benchmarks still differentiate frontier models?