benchmark.darvinyi.com

Agent Evaluations

Benchmarks that test AI on real human work — not synthetic tasks. How much economic value can AI agents actually deliver?

 

Why real-work benchmarks are different

Standard benchmarks test isolated capabilities — a math problem, a code snippet, a multiple-choice question. Real-work benchmarks test complete, end-to-end tasks: drafting a legal memo, completing a freelance animation project, or analyzing patient records and producing a care plan. The gap between the two is enormous. Models scoring 80%+ on SWE-bench Verified complete fewer than 4% of real Upwork projects to client-acceptable quality.

All Evaluations

Mercor (San Francisco)
2025

Mercor APEX + APEX-Agents

Professional knowledge work benchmark across investment banking, consulting, law, and medicine.

World-building infrastructure: 33 environments containing ~166 real files each in authentic Google Workspace + email + calendar + file storage — no other benchmark deploys this level of workplace fidelity.

880 tasks
67.2% ± 2.4%
OpenAI
2025

GDPval

AI on real professional tasks from 44 occupations covering $3 trillion in annual wages.

GDP-grounded occupational selection: the only benchmark that uses Federal Reserve GDP data + BLS wage data to determine which occupations and tasks to include — a scientifically defensible sampling frame.

1320 tasks
47.6% win+tie
CAIS / Scale AI
2025

Remote Labor Index (RLI)

Real Upwork freelance projects: $143,991 of actual paid work, 2.5% maximum automation.

True economic grounding: every project has a real dollar value set by the real professional who did the work — no synthetic cost estimates.

240 tasks
3.75%
Upwork Inc.
2025

Upwork HAPI

The first benchmark measuring how human expertise amplifies AI agent performance.

The only benchmark that measures human+agent collaboration as its primary signal, not agent performance in isolation.

322 tasks
93% (with human)
METR (nonprofit, Berkeley-based)
2025

METR Time Horizon

How long can AI agents work autonomously? The 50%-success time horizon, doubling every 3 months.

Single interpretable number: 'This model can complete 50%-probability tasks that take a human expert X hours' is more meaningful to policymakers than '74.2% on MMLU'.

228 tasks
~14.5 hours
Harvey AI
2024

BigLaw Bench

Real legal work quality — what percent of a lawyer-quality deliverable does AI produce?

Negative-point scoring: hallucinating legal citations actively hurts your score — unlike most benchmarks where wrong answers simply fail to earn points.

— tasks
89.22%
UC Berkeley RDI (Center for Responsible, Decentralized Intelligence)
2026

Agents' Last Exam (ALE)

UC Berkeley RDI benchmark evaluating computer-use agents on 1,500+ expert-sourced, real professional-project tasks across 55 subdomains and 13 industries, with deterministic outcome grading and a brutal 'Last-Exam' tier where frontier agents score ~0%.

Built with 250-300+ industry experts across 100+ institutions, sourcing tasks from real completed professional projects rather than synthetic or crowdsourced scenarios.

1500 tasks
24.0%
Zapier
2026

AutomationBench

Zapier's benchmark testing whether AI agents can autonomously orchestrate real cross-application business workflows (CRM, inbox, calendar, ticketing) via REST APIs, under real business policy constraints and misleading data.

Built directly from real workflow patterns observed on Zapier's own automation platform — grounded in actual production business-automation demand rather than synthetic scenarios.

657 tasks
51.4%
Frontis.ai / Tsinghua University
2026

EnterpriseClawBench

Benchmark constructed from sanitized real enterprise agent sessions, testing coding/computer-use agents on genuine workplace deliverables (documents, spreadsheets, web pages) with role-, skill-, and rubric-tagged grading; dataset kept private for confidentiality.

Sourced from a large archive of real, proprietary enterprise agent sessions rather than crowdsourced or synthetic tasks, aiming to reflect the genuine distribution of real workplace requests.

852 tasks
0.663
University of Hong Kong (XLANG Lab)NEW
2026

OSWorld 2.0

Benchmark of 108 long-horizon, real computer-use workflows (median 1.6 hrs, up to 318 tool calls) across 7 professional domains, testing whether agents can complete full end-to-end desktop/web jobs rather than short GUI clicks.

Focuses on long-horizon professional workflows (median 1.6 hrs) rather than short GUI manipulation tasks.

108 tasks
20.6%
Microsoft ResearchNEW
2026

HealthAgentBench

Microsoft Research suite of 54 agentic healthcare tasks across 7 real clinical/biomedical workflows (EHR auditing, X-ray report correction, clinical trial matching, pathology, CT), scored by task-specific verifiers with cost/time tracking.

Spans the full patient-journey and data-modality range (2D/3D imaging, whole-slide pathology, structured EHR, free text) in one suite, unlike single-modality medical benchmarks.

54 tasks
42%
Multi-institution research consortiumNEW
2026

ResearchClawBench

End-to-end autonomous scientific research benchmark: 40 tasks across 10 scientific domains where agents must rediscover a real published paper's findings from raw data and literature, graded by expert multimodal rubrics.

Grounds every task in a real published paper's actual data and literature, hiding the target paper to test genuine rediscovery rather than paraphrase.

40 tasks
20.7/100