benchmark.darvinyi.com
← Back to Benchmarks
Reasoning

ARC-AGI 2

An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.

Year2025

This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.

What It Tests

An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.

Discovery notes

Notes the discovery agent wrote when proposing this benchmark.

Successor to 2019 ARC-Challenge. Submitted May 2025. All frontier models scored 0% on launch (March 2025); GPT-5.2 reached 54.2% by December 2025, requiring refinement loops. Tests genuine problem-solving beyond memorization.

Links