ARC-AGI 2
An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.
This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.
What It Tests
An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.
Discovery notes
Notes the discovery agent wrote when proposing this benchmark.
Successor to 2019 ARC-Challenge. Submitted May 2025. All frontier models scored 0% on launch (March 2025); GPT-5.2 reached 54.2% by December 2025, requiring refinement loops. Tests genuine problem-solving beyond memorization.