THINK-Bench
A 1,375-question benchmark with 13,311 manually curated reasoning steps, specifically evaluating thinking efficiency and Chain-of-Thought quality in Large Reasoning Models across math, physics, and chemistry.
This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.
What It Tests
A 1,375-question benchmark with 13,311 manually curated reasoning steps, specifically evaluating thinking efficiency and Chain-of-Thought quality in Large Reasoning Models across math, physics, and chemistry.
Discovery notes
Notes the discovery agent wrote when proposing this benchmark.
First benchmark to systematically measure overthinking in reasoning models. Introduces efficiency metrics (tokens, first-correct-tokens, efficiency score). Reveals most LRMs waste tokens on easy questions.