Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
An upgraded abstract reasoning benchmark providing finer-grained evaluation of visual reasoning at higher cognitive complexity. Maintains input-output pair format with newly curated, more challenging tasks resistant to frontier AI systems.
Grade-school science questions that simple retrieval systems can't answer.
23 hard reasoning tasks where chain-of-thought is required to exceed human performance.
PhD-level science questions so hard that even experts with Google still struggle.
Commonsense reasoning — pick the most plausible continuation of an everyday activity.
A 1,375-question benchmark with 13,311 manually curated reasoning steps, specifically evaluating thinking efficiency and Chain-of-Thought quality in Large Reasoning Models across math, physics, and chemistry.
Can AI avoid repeating common myths and falsehoods that pervade its training data?
A comprehensive multi-task, multi-modal time series reasoning benchmark covering 4 dimensions (Perception, Reasoning, Prediction, Decision-Making) across 15 tasks and 14 distinct domains in finance, healthcare, industrial systems.