Benchmarks
Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
Sort:
MultimodalEnriching…
2026Vision-DeepResearch Benchmark
A 2,000-question multimodal benchmark for evaluating visual-textual search capabilities in MLLMs, requiring actual image analysis beyond text cues for complex fact-finding tasks with minimal dependence on general knowledge.
Pending curation
MultimodalActive
2025VS-Bench (Visual Strategic Bench)
A CVPR 2026 Oral benchmark testing whether vision-language models can perceive, reason strategically (theory-of-mind), and make good decisions inside 10 vision-grounded multi-agent games.
8,000 tasks
84.9%
MultimodalActive
2026WorldBench
A 2,000-question multimodal reasoning benchmark built to maximize visual diversity, not just task diversity, across seven real-world visual domains.
2,000 tasks
64.0% avg accuracy