Vision-DeepResearch Benchmark
A 2,000-question multimodal benchmark for evaluating visual-textual search capabilities in MLLMs, requiring actual image analysis beyond text cues for complex fact-finding tasks with minimal dependence on general knowledge.
This entry was automatically discovered and hasn't been researched yet. Sections below fill in as enrichment completes. Discovered 7/13/2026.
What It Tests
A 2,000-question multimodal benchmark for evaluating visual-textual search capabilities in MLLMs, requiring actual image analysis beyond text cues for complex fact-finding tasks with minimal dependence on general knowledge.
Discovery notes
Notes the discovery agent wrote when proposing this benchmark.
Addresses gap in multimodal benchmarks by focusing on realistic vision-search scenarios. Evaluates MLLMs on visual retrieval requiring deep image understanding, not just VQA. Proposes multi-round cropped-search workflow.