benchmark.darvinyi.com

Compare Models

Select 2–4 models to compare side-by-side across all benchmarks.

3/4 models selected — click to add or remove

Capability Profile

Average score across benchmarks in each category. Higher = better.

Coding: SWE-bench, LiveCodeBench, HumanEval / HumanEval+ · Math: MATH Benchmark, AIME, GSM8K · Reasoning: GPQA Diamond, BIG-Bench Hard, TruthfulQA · Knowledge: MMLU / MMLU-Pro · Agent: τ-bench (tau-bench), WebArena / VisualWebArena, AgentBench · Multimodal: WorldBench, VS-Bench (Visual Strategic Bench)

Pricing (per million tokens)

Kimi K2.6
Input$0.60
Output$4.00
Claude Opus 4.7
Not available
Gemma 4 31B
Not available
BenchmarkKimi K2.6Claude Opus 4.7Gemma 4 31B
80.2%
89.6%
~86%
53.7% avg accuracy49.7% avg accuracy

26 more benchmarks have no scores for the selected models.