Compare Models
Select 2–4 models to compare side-by-side across all benchmarks.
3/4 models selected — click to add or remove
Capability Profile
Average score across benchmarks in each category. Higher = better.
Coding: SWE-bench, LiveCodeBench, HumanEval / HumanEval+ · Math: MATH Benchmark, AIME, GSM8K · Reasoning: GPQA Diamond, BIG-Bench Hard, TruthfulQA · Knowledge: MMLU / MMLU-Pro · Agent: τ-bench (tau-bench), WebArena / VisualWebArena, AgentBench · Multimodal: WorldBench, VS-Bench (Visual Strategic Bench)
Pricing (per million tokens)
Kimi K2.6
Input$0.60
Output$4.00
Claude Opus 4.7
Not availableGemma 4 31B
Not available| Benchmark | Kimi K2.6 | Claude Opus 4.7 | Gemma 4 31B |
|---|---|---|---|
| 80.2% | — | — | |
| 89.6% | — | — | |
| ~86% | — | — | |
| — | 53.7% avg accuracy | 49.7% avg accuracy |
26 more benchmarks have no scores for the selected models.