Every major LLM benchmark explained — what it tests, how tasks work, and where models stand.
Crowdsourced human preference Elo ratings from millions of real user comparisons.