Claude Opus 5 leads on intelligence at 60.7, no single lab wins every category — the 8 rankings below are led by 5 different makers, and the cheapest usable model is $0.010 per million input tokens.
Every number on this page is read from the live catalog when you load it, so it is current by construction rather than by somebody remembering to edit it. That matters more than it sounds: the previous version of this post was written in July, and its headline claim about the smartest model was wrong within two weeks.
Claude Opus 5 leads at 60.7 on the Artificial Analysis Intelligence Index, followed by Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9). See the full ranking →
Claude Opus 5 leads at 78.0 on the Coding Index, followed by GPT-5.6 Sol (77.4) and GPT-5.6 Terra (76.7). See the full ranking →
Kimi K3 leads at 1379 on Design Arena's website board, followed by GPT-5.6 Sol (1344) and Claude Opus 5 (1343). See the full ranking →
Claude Opus 5 leads at — on the Agentic Index, followed by Kimi K3 (—) and Grok 4.5 (—). See the full ranking →
Morph V3 Fast leads at 5.0k t/s in output tokens per second, followed by Morph V3 Large (2.6k t/s) and Relace Apply 3 (1.7k t/s). See the full ranking →
Ling-2.6-flash leads at $0.010 per million input tokens, followed by Granite 4.0 Micro ($0.017) and Mistral Nemo ($0.019). See the full ranking →
Kimi K3 leads at 57.1 among models you can download and self-host, followed by GLM 5.2 (51.1) and DeepSeek V4 Flash 0731 (49.9). See the full ranking →
Grok 4.20 Multi-Agent leads at 2M in context window size, followed by Grok 4.20 (2M) and Llama 4 Scout (1.3M). See the full ranking →
The spread across categories is the point. 5 different labs lead the 8 rankings above, which means "what is the best AI model" has no answer without a second half to the question. The model that tops the intelligence index is rarely the one you should put behind a high-volume endpoint, and the one that wins on design is chosen by human voters rather than a benchmark harness.
If you want the adoption view instead of the benchmark view, the most-used rankings show which models people actually run in production, by task. Those two orderings disagree more often than you would expect, because price and availability decide as much as capability does.