Across 8 task categories, the overlap between a benchmark's top 10 and the 10 models most used for the work that benchmark measures averages 0.4 out of 10. The models leading benchmarks cost a median $1.20 per million input tokens; the models people actually deploy cost $0.145 — 8× less, for only 9% less measured intelligence.
Every leaderboard, this site's included, answers "which model scores highest." None of them can tell you whether that is the model anyone ends up running, because that needs traffic data rather than a benchmark harness. We hold both, so it is measurable — and the two orderings turn out to be almost disjoint.
The table pairs each task with the benchmark a buyer would reasonably consult before picking a model for it, then checks how many of that benchmark's top ten appear among the ten models actually handling that work in production.
| Task | Benchmark consulted | Benchmark #1 | Most used | Overlap |
|---|---|---|---|---|
| Code Generation | AA Coding Index | 1/10 | ||
| Debugging | AA Coding Index | 0/10 | ||
| Frontend & UI | AA Coding Index | 0/10 | ||
| Tool Dispatch | Tau²-Bench | 1/10 | ||
| Math | AA Math Index | 1/10 | ||
| Data Extraction | AA Intelligence Index | 0/10 | ||
| Summarization | AA Intelligence Index | 0/10 | ||
| Translation | AA Intelligence Index | 0/10 |
The tempting read is that the benchmarks are broken. They aren't. They measure a ceiling — what a model can do on a hard problem given no cost constraint — and they measure it carefully. The gap appears because almost nobody buys the ceiling.
Put the two populations side by side and the shape is obvious. The models leading benchmarks carry a median input price of $1.20 per million tokens and a median intelligence of 39.5. The models people actually deploy across all 29 tracked tasks cost $0.145 and score 35.9. That is 8× the price for 9% more intelligence.
Most production work is not hard. Classification, extraction, summarization and routine code sit well below the frontier, so the marginal points a flagship buys you are points you were never going to spend. What a benchmark leaderboard cannot show you is where the knee of that curve sits for your workload — and the knee is the entire decision.
Two honest caveats. This traffic is OpenRouter's, which skews toward developers and indie builders paying their own bills rather than enterprises on committed spend, so the price sensitivity here is real but not universal. And usage is a lagging signal: it reflects what people chose weeks ago, including out of habit, which is exactly the thing benchmarks are good at correcting.
Read the benchmark to find the smallest model that clears your quality bar, not the model at the top. Then read usage to see what everyone solving your problem has already settled on — a model carrying a large share of real traffic for your task has been load-tested by thousands of people in a way no eval can replicate.
Where they agree, the choice is easy. Where they disagree — which, per the table above, is nearly always — the disagreement is telling you the frontier model is overkill for that task.
Benchmarks from Artificial Analysis. Usage: Source: OpenRouter (openrouter.ai/rankings), as of 2026-08-04, a trailing 7-day window of classified, sampled traffic. Every figure on this page is computed at request time — re-check it whenever you like.