gpt-oss-120b is the most-used model for summarization on OpenRouter, handling 15.9% of requests classified as summarization over the last 7 days.
This ranks models by how often they are actually chosen for summarization in production traffic, not by benchmark score. Adoption reflects price, latency and availability as much as raw capability — which is exactly why it often disagrees with the benchmark leaderboards, and why it is worth reading alongside them.
| # | Model | Request share | Token share | Input /M | Context |
|---|---|---|---|---|---|
| 1 | 15.9% | 15.8% | $0.037 | 131K | |
| 2 | 9.1% | 7.8% | $0.140 | 1.0M | |
| 3 | ILing-2.6-flash | 6.0% | 2.1% | $0.010 | 262K |
| 4 | 5.9% | 1.7% | $0.300 | 1.0M | |
| 5 | 5.6% | 2.0% | $0.100 | 1.0M | |
| 6 | 4.1% | 8.8% | $0.500 | 1.0M | |
| 7 | 3.7% | 3.8% | $0.100 | 1.1M | |
| 8 | 3.1% | 1.3% | $0.070 | 262K | |
| 9 | 3.0% | 1.1% | $1.00 | 200K | |
| 10 | 2.9% | 0.90% | $0.250 | 1.0M |
Reading the two columns together: Gemini 3 Flash Preview takes a noticeably larger share of tokens than of requests — meaning it is being used for the longer, heavier summarization jobs rather than quick one-shot calls.
gpt-oss-120b is the most-used model for summarization on OpenRouter, handling 15.9% of requests classified as summarization over the last 7 days. It is followed by DeepSeek V4 Flash 0423 (9.1%) and Ling-2.6-flash (6.0%).
No. This ranks by how often each model is actually chosen for summarization in production traffic through OpenRouter — real-world adoption, which reflects price and availability as much as capability. For capability-based rankings, see the benchmark leaderboards.
Summarization accounts for 2.7% of classified requests and 1.7% of classified tokens on OpenRouter over the trailing 7 days.
Source: OpenRouter (openrouter.ai/rankings), as of 2026-08-04. Shares are of classified, sampled traffic over a trailing 7-day window; absolute volumes are not published.