modelgrep is a Chatbot Arena (LMArena) alternative that ranks models by independent benchmarks instead of crowd votes — currently led by Claude Fable 5 at 59.9 on the Intelligence Index — and adds what arena votes can't tell you: live speed, latency and per-token price for every model.
| Dimension | Chatbot Arena / LMArena | modelgrep |
|---|---|---|
| Ranking signal | Crowd votes on blind chat pairs (Elo) | Independent benchmark suites (Intelligence, Coding, Agentic indexes) + Design Arena human-preference Elo for UI output |
| Speed & latency | Not measured | Live tokens/sec and time-to-first-token, refreshed hourly per provider |
| Pricing | Not shown | Per-token input/output/cache pricing on every row, plus a cost calculator |
| Use-case rankings | Category boards for chat styles | Coding, design, writing, math, RAG, SQL, agents, roleplay, local, and per-maker cuts of each |
| Coverage | Models opted into the arena | Every model served by a tracked provider — 300+ LLMs plus 1,400+ image/video/voice models |
| Data access | Web UI | Free, no-key JSON API for all benchmark, speed and price data |
It answers the same question — which model should I use? — with a different signal. Arena Elo measures which responses people prefer in blind chat; modelgrep measures benchmark performance plus the operational data (speed, latency, price) that determines whether a model works in production. Many people use both.
Arena votes reward style and formatting as much as correctness, and can't measure coding ability, tool use, latency or cost at all. Benchmark suites like GPQA and SWE-bench test verifiable tasks, and human-preference Elo still covers the subjective axis where it belongs — modelgrep uses Design Arena Elo for UI generation quality.
Yes, and more of them: coding, design/frontend, writing, math, RAG, SQL, agents, roleplay, vision, open-source, local, plus speed, latency, price and context rankings — each also sliceable per maker (e.g. best Anthropic model for coding).