Provider data is temporarily unavailable — check back shortly.
Every inference provider serving the tracked catalog, ranked by how many models each hosts, with the strongest models it serves. Per-model speed, uptime and price for each provider live on the individual model pages; compare total costs with the pricing calculator.
It depends on the model and the metric. Frontier models (GPT, Claude, Gemini) are served by their labs plus resellers; open-weight models are served by many providers at very different speeds and prices. modelgrep lists each model at its best live provider rate, and every model page breaks pricing and uptime down per provider.
Providers run different hardware, quantizations and margins. A 70B open-weight model can stream at 40 tokens/sec on commodity GPUs or 400+ on specialized accelerators, at prices that vary just as widely. That spread is why per-provider comparison matters more than a single list price.
No — aggregators like OpenRouter expose most of these providers behind one OpenAI-compatible API and key, routing each request to the provider you pick (or the cheapest/fastest available). That's also where modelgrep's live pricing and performance data comes from.