Describe your workload once — tokens per request and monthly volume — and see the exact monthly bill for every model, cheapest first. Prices are live, per-token rates from the cheapest provider serving each model.
| Model | Input $/M | Output $/M | Per request | Per month |
|---|---|---|---|---|
| ling-3.0-flash:free | Free | Free | $0 | $0 |
| laguna-s-2.1:free | Free | Free | $0 | $0 |
| laguna-xs-2.1:free | Free | Free | $0 | $0 |
| north-mini-code:free | Free | Free | $0 | $0 |
| nemotron-3.5-content-safety:free | Free | Free | $0 | $0 |
| nemotron-3-ultra-550b-a55b:free | Free | Free | $0 | $0 |
| nemotron-3-nano-omni-30b-a3b-reasoning:free | Free | Free | $0 | $0 |
| laguna-m.1:free | Free | Free | $0 | $0 |
| gemma-4-26b-a4b-it:free | Free | Free | $0 | $0 |
| gemma-4-31b-it:free | Free | Free | $0 | $0 |
| lyria-3-pro-preview | Free | Free | $0 | $0 |
| lyria-3-clip-preview | Free | Free | $0 | $0 |
| nemotron-3-super-120b-a12b:free | Free | Free | $0 | $0 |
| nemotron-3-nano-30b-a3b:free | Free | Free | $0 | $0 |
| nemotron-nano-12b-v2-vl:free | Free | Free | $0 | $0 |
| nemotron-nano-9b-v2:free | Free | Free | $0 | $0 |
| gpt-oss-20b:free | Free | Free | $0 | $0 |
| ling-2.6-flash | $0.010 | $0.030 | $0.00003 | $0.300 |
| mistral-nemo | $0.019 | $0.030 | $0.00004 | $0.435 |
| granite-4.0-h-micro | $0.017 | $0.112 | $0.00008 | $0.815 |
| l3-lunaris-8b | $0.040 | $0.050 | $0.00009 | $0.850 |
| nex-n2-mini | $0.025 | $0.100 | $0.00009 | $0.875 |
| qwen-2.5-7b-instruct | $0.040 | $0.100 | $0.00011 | $1.10 |
| gpt-oss-20b | $0.029 | $0.140 | $0.00011 | $1.14 |
| mistral-small-24b-instruct-2501 | $0.050 | $0.080 | $0.00012 | $1.15 |
| llama-3.1-8b-instruct | $0.050 | $0.080 | $0.00012 | $1.15 |
| mythomax-l2-13b | $0.060 | $0.060 | $0.00012 | $1.20 |
| nova-micro-v1 | $0.035 | $0.140 | $0.00012 | $1.22 |
| granite-4.1-8b | $0.050 | $0.100 | $0.00013 | $1.25 |
| gemma-3-4b-it | $0.050 | $0.100 | $0.00013 | $1.25 |
| command-r7b-12-2024 | $0.037 | $0.150 | $0.00013 | $1.31 |
| gpt-oss-120b | $0.037 | $0.170 | $0.00014 | $1.41 |
| llama-3.2-1b-instruct | $0.027 | $0.201 | $0.00014 | $1.41 |
| laguna-xs-2.1 | $0.060 | $0.120 | $0.00015 | $1.50 |
| gemma-3n-e4b-it | $0.060 | $0.120 | $0.00015 | $1.50 |
| gemma-3-12b-it | $0.050 | $0.150 | $0.00015 | $1.50 |
| nemotron-3-nano-30b-a3b | $0.050 | $0.200 | $0.00017 | $1.75 |
| phi-4 | $0.070 | $0.140 | $0.00017 | $1.75 |
| hy3-preview | $0.063 | $0.210 | $0.00020 | $1.99 |
| reka-edge | $0.100 | $0.100 | $0.00020 | $2.00 |
| ministral-3b-2512 | $0.100 | $0.100 | $0.00020 | $2.00 |
| nova-lite-v1 | $0.060 | $0.240 | $0.00021 | $2.10 |
| qwen3.5-9b | $0.100 | $0.150 | $0.00022 | $2.25 |
| qwen3.5-flash-02-23 | $0.065 | $0.260 | $0.00023 | $2.27 |
| qwen3-coder-30b-a3b-instruct | $0.070 | $0.270 | $0.00024 | $2.40 |
| llama-3.2-3b-instruct | $0.050 | $0.330 | $0.00024 | $2.40 |
| deepseek-v4-flash | $0.098 | $0.196 | $0.00024 | $2.45 |
| laguna-s-2.1 | $0.100 | $0.200 | $0.00025 | $2.50 |
| ui-tars-1.5-7b | $0.100 | $0.200 | $0.00025 | $2.50 |
| reka-flash-3 | $0.100 | $0.200 | $0.00025 | $2.50 |
| qwen3-32b | $0.080 | $0.280 | $0.00026 | $2.60 |
| seed-1.6-flash | $0.075 | $0.300 | $0.00026 | $2.63 |
| gpt-oss-safeguard-20b | $0.075 | $0.300 | $0.00026 | $2.63 |
| gpt-5-nano | $0.050 | $0.400 | $0.00028 | $2.75 |
| glm-4.7-flash | $0.060 | $0.400 | $0.00029 | $2.90 |
| step-3.5-flash | $0.100 | $0.300 | $0.00030 | $3.00 |
| ministral-8b-2512 | $0.150 | $0.150 | $0.00030 | $3.00 |
| voxtral-small-24b-2507 | $0.100 | $0.300 | $0.00030 | $3.00 |
| qwen3-30b-a3b-instruct-2507 | $0.100 | $0.300 | $0.00030 | $3.00 |
| mistral-small-3.2-24b-instruct | $0.100 | $0.300 | $0.00030 | $3.00 |
| llama-4-scout | $0.100 | $0.300 | $0.00030 | $3.00 |
| nemotron-3-super-120b-a12b | $0.080 | $0.450 | $0.00034 | $3.45 |
| gemma-3-27b-it | $0.080 | $0.450 | $0.00034 | $3.45 |
| mimo-v2.5 | $0.140 | $0.280 | $0.00035 | $3.50 |
| seed-2.0-mini | $0.100 | $0.400 | $0.00035 | $3.50 |
| gemini-2.5-flash-lite | $0.100 | $0.400 | $0.00035 | $3.50 |
| gpt-4.1-nano | $0.100 | $0.400 | $0.00035 | $3.50 |
| gemma-4-26b-a4b-it | $0.120 | $0.350 | $0.00036 | $3.55 |
| llama-guard-4-12b | $0.180 | $0.180 | $0.00036 | $3.60 |
| qwen3-vl-32b-instruct | $0.104 | $0.416 | $0.00036 | $3.64 |
| hermes-4-70b | $0.130 | $0.400 | $0.00040 | $3.95 |
| llama-3.3-70b-instruct | $0.130 | $0.400 | $0.00040 | $3.95 |
| ministral-14b-2512 | $0.200 | $0.200 | $0.00040 | $4.00 |
| qwen3-vl-8b-instruct | $0.117 | $0.455 | $0.00040 | $4.03 |
| qwen3-8b | $0.117 | $0.455 | $0.00040 | $4.03 |
| gemma-4-31b-it | $0.140 | $0.400 | $0.00041 | $4.10 |
| qwen3-235b-a22b-2507 | $0.090 | $0.550 | $0.00041 | $4.10 |
| ring-2.6-1t | $0.075 | $0.625 | $0.00042 | $4.25 |
| ling-2.6-1t | $0.075 | $0.625 | $0.00042 | $4.25 |
| qwen3-30b-a3b | $0.130 | $0.520 | $0.00046 | $4.55 |
| hy3 | $0.132 | $0.528 | $0.00046 | $4.62 |
| olmo-3-32b-think | $0.150 | $0.500 | $0.00047 | $4.75 |
| hunyuan-a13b-instruct | $0.140 | $0.570 | $0.00049 | $4.95 |
| laguna-m.1 | $0.200 | $0.400 | $0.00050 | $5.00 |
| kat-coder-air-v2.5 | $0.150 | $0.600 | $0.00052 | $5.25 |
| mistral-small-2603 | $0.150 | $0.600 | $0.00052 | $5.25 |
| solar-pro-3 | $0.150 | $0.600 | $0.00052 | $5.25 |
| qwen3-vl-30b-a3b-instruct | $0.150 | $0.600 | $0.00052 | $5.25 |
| gpt-4o-mini-search-preview | $0.150 | $0.600 | $0.00052 | $5.25 |
| command-r-08-2024 | $0.150 | $0.600 | $0.00052 | $5.25 |
| gpt-4o-mini | $0.150 | $0.600 | $0.00052 | $5.25 |
| gpt-4o-mini-2024-07-18 | $0.150 | $0.600 | $0.00052 | $5.25 |
| qwen3-next-80b-a3b-thinking | $0.098 | $0.780 | $0.00054 | $5.36 |
| qwen3-coder-next | $0.110 | $0.800 | $0.00056 | $5.65 |
| mistral-saba | $0.200 | $0.600 | $0.00060 | $6.00 |
| deepseek-v3.2 | $0.269 | $0.400 | $0.00060 | $6.04 |
| deepseek-v3.2-exp | $0.270 | $0.410 | $0.00061 | $6.10 |
| glm-4.5-air | $0.130 | $0.850 | $0.00062 | $6.20 |
| rocinante-12b | $0.250 | $0.500 | $0.00063 | $6.25 |
| minimax-m2.5 | $0.150 | $0.900 | $0.00068 | $6.75 |
Multiply your input tokens per request by the model's input price, add output tokens times the output price, divide by one million, then multiply by monthly request volume. This calculator does exactly that against live prices for every tracked model.
A rule of thumb: 1,000 tokens is about 750 words. A chat message with a system prompt commonly runs 500–3,000 input tokens; responses run 100–1,000 output tokens. Code and RAG contexts run much larger.
Providers price the same open-weight model differently. modelgrep lists each model at the cheapest live provider rate via OpenRouter, so your effective cost can be lower than a lab's list price.