modelgrep

Token counter & cost calculator

Match · Updated August 2026

Paste any prompt to estimate its token count and see what it costs on 321 models at once — input and output priced separately, with the models whose context window can't hold it marked. Free, no signup, no API key.

321models priced
nonesignup
Estimated input
100
tokens · 434 chars · 79 words
Estimated, and it has to be — every model family tokenizes differently. Within ~6% on English prose; code and JSON read high.
500

Output is usually the expensive half — most models charge 2–5× more for it than for input.

Show cost

100 input + 500 output tokens fits 321 of 321 models. The cheapest that fits is Ling-2.6-flash at $0.0160 per 1,000 requests.

ModelInput /MOutput /MContextCost per 1,000 requests
north-mini-code:freeFreeFree256K$0
nemotron-3-nano-omni-30b-a3b-reasoning:freeFreeFree256K$0
gpt-oss-20b:freeFreeFree131K$0
ling-3.0-flash:freeFreeFree262K$0
laguna-s-2.1:freeFreeFree262K$0
laguna-xs-2.1:freeFreeFree262K$0
nemotron-3.5-content-safety:freeFreeFree128K$0
nemotron-3-ultra-550b-a55b:freeFreeFree1M$0
gemma-4-26b-a4b-it:freeFreeFree262K$0
gemma-4-31b-it:freeFreeFree262K$0
lyria-3-pro-previewFreeFree1.0M$0
lyria-3-clip-previewFreeFree1.0M$0
nemotron-3-super-120b-a12b:freeFreeFree262K$0
nemotron-3-nano-30b-a3b:freeFreeFree256K$0
nemotron-nano-12b-v2-vl:freeFreeFree128K$0
nemotron-nano-9b-v2:freeFreeFree128K$0
ling-2.6-flash$0.010$0.030262K$0.0160
mistral-nemo$0.019$0.030131K$0.0169
l3-lunaris-8b$0.040$0.0508K$0.0290
mistral-small-24b-instruct-2501$0.050$0.08033K$0.0450
llama-3.1-8b-instruct$0.050$0.080131K$0.0450
nex-n2-mini$0.025$0.100262K$0.0525
granite-4.1-8b$0.050$0.100131K$0.0550
gemma-3-4b-it$0.050$0.100131K$0.0550
granite-4.0-h-micro$0.017$0.112131K$0.0577
reka-edge$0.100$0.10016K$0.0600
ministral-3b-2512$0.100$0.100131K$0.0600
mythomax-l2-13b$0.080$0.1108K$0.0630
laguna-xs-2.1$0.060$0.120262K$0.0660
gemma-3n-e4b-it$0.060$0.12033K$0.0660
gpt-oss-20b$0.030$0.130131K$0.0680
qwen3.7-flash$0.030$0.1301M$0.0680
nova-micro-v1$0.035$0.140128K$0.0735
phi-4$0.070$0.14016K$0.0770
command-r7b-12-2024$0.037$0.150128K$0.0787
gemma-3-12b-it$0.050$0.150131K$0.0800
qwen3.5-9b$0.100$0.150262K$0.0850
gpt-oss-120b$0.037$0.170131K$0.0887
ministral-8b-2512$0.150$0.150262K$0.0900
deepseek-v4-flash-0731$0.090$0.1801.0M$0.0990
laguna-s-2.1$0.090$0.1801.0M$0.0990
qwen3-30b-a3b-instruct-2507$0.048$0.193262K$0.1014
llama-3.2-1b-instruct$0.027$0.20160K$0.1032
nemotron-3-nano-30b-a3b$0.050$0.200262K$0.1050
llama-guard-4-12b$0.180$0.1801.0M$0.1080
reka-flash-3$0.100$0.20066K$0.1100
ui-tars-1.5-7b$0.100$0.200128K$0.1100
qwen-2.5-7b-instruct$0.100$0.20033K$0.1100
hy3-preview$0.063$0.210262K$0.1113
ministral-14b-2512$0.200$0.200262K$0.1200
nova-lite-v1$0.060$0.240300K$0.1260
mistral-small-3.2-24b-instruct$0.094$0.250256K$0.1344
qwen3.5-flash-02-23$0.065$0.2601M$0.1365
qwen3-coder-30b-a3b-instruct$0.070$0.270262K$0.1420
qwen3-32b$0.080$0.280131K$0.1480
deepseek-v4-flash$0.140$0.2801.0M$0.1540
mimo-v2.5$0.140$0.2801.1M$0.1540
seed-1.6-flash$0.075$0.300262K$0.1575
gpt-oss-safeguard-20b$0.075$0.300131K$0.1575
step-3.5-flash$0.100$0.300262K$0.1600
llama-4-scout$0.100$0.3001.3M$0.1600
voxtral-small-24b-2507$0.100$0.30032K$0.1600
llama-3.3-70b-instruct$0.100$0.320131K$0.1700
llama-3.2-3b-instruct$0.050$0.330131K$0.1700
gemma-4-26b-a4b-it$0.070$0.340262K$0.1770
gemma-4-31b-it$0.100$0.340262K$0.1800
gpt-5-nano$0.050$0.400400K$0.2050
glm-4.7-flash$0.060$0.400203K$0.2060
nemotron-3-super-120b-a12b$0.085$0.4001M$0.2085
gpt-4.1-nano$0.100$0.4001.0M$0.2100
gemini-2.5-flash-lite$0.100$0.4001.0M$0.2100
seed-2.0-mini$0.100$0.400262K$0.2100
hermes-4-70b$0.130$0.400131K$0.2130
qwen3-vl-32b-instruct$0.104$0.416131K$0.2184
deepseek-v3.2$0.269$0.400164K$0.2269
deepseek-v3.2-exp$0.270$0.410164K$0.2320
gemma-3-27b-it$0.080$0.450262K$0.2330
qwen-2.5-72b-instruct$0.360$0.40033K$0.2360
qwen3-vl-8b-instruct$0.117$0.455262K$0.2392
qwen3-8b$0.117$0.455131K$0.2392
unslopnemo-12b$0.400$0.4001.0M$0.2400
llama-3.1-70b-instruct$0.400$0.400131K$0.2400
qwen3-30b-a3b$0.120$0.500131K$0.2620
olmo-3-32b-think$0.150$0.50066K$0.2650
rocinante-12b$0.250$0.50066K$0.2750
hy3$0.132$0.528262K$0.2772
cydonia-24b-v4.1$0.300$0.500131K$0.2800
hunyuan-a13b-instruct$0.140$0.570131K$0.2990
gpt-5.6-luna$0.100$0.6001.1M$0.3100
gpt-5.6-luna-pro$0.100$0.6001.1M$0.3100
mistral-small-3.1-24b-instruct$0.351$0.555128K$0.3126
qwen3-235b-a22b-2507$0.149$0.598262K$0.3140
solar-pro-3$0.150$0.600128K$0.3150
qwen3-vl-30b-a3b-instruct$0.150$0.600262K$0.3150
gpt-4o-mini$0.150$0.600128K$0.3150
gpt-4o-mini-2024-07-18$0.150$0.600128K$0.3150
kat-coder-air-v2.5$0.150$0.600256K$0.3150
mistral-small-2603$0.150$0.600262K$0.3150
command-r-08-2024$0.150$0.600128K$0.3150
mistral-saba$0.200$0.60033K$0.3200
ring-2.6-1t$0.075$0.625262K$0.3200
ling-2.6-1t$0.075$0.625262K$0.3200
remm-slerp-l2-13b$0.450$0.6506K$0.3700
wizardlm-2-8x22b$0.620$0.62066K$0.3720
gemma-2-27b-it$0.650$0.6508K$0.3900
mercury-2$0.250$0.750128K$0.4000
qwen2.5-vl-72b-instruct$0.250$0.750128K$0.4000
qwen3-coder-next$0.120$0.800262K$0.4120
qwen-plus-2025-07-28$0.260$0.7801M$0.4160
qwen-plus$0.260$0.7801M$0.4160
llama-4-maverick$0.200$0.8001.0M$0.4200
hermes-3-llama-3.1-70b$0.700$0.700131K$0.4200
weaver$0.500$0.7508K$0.4250
glm-4.5-air$0.130$0.850131K$0.4380
l3.3-euryale-70b$0.650$0.750131K$0.4400
trinity-large-thinking$0.220$0.850262K$0.4470
skyfall-36b-v2$0.550$0.80033K$0.4550
minimax-m2.5$0.150$0.900205K$0.4650
dolphin-mistral-24b-venice-edition$0.200$0.900128K$0.4700
qwen3-14b$0.228$0.910131K$0.4778
deepseek-v4-pro$0.435$0.8701.0M$0.4785
mimo-v2.5-pro$0.435$0.8701.1M$0.4785
deepseek-r1-distill-llama-70b$0.800$0.8008K$0.4800
glm-4.6v$0.300$0.900131K$0.4800
codestral-2508$0.300$0.900256K$0.4800
deepseek-chat-v3.1$0.250$0.950164K$0.5000
qwen3-coder-flash$0.195$0.9751M$0.5070
l3.1-euryale-70b$0.850$0.850131K$0.5100
qwen3.6-35b-a3b$0.140$1.00262K$0.5140
qwen3.5-35b-a3b$0.140$1.00262K$0.5140
nex-n2-pro$0.250$1.00262K$0.5250
deepseek-v3.1-terminus$0.270$1.00164K$0.5270
qwen3-coder$0.300$1.00262K$0.5300
minimax-m2$0.255$1.02205K$0.5355
deepseek-chat$0.257$1.03164K$0.5401
qwen3-next-80b-a3b-instruct$0.090$1.10262K$0.5590
qwen-2.5-coder-32b-instruct$0.660$1.0033K$0.5660
minimax-m2.7$0.270$1.08205K$0.5670
minimax-01$0.200$1.101.0M$0.5700
qwen3.6-flash$0.188$1.131M$0.5813
deepseek-chat-v3-0324$0.270$1.12164K$0.5870
step-3.7-flash$0.200$1.15262K$0.5950
sonar$1.00$1.00127K$0.6000
hermes-3-llama-3.1-405b$1.00$1.00131K$0.6000
qwen3-next-80b-a3b-thinking$0.150$1.20262K$0.6150
minimax-m3$0.300$1.201.0M$0.6300
kat-coder-pro-v2$0.300$1.20262K$0.6300
longcat-2.0$0.300$1.201.0M$0.6300
minimax-m2.1$0.300$1.20205K$0.6300
minimax-m2-her$0.300$1.2066K$0.6300
qwen-plus-2025-07-28:thinking$0.400$1.201M$0.6400
gpt-5.4-nano$0.200$1.25400K$0.6450
inkling-small$0.500$1.20524K$0.6500
claude-3-haiku$0.250$1.25200K$0.6500
ernie-4.5-vl-424b-a47b$0.420$1.25123K$0.6670
qwen3.7-plus$0.320$1.281M$0.6720
virtuoso-large$0.750$1.20131K$0.6750
morph-v3-fast$0.800$1.2082K$0.6800
relace-apply-3$0.850$1.25256K$0.7100
cogito-v2.1-671b$1.25$1.25128K$0.7500
perceptron-mk1$0.150$1.5033K$0.7650
aion-3.0-mini$0.700$1.40131K$0.7700
gemini-3.1-flash-lite$0.250$1.501.0M$0.7750
gemini-3.1-flash-lite-preview$0.250$1.501.0M$0.7750
gemini-3.1-flash-lite-image$0.250$1.5066K$0.7750
qwen3.5-27b$0.195$1.56262K$0.7995
gpt-3.5-turbo$0.500$1.5016K$0.8000
mistral-large-2512$0.500$1.50262K$0.8000
qwen3.5-plus-02-15$0.260$1.561M$0.8060
gpt-4.1-mini$0.400$1.601.0M$0.8400
aion-2.0$0.800$1.60131K$0.8800
aion-rp-llama-3.1-8b$0.800$1.6033K$0.8800
glm-4.7$0.400$1.75205K$0.9150
qwen3.5-plus-20260420$0.300$1.801M$0.9300
qwen3-235b-a22b$0.455$1.82131K$0.9555
glm-4.5v$0.600$1.8066K$0.9600
qwen3-vl-235b-a22b-instruct$0.210$1.90262K$0.9710
qwen3.6-plus$0.325$1.951M$1.01
gpt-5.1-codex-mini$0.250$2.00400K$1.03
gpt-5-mini$0.250$2.00400K$1.03
seed-2.0-lite$0.250$2.00262K$1.03
seed-1.6$0.250$2.00262K$1.03
morph-v3-large$0.900$1.90262K$1.04
mistral-medium-3.1$0.400$2.00131K$1.04
mistral-medium-3$0.400$2.00131K$1.04
glm-4.6$0.500$2.00205K$1.05
qwen3.5-122b-a10b$0.260$2.08262K$1.07
qwen3-vl-8b-thinking$0.180$2.10131K$1.07
grok-build-0.1$1.00$2.00256K$1.10
gpt-3.5-turbo-0613$1.00$2.004K$1.10
deepseek-r1-0528$0.500$2.15164K$1.13
gpt-3.5-turbo-instruct$1.50$2.004K$1.15
minimax-m1$0.550$2.201M$1.16
glm-4.5$0.600$2.20131K$1.16
qwen3-235b-a22b-thinking-2507$0.230$2.30262K$1.17
kimi-k2$0.570$2.30131K$1.21
qwen3.5-397b-a17b$0.390$2.34262K$1.21
qwen3-vl-30b-a3b-thinking$0.200$2.40262K$1.22
qwen3-30b-a3b-thinking-2507$0.200$2.4082K$1.22
qwen3.6-27b$0.289$2.40262K$1.23
gpt-5-image-mini$2.50$2.00400K$1.25
gpt-audio-mini$0.600$2.40128K$1.26
gemini-3.5-flash-lite$0.300$2.501.0M$1.28
gemini-2.5-flash$0.300$2.501.0M$1.28
nova-2-lite-v1$0.300$2.501M$1.28
gemini-2.5-flash-image$0.300$2.5033K$1.28
glm-5.2$0.760$2.421.0M$1.29
kimi-k2.6$0.589$2.48262K$1.30
kimi-k2-thinking$0.600$2.50262K$1.31
kimi-k2-0905$0.600$2.50262K$1.31
deepseek-r1$0.700$2.50164K$1.32
glm-5$0.950$2.55205K$1.37
grok-4.3$1.25$2.501M$1.38
grok-4.20$1.25$2.502M$1.38
grok-4.20-multi-agent$1.25$2.502M$1.38
kimi-k2.5$0.570$2.85262K$1.48
gemini-3.1-flash-image$0.500$3.00131K$1.55
gemini-3.1-flash-image-preview$0.500$3.0066K$1.55
gemini-3-flash-preview$0.500$3.001.0M$1.55
kat-coder-pro-v2.5$0.740$2.96256K$1.55
relace-search$1.00$3.00256K$1.60
hermes-4-405b$1.00$3.00131K$1.60
glm-5.1$0.966$3.04205K$1.61
nova-pro-v1$0.800$3.20300K$1.68
qwen3-coder-plus$0.650$3.251M$1.69
kimi-k2.7-code$0.730$3.50262K$1.82
nemotron-3-ultra-550b-a55b$0.600$3.60512K$1.86
qwen3-max-thinking$0.780$3.90262K$2.03
qwen3-max$0.780$3.90262K$2.03
qwen3-vl-235b-a22b-thinking$0.400$4.00131K$2.04
glm-5-turbo$1.20$4.00203K$2.12
glm-5v-turbo$1.20$4.00203K$2.12
inkling$1.00$4.051.0M$2.12
muse-spark-1.1$1.25$4.251.0M$2.25
gpt-3.5-turbo-16k$3.00$4.0016K$2.30
o4-mini-high$1.10$4.40200K$2.31
o4-mini$1.10$4.40200K$2.31
o3-mini$1.10$4.40200K$2.31
o3-mini-high$1.10$4.40200K$2.31
gpt-5.4-mini$0.750$4.50400K$2.33
qwen3.7-max$1.48$4.421M$2.36
claude-haiku-4.5$1.00$5.00200K$2.60
magnum-v4-72b$3.00$5.0016K$2.80
palmyra-x5$0.600$6.001.0M$3.06
gpt-5.6-terra$1.00$6.001.1M$3.10
gpt-5.6-terra-pro$1.00$6.001.1M$3.10
qwen3.6-max-preview$1.03$6.16262K$3.18
grok-4.5$2.00$6.00500K$3.20
mistral-large-2407$2.00$6.00131K$3.20
mixtral-8x22b-instruct$2.00$6.0066K$3.20
mistral-large$2.00$6.00128K$3.20
qwen3.8-max$2.00$6.001M$3.20
aion-3.0$3.00$6.00131K$3.30
gemini-3.6-flash$1.50$7.501.0M$3.90
mistral-medium-3-5$1.50$7.50262K$3.90
o3$2.00$8.00200K$4.20
gpt-4.1$2.00$8.001.0M$4.20
sonar-reasoning-pro$2.00$8.00128K$4.20
jamba-large-1.7$2.00$8.00256K$4.20
sonar-deep-research$2.00$8.00128K$4.20
gemini-3.5-flash$1.50$9.001.0M$4.65
gpt-5.1$1.25$10.00400K$5.13
gpt-5.1-codex$1.25$10.00400K$5.13
gpt-5$1.25$10.00400K$5.13
gemini-2.5-pro$1.25$10.001.0M$5.13
gpt-5.1-codex-max$1.25$10.00400K$5.13
gemini-2.5-pro-preview$1.25$10.001.0M$5.13
gemini-2.5-pro-preview-05-06$1.25$10.001.0M$5.13
claude-sonnet-5$2.00$10.001M$5.20
gpt-4o-2024-11-20$2.50$10.00128K$5.25
gpt-4o$2.50$10.00128K$5.25
gpt-4o-2024-08-06$2.50$10.00128K$5.25
command-a$2.50$10.00256K$5.25
gpt-audio$2.50$10.00128K$5.25
command-r-plus-08-2024$2.50$10.00128K$5.25
gpt-5-image$10.00$10.00400K$6.00
gemini-3.1-pro-preview$2.00$12.001.0M$6.20
gemini-3-pro-image$2.00$12.00131K$6.20
gemini-3.1-pro-preview-customtools$2.00$12.001.0M$6.20
gemini-3-pro-image-preview$2.00$12.0066K$6.20
nova-premier-v1$2.50$12.501M$6.50
gpt-5.3-codex$1.75$14.00400K$7.17
gpt-5.2$1.75$14.00400K$7.17
gpt-5.2-codex$1.75$14.00400K$7.17
gpt-5.3-chat$1.75$14.00128K$7.17
gpt-5.2-chat$1.75$14.00128K$7.17
gpt-5.4$2.50$15.001.1M$7.75
kimi-k3$3.00$15.001.0M$7.80
claude-sonnet-4.6$3.00$15.001M$7.80
claude-sonnet-4.5$3.00$15.001M$7.80
claude-sonnet-4$3.00$15.001M$7.80
sonar-pro$3.00$15.00200K$7.80
sonar-pro-search$3.00$15.00200K$7.80
gpt-4o-2024-05-13$5.00$15.00128K$8.00
gpt-5.4-image-2$8.00$15.00272K$8.30
claude-opus-5$5.00$25.001M$13.00
claude-opus-4.8$5.00$25.001M$13.00
claude-opus-4.7$5.00$25.001M$13.00
claude-opus-4.6$5.00$25.001M$13.00
claude-opus-4.5$5.00$25.00200K$13.00
gpt-5.6-sol$5.00$30.001.1M$15.50
gpt-5.5$5.00$30.001.1M$15.50
gpt-5.6-sol-pro$5.00$30.001.1M$15.50
fugu-ultra$5.00$30.001M$15.50
gpt-chat-latest$5.00$30.00400K$15.50
gpt-4-turbo$10.00$30.00128K$16.00
gpt-4-turbo-preview$10.00$30.00128K$16.00
claude-fable-5$10.00$50.001M$26.00
claude-opus-5-fast$10.00$50.001M$26.00
claude-opus-4.8-fast$10.00$50.001M$26.00
o1$15.00$60.00200K$31.50
gpt-4$30.00$60.008K$33.00
claude-opus-4.1$15.00$75.00200K$39.00
claude-opus-4$15.00$75.00200K$39.00
o3-pro$20.00$80.00200K$42.00
gpt-5-pro$15.00$120.00400K$61.50
claude-opus-4.7-fast$30.00$150.001M$78.00
gpt-5.2-pro$21.00$168.00400K$86.10
gpt-5.5-pro$30.00$180.001.1M$93.00
gpt-5.4-pro$30.00$180.001.1M$93.00
o1-pro$150.00$600.00200K$315

How this works

Language models don't read characters — they read tokens, chunks of roughly four characters produced by a tokenizer that has learned which character sequences occur together. Common words are a single token; rare or long ones get split into several. That is why normalization costs one token while a made-up word of the same length costs four.

The count here is an estimate, deliberately. Each model family ships its own tokenizer and they disagree on the same input — so a single "exact" number would be exact for one family and wrong for every other model on this page. The estimate lands within a few percent on ordinary prose and over-counts dense JSON, which at least errs toward over-budgeting. If you know your real count, type it into the override and every price recalculates from it.

The pricing itself is not estimated. It is the live per-million rate from the cheapest provider serving each model, refreshed hourly — the same data behind the pricing comparison and cheapest-models ranking.

Frequently asked

How many tokens is my prompt?

Paste it above and the counter estimates it as you type. As a rule of thumb, English prose runs about 0.75 words per token — roughly 4 characters — so 1,000 words is near 1,300 tokens. Code, JSON and non-Latin scripts are denser and cost more tokens per character.

Why is the count an estimate rather than exact?

Because there is no single correct answer. Every model family ships its own tokenizer and they disagree on the same text — measured against real implementations, OpenAI's own o200k and cl100k tokenizers differ from each other by up to 76% on Chinese text. An 'exact' number would be exact for at most one family and quietly wrong for every other model priced here. This estimator is tuned for English prose, where it lands within about 6%; code and JSON run roughly 15% high, and CJK higher still. It always errs upward, so you over-budget rather than under. If you know your real count, type it into the override and every price recalculates from it.

Why does output cost more than input?

Input tokens are processed in parallel in a single prefill pass, while output tokens are generated one at a time, each requiring a full forward pass through the model. That serial cost is why most providers charge 2–5× more per output token — and why response length usually drives your bill more than prompt length.

What does 'won't fit' mean?

The model's context window is smaller than your input plus expected output combined. The request would be rejected or truncated. Context has to hold the prompt, any documents, the conversation history and the response — all from the same budget.

How can I reduce token costs?

Three things move the needle most: shorten expected output (it's the expensive half), use prompt caching if you resend the same context repeatedly (often 10× cheaper on the cached portion), and drop to a smaller model for the high-volume, low-difficulty parts of your pipeline. The cost table above is sorted cheapest-first, so the tradeoff is visible directly.

Related