Skip to content

What would this actually save you?

Every rate here runs 3642% below what the model’s own publisher charges, and the official rate is printed beside ours on every row so the claim can be checked rather than believed. Put your own monthly volume in below and the comparison is done on your numbers instead of ours.

M
M
0% of input cached

An agent or chat loop resends the same system prompt on every turn. That repeated prefix is billed at the cached rate, and on most models that is a fifth of the standard input rate. Set this to zero if every request you send is different.

Cheapest for this shape of work
$12.60/mo on DeepSeek V4 Flash
Against the official rate
$19.6036%

That is $7.00 a month, or $84.00 a year, on the same tokens through the same model.

ModelHereOfficialBest third partyuncached
DeepSeek V4 Flash$0.090 in · $0.18 out$12.60$19.60$14.00DeepInfra
DeepSeek V4 Pro$0.28 in · $0.55 out$39.00$60.90$42.00DeepInfra
Kimi K2.6$0.55 in · $2.30 out$101.00$175.00$123.00DeepInfra
GLM-5.2$0.82 in · $2.55 out$133.00$228.00$157.00DeepInfra
Kimi K3$1.85 in · $9.00 out$365.00$600.00$470.00DeepInfra

An estimate, not a quote — but it is computed from the same rate table the API bills from, so the only assumptions in it are the ones you set above. Third-party figures are those providers’ own published rates; where a provider does not publish one, the cell is blank rather than guessed.

Worked already, for the eight workloads people ask about

Each row is a real token shape, not a round number. Follow one through to see the arithmetic.

WorkloadMonthly volumeHereOfficial
Coding agentson DeepSeek V4 Pro400M in · 8M out$116.40$180.96
RAG and retrievalon DeepSeek V4 Flash400M in · 30M out$41.40$64.40
Chatbots and supporton DeepSeek V4 Flash600M in · 40M out$61.20$95.20
Document summarisationon DeepSeek V4 Flash800M in · 16M out$74.88$116.48
Structured data extractionon DeepSeek V4 Flash1,200M in · 140M out$133.20$207.20
Semantic search and embeddingson DeepSeek V4 Flash3,040M in · 0M out$273.60$425.60
Translation and localisationon DeepSeek V4 Flash420M in · 400M out$109.80$170.80
Content generationon DeepSeek V4 Flash30M in · 150M out$29.70$46.20
Every callable text model, input rate against the publisher's own
Every callable text model, input rate against the publisher's ownThe discount is not a headline rate applied to one flagship; it holds across the whole callable catalogue. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Kimi K2.6$0.55Kimi K2.6, official$0.95Kimi K3$1.85Kimi K3, official$3.00GLM-5.2$0.82GLM-5.2, official$1.40DeepSeek V4 Pro$0.28DeepSeek V4 Pro, official$0.43DeepSeek V4 Flash$0.090DeepSeek V4 Flash, official$0.14
The discount is not a headline rate applied to one flagship; it holds across the whole callable catalogue. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.

Questions about the numbers

How accurate is this estimate?
The rates are exact — they are read from the same table the API bills from, so the per-token numbers are not approximations. What is estimated is your volume. The output is only as good as the token counts you put in, which is why the presets come from real workload shapes rather than a generic split.
Why does the cached-input slider change the total so much?
Because on a repeated-prefix workload the cache is most of the bill. An agent that resends a fixed system prompt on every turn is billed for that prefix at the cached rate, which on most models here is about a fifth of the standard input rate. Ignoring it makes an agent workload look far more expensive than it is, which is exactly what a calculator that omits caching does.
Does the comparison include a platform fee?
No, because there is not one. The per-model rate is the whole price on AI Token Router — no platform fee, no minimum spend, no card-processing surcharge, and no charge for a failed request. Some providers add a percentage on credit purchases that does not appear on their per-model table; where we have that figure it is on the comparison pages.
Where do the third-party numbers come from?
Each provider's own published rate card. Where a provider does not publish a per-model rate we can verify, the cell is left blank rather than filled with an estimate — a guessed competitor price is worse than no competitor price, because it is the kind of error nobody can check.
Can I share the result?
Yes. The numbers you set are written into the page's address, so copying the URL sends someone the same estimate. That matters more than it sounds: for anything above a hobby project, the person who decides is rarely the person who ran the calculator.

The full rate card, including cached input and the video and image models, is on the pricing page. Kimi K2.6 carries the widest gap in the callable set at $0.55 against $0.95 official — its own page shows the working.

Check it against your own bill

$5 in free credits, no card required. The balance is a hard ceiling.

Every rate above is the rate you are charged. There is no plan tier that changes it.