What would this actually save you?
Every rate here runs 36–42% below what the model’s own publisher charges, and the official rate is printed beside ours on every row so the claim can be checked rather than believed. Put your own monthly volume in below and the comparison is done on your numbers instead of ours.
An agent or chat loop resends the same system prompt on every turn. That repeated prefix is billed at the cached rate, and on most models that is a fifth of the standard input rate. Set this to zero if every request you send is different.
That is $7.00 a month, or $84.00 a year, on the same tokens through the same model.
| Model | Here | Official | Best third partyuncached |
|---|---|---|---|
| DeepSeek V4 Flash$0.090 in · $0.18 out | $12.60 | $19.60 | $14.00DeepInfra |
| DeepSeek V4 Pro$0.28 in · $0.55 out | $39.00 | $60.90 | $42.00DeepInfra |
| Kimi K2.6$0.55 in · $2.30 out | $101.00 | $175.00 | $123.00DeepInfra |
| GLM-5.2$0.82 in · $2.55 out | $133.00 | $228.00 | $157.00DeepInfra |
| Kimi K3$1.85 in · $9.00 out | $365.00 | $600.00 | $470.00DeepInfra |
An estimate, not a quote — but it is computed from the same rate table the API bills from, so the only assumptions in it are the ones you set above. Third-party figures are those providers’ own published rates; where a provider does not publish one, the cell is blank rather than guessed.
Worked already, for the eight workloads people ask about
Each row is a real token shape, not a round number. Follow one through to see the arithmetic.
| Workload | Monthly volume | Here | Official |
|---|---|---|---|
| Coding agentson DeepSeek V4 Pro | 400M in · 8M out | $116.40 | $180.96 |
| RAG and retrievalon DeepSeek V4 Flash | 400M in · 30M out | $41.40 | $64.40 |
| Chatbots and supporton DeepSeek V4 Flash | 600M in · 40M out | $61.20 | $95.20 |
| Document summarisationon DeepSeek V4 Flash | 800M in · 16M out | $74.88 | $116.48 |
| Structured data extractionon DeepSeek V4 Flash | 1,200M in · 140M out | $133.20 | $207.20 |
| Semantic search and embeddingson DeepSeek V4 Flash | 3,040M in · 0M out | $273.60 | $425.60 |
| Translation and localisationon DeepSeek V4 Flash | 420M in · 400M out | $109.80 | $170.80 |
| Content generationon DeepSeek V4 Flash | 30M in · 150M out | $29.70 | $46.20 |
Questions about the numbers
- How accurate is this estimate?
- The rates are exact — they are read from the same table the API bills from, so the per-token numbers are not approximations. What is estimated is your volume. The output is only as good as the token counts you put in, which is why the presets come from real workload shapes rather than a generic split.
- Why does the cached-input slider change the total so much?
- Because on a repeated-prefix workload the cache is most of the bill. An agent that resends a fixed system prompt on every turn is billed for that prefix at the cached rate, which on most models here is about a fifth of the standard input rate. Ignoring it makes an agent workload look far more expensive than it is, which is exactly what a calculator that omits caching does.
- Does the comparison include a platform fee?
- No, because there is not one. The per-model rate is the whole price on AI Token Router — no platform fee, no minimum spend, no card-processing surcharge, and no charge for a failed request. Some providers add a percentage on credit purchases that does not appear on their per-model table; where we have that figure it is on the comparison pages.
- Where do the third-party numbers come from?
- Each provider's own published rate card. Where a provider does not publish a per-model rate we can verify, the cell is left blank rather than filled with an estimate — a guessed competitor price is worse than no competitor price, because it is the kind of error nobody can check.
- Can I share the result?
- Yes. The numbers you set are written into the page's address, so copying the URL sends someone the same estimate. That matters more than it sounds: for anything above a hobby project, the person who decides is rarely the person who ran the calculator.
The full rate card, including cached input and the video and image models, is on the pricing page. Kimi K2.6 carries the widest gap in the callable set at $0.55 against $0.95 official — its own page shows the working.
Check it against your own bill
$5 in free credits, no card required. The balance is a hard ceiling.
Every rate above is the rate you are charged. There is no plan tier that changes it.