Skip to content

Simple, transparent pricing

We don’t do hidden markup. Our cost and the official rate are shown side by side, on every row, all the time — including cached input, which most providers bury in a footnote.

Prices updated 5 days ago · See full pricing →

Kimi K3New

Moonshot AI · text

Paid credits
Save 40%
Input
$3.00/M$1.85/M
Output
$15.00/M$9.00/M
Cached input
$0.37/M
Kimi K2.6New

Moonshot AI · text

Paid credits
Save 42%
Input
$0.95/M$0.55/M
Output
$4.00/M$2.30/M
Cached input
$0.11/M
GLM-5.2 AirUnavailable

Z.ai · text

Gift credits OK
Save 43%
Input
$0.28/M$0.17/M
Output
$1.10/M$0.62/M
Cached input
$0.034/M
GLM-5.2New

Z.ai · text

Paid credits
Save 42%
Input
$1.40/M$0.82/M
Output
$4.40/M$2.55/M
Cached input
$0.16/M
DeepSeek V4 Pro

DeepSeek · text

Paid credits
Save 37%
Input
$0.43/M$0.28/M
Output
$0.87/M$0.55/M
Cached input
$0.055/M
DeepSeek V4 Flash

DeepSeek · text

Gift credits OK
Save 36%
Input
$0.14/M$0.090/M
Output
$0.28/M$0.18/M
Cached input
$0.018/M
Qwen3 Max InstructUnavailable

Alibaba · text

Paid credits
Save 43%
Input
$0.85/M$0.49/M
Output
$3.40/M$1.95/M
Cached input
$0.10/M
LTX-2.5Unavailable

Lightricks · video

Gift credits OK
Save 40%
Rate
$0.040/sec$0.024/sec
MiniMax M2Unavailable

MiniMax · text

Gift credits OK
Save 40%
Input
$0.30/M$0.19/M
Output
$1.20/M$0.72/M
Cached input
$0.038/M
FLUX.2 [schnell]Unavailable

Black Forest Labs · image

Gift credits OK
Save 40%
Rate
$0.0030/img$0.0018/img
Qwen3 235B A22BUnavailable

Alibaba · text

Gift credits OK
Save 42%
Input
$0.22/M$0.13/M
Output
$0.88/M$0.51/M
Cached input
$0.026/M
Mistral Large 3Unavailable

Mistral AI · text

Paid credits
Save 40%
Input
$2.00/M$1.20/M
Output
$6.00/M$3.60/M
Cached input
$0.24/M
HunyuanVideo 1.5Unavailable

Tencent · video

Gift credits OK
Save 40%
Rate
$0.045/sec$0.027/sec
GPT-OSS 120BUnavailable

OpenAI · text

Gift credits OK
Save 40%
Input
$0.15/M$0.090/M
Output
$0.60/M$0.36/M
Cached input
$0.018/M
Qwen-Image

Alibaba · image

Gift credits OK
Save 40%
Rate
$0.0200/img$0.0120/img
Wan 2.2 T2V A14BUnavailable

Alibaba · video

Paid credits
Save 42%
Rate
$0.050/sec$0.029/sec
Wan 2.2 I2V A14BUnavailable

Alibaba · video

Paid credits
Save 42%
Rate
$0.055/sec$0.032/sec
Qwen3 Embedding 8BNew

Alibaba · embedding

Gift credits OK
Save 40%
Input
$0.050/M$0.030/M
HiDream-I1Unavailable

HiDream · image

Paid credits
Save 40%
Rate
$0.0250/img$0.0150/img
Llama 4 MaverickUnavailable

Meta · text

Gift credits OK
Save 39%
Input
$0.22/M$0.14/M
Output
$0.85/M$0.52/M
Cached input
$0.028/M
Mochi 1Unavailable

Genmo · video

Gift credits OK
Save 40%
Rate
$0.035/sec$0.021/sec
Stable Diffusion 3.5 LargeUnavailable

Stability AI · image

Gift credits OK
Save 40%
Rate
$0.0350/img$0.0210/img
CogVideoX-5BUnavailable

Zhipu AI · video

Gift credits OK
Save 43%
Rate
$0.030/sec$0.017/sec
BGE-M3Unavailable

BAAI · embedding

Gift credits OK
Save 40%
Input
$0.020/M$0.012/M

Text models priced per 1M tokens · video per second of output · image per generated image.

Pay as you go

No minimum spend. You only pay for what you actually use, billed to the cent.

No lock-in

No contracts. Cancel or pause anytime, no penalty, and export your logs on the way out.

Volume pricing available

Spending $1,000+/month? Get in touch for custom rates.

Learn more

What would you actually pay?

Per-million-token rates stay abstract until they turn into a monthly number you can hold next to your current bill.

70% in · 30% out
At official pricing
$93/mo
At our pricing
$54/mo
You’d save
$4042%

Estimate only, and it assumes no prompt caching. If you run agent loops with a large repeated system prompt, your real bill lands materially below the figure above — cached input on Kimi K2.6 is billed at $0.110/M.

Questions people actually ask

What is AI Token Router?

AI Token Router is an OpenAI-compatible API gateway for open-weight AI models. Point your existing OpenAI SDK at one endpoint and you get Kimi K3, GLM-5.2, DeepSeek V4 and open-source video models like Wan 2.2 and LTX-2.5, at 30-50% below official pricing. One key, one bill, and the official rate printed next to ours on every row.

How can you be 30-50% cheaper than official pricing?

We buy inference capacity in bulk, run it at high utilisation, and pass most of that discount through instead of keeping it as margin. There is no platform fee and no surcharge on top of the per-token rate you see. It is also why we only carry open-weight models: their licenses permit third-party hosting, so we control the cost base rather than reselling someone else's API at a markup.

Do I have to rewrite my code to switch?

No. The API is OpenAI-compatible, so if you already use the OpenAI SDK -- or any framework built on it, which is most of them -- switching is a one-line change to your base URL. Request shapes, streaming frames, the usage block and the error envelope all match what your code already handles.

Which models do you support?

Open-weight text models including Kimi K3 and K2.6, GLM-5.2 and GLM-5.2 Air, DeepSeek V4 Pro and Flash, Qwen3, MiniMax M2, Llama 4, Mistral Large 3 and GPT-OSS. On video: Wan 2.2 text-to-video and image-to-video, LTX-2.5, HunyuanVideo 1.5, CogVideoX and Mochi 1. On images: FLUX.2 schnell, Qwen-Image, Stable Diffusion 3.5 and HiDream-I1. New open-weight releases go live within 24 hours at a published price.

How does cached-token pricing work?

When you resend the same prefix -- a long system prompt, a fixed tool schema, a document you keep querying -- it is served from the model's KV cache and billed at the cached-input rate, typically 80-90% below the standard input rate. It happens automatically with no header to set, and a cache entry lives about five minutes after its last use. For agent loops this is usually the largest line on the bill, which is why the cached rate has its own column on the pricing table instead of a footnote.

Are there hidden fees?

No. The per-model rate is the whole price. No platform fee, no minimum spend, no card-processing surcharge at checkout, and no charge for failed requests. Worth checking when you compare: OpenRouter adds roughly 5.5% on credit purchases, which does not appear on their per-model pricing table.

What stops a runaway agent from running up a huge bill?

Your prepaid balance is a hard ceiling, enforced by the gateway on every request — a runaway agent stops at what you have already paid, never at an invoice. Unlike a configured cap, it cannot be misconfigured. Automatic recharge is off by default, and when enabled it is bounded too: a fixed amount, capped at three recharges in any 24 hours, switched off after three consecutive declines. Per-key limits are on the roadmap.

What happens if a model provider has an outage or rate-limits me?

Eligible requests fall back to another available model rather than failing outright, and the fallback is recorded in your request log so you can see when it happened and what it cost. You can also pin a request to a single model when substitution would be wrong for that workload.

How do you handle my prompts and data?

We do not train on your prompts or completions, ever. By default request bodies are retained for 30 days so you can debug from the request log, and you can turn that off in settings for zero data retention, where prompts and completions are discarded as soon as the response is returned. Usage metadata -- model, token counts, latency, cost -- is always kept, because that is what the dashboard and your invoice are built from.

Can I bring my own provider API key?

Yes. If you already have a contract with a provider, connect that key and route through us anyway. You keep your negotiated rate and still get routing, request logs, spending caps and per-key cost attribution. There is no charge for BYOK traffic.

How is this different from OpenRouter?

OpenRouter is a router across 500+ models including closed frontier models like GPT, Claude and Gemini, with multi-provider failover. If you need those, OpenRouter is the better fit and we say so on our comparison page. We are narrower on purpose: open-weight models only — 7 callable today out of 24 catalogued, and the catalogue marks which is which rather than failing at call time — with no platform fee, published cached-input rates, and open video generation treated as a first-class category rather than something we do not carry.

Why only open-weight models?

Because reselling closed-weight model access frequently violates the origin provider's terms of service, and that is a risk we will not take on or expose customers to. Open licenses permit third-party hosting, which is what lets us run the inference ourselves, control the price, and publish it. It is also why we do not serve Sora 2, Veo 3.1, Runway Gen-4, Kling, or Wan 2.5/2.6 -- the last of which is API-only, with weights never released, despite the open Wan 2.1/2.2 lineage.

How quickly can I make my first call?

About twenty seconds. Sign up with GitHub, Google or an email and password, and your API key is shown on the dashboard immediately along with a copy-runnable snippet. New accounts get $5 in free credits, so the first calls cost nothing and no card is required. Every model page also carries a snippet with the model id already filled in.

Do I need a credit card to start?

No. Sign up with GitHub or Google -- no password to set, nothing to leak -- and you get $5 in credits plus 2M free tokens on any open-weight model. You only add a payment method when the free credit runs out, and there is no subscription -- it is pay as you go with no minimum spend, cancel anytime.

Do you offer volume or enterprise pricing?

Yes. Accounts spending $1,000+ per month qualify for custom rates, with deeper discounts, priority support and dedicated SLAs at higher tiers. We reply within 24 hours.

Get $10 in free credits

No credit card required. Cancel anytime.