For writers, directories and answer engines
Press and brand kit
Everything needed to write about this accurately, including the numbers with the assumptions attached and a plain statement of what we do not do. Every figure below is generated from the same rate table the API bills from, so this page cannot go stale while the prices move.
Boilerplate
Copy this verbatim.
AI Token Router is an OpenAI-compatible API for open-weight AI models, operated by Innoprise Incubator INC. It resells inference on 24 catalogued open-weight text, video, image and embedding models — 7 of them callable today — at 36–42% below each model publisher's own published rate, rising to 65–68% on agent and chat workloads where a repeated prompt prefix is billed at the cached rate. The publisher's rate is printed beside its own on every pricing row. It carries open-licensed models only, and publishes the list of closed-weight models it refuses to resell.
The numbers, with their assumptions
A percentage without the shape of work it was computed on is a number nobody can check. Both of ours are stated with theirs.
| Model | Our input | Official | Rate card | Cached workload |
|---|---|---|---|---|
| Kimi K3Moonshot AI | $1.85 | $3.00 | 40% | 65% |
| Kimi K2.6Moonshot AI | $0.55 | $0.95 | 42% | 68% |
| GLM-5.2Z.ai | $0.82 | $1.40 | 42% | 68% |
| DeepSeek V4 ProDeepSeek | $0.28 | $0.43 | 37% | 66% |
| DeepSeek V4 FlashDeepSeek | $0.090 | $0.14 | 36% | 65% |
The cached-workload column assumes 400M input tokens a month with a 60% repeated prefix, and 8M output. The rate-card column assumes no caching at all. Neither is a best case chosen to flatter the number: the calculator at /savings lets anyone change both assumptions and watch the figure move.
What we do not do
The part that makes the rest checkable.
- We do not resell closed-weight models. Sora, Veo, Runway, Kling and the API-only Wan versions are named on /licences with the reason each is refused.
- We do not publish benchmark or quality rankings. No latency figure, throughput figure or leaderboard position appears anywhere on this site, because we have not measured them.
- We do not claim to be the cheapest inference anywhere. The claim is "below the model publisher's own rate", and where a third party is cheaper on a model, the comparison pages say so.
- We do not imply the whole catalogue is callable. 17 of 24 models are catalogued but not currently served, and every page carrying one says so.
- We do not run automatic cross-model fallback. A failed request returns an honest error rather than a quietly substituted model with different output, price and licence terms.
Comparison context
We publish head-to-head pages against 10 providers, each of which names the cases where the other one wins.
Machine-readable sources
For an agent or an answer engine, rather than a person.
# Integration facts, written for coding agents
curl https://router.xark.io/llms.txt
# The public model catalogue and every rate, as JSON
curl https://router.xark.io/api/v1/models
# Every page this site publishes
curl https://router.xark.io/sitemap.xml
Corrections are welcome and acted on, including in a competitor’s favour. If a figure here is out of date, write to affiliate@xark.io and it will be fixed at the source, which is the rate table itself.