36–42% below official pricing on the rate card, 65–68% on 400m input tokens a month with a 60% repeated prefix, and 8m output. Every ranking below reads from the same rate table the API bills from — price, discount, context window and licence, not a quality score we have not measured.
Every callable model beats its own publisher by a wide, comfortable margin on a plain one-shot request. Kimi K2.6 carries the deepest discount in this set at 42%, against 36% for DeepSeek V4 Flash. Figures are the rate-card discount, with no caching assumed.Computed from our published rates against the official rate.
Agent-workload discount, ranked
The discount widens for every model once a repeated prompt prefix is billed at the cached rate, and the ranking order is not identical to the rate-card version. Kimi K2.6 carries the deepest discount in this set at 68%, against 65% for DeepSeek V4 Flash. Figures assume 400M input tokens a month with a 60% repeated prefix, and 8M output.Computed from our published rates against the official rate, on the stated workload.
Price rankings, per token
Input rate, ranked low to high
Input rates cluster more tightly than output rates do across this catalogue. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.
Output rate, ranked low to high
This is where the models actually separate, and it is why a workload's output share decides more of the bill than the model choice alone. On output, DeepSeek V4 Flash is the lowest at $0.18 per 1M tokens and Kimi K3 the highest at $9.00 per 1M tokens, a 50x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.
Cached input rate, ranked low to high
This is the rate a repeated agent or chat prefix is actually billed at, and it is a fraction of every model's standard input rate. On cached input, DeepSeek V4 Flash is the lowest at $0.018 per 1M tokens and Kimi K3 the highest at $0.37 per 1M tokens, a 20.6x spread.Our published rates, read from the same table the API bills from.
Context window ranking
Across all 24 catalogued models, callable and not — a wider window is a real property of the model, independent of whether an upstream currently serves it.
Context window, ranked widest to narrowest
The widest windows in the catalogue belong to models not yet callable, which is worth knowing before architecting around one that is not being served today. Llama 4 Maverick leads at 1M, against 8K for BGE-M3.Read from the same catalogue every price on this site comes from.
Media model rankings
Honestly: none of the video models below are callable today, and only Qwen-Image is callable among the image models. The prices are real and set; the availability is not there yet. See video generation for the full picture.
Video generation, per second, ranked (none callable today)
Every video model in the catalogue is priced and none is currently servable — the ranking is real, the availability is not. On input, CogVideoX-5B is the lowest at $0.017 per second and Wan 2.2 I2V A14B the highest at $0.032 per second, a 1.9x spread. Every rate here runs 40–43% below the model publisher's own.Our published rates, read from the same table the API bills from.
Image generation, per image, ranked
Only one of these four is callable today; the rest are priced and catalogued ahead of an upstream serving them. On input, FLUX.2 [schnell] is the lowest at $0.0018 per image and Stable Diffusion 3.5 Large the highest at $0.0210 per image, a 11.7x spread. That is 40% below the model publisher's own rate.Our published rates, read from the same table the API bills from.
Embedding model ranking
Embedding rate, ranked
Embedding models have no output side to price -- the input rate is the whole bill. On input, BGE-M3 is the lowest at $0.012 per 1M tokens and Qwen3 Embedding 8B the highest at $0.030 per 1M tokens, a 2.5x spread. That is 40% below the model publisher's own rate.Our published rates, read from the same table the API bills from.
Licence family distribution
Catalogue by licence family
24 models
Apache 2.011 models46%
MIT7 models29%
Modified MIT2 models8%
Llama 4 Community1 model4%
LTXV Open Weights1 model4%
Stability Community1 model4%
Tencent Hunyuan Community1 model4%
Apache 2.0 and MIT together account for most of the catalogue -- the two licence families with no per-model ambiguity to read.Computed from the licence field on every catalogued model, reconciled 2026-09-02T14:00:00Z. See /licences for the full reference.
Segment counts read from the licence reference at build time; they are not independently maintained here.
Further reading
The numbers above, argued out in full and cited against outside reporting.
Three realistic budget tiers, worked by hand against this catalogue's own rate table, at one stated request shape -- how many requests and tokens $10, $50 and $250 a month actually buys on three callable models.
Closed frontier models measurably lead reasoning-heavy benchmarks as of September 2026. Where that lead and a simpler operational model are worth the higher price -- and why our catalogue is not the answer for that reader.
Where the self-host breakeven actually sits, what self-hosting really costs once engineering time is priced in, and the honest cases where self-hosting wins -- this is not a blanket argument for a managed API.
Why an agent's bill doesn't look like a chat bill -- a closed-frontier full-day usage pattern reported near $594/month, agent loops burning 5-30x an equivalent chat interaction, and one study's finding that 59.4% of an agent's tokens go to review, not writing.
Questions about these rankings
Do these rankings measure model quality?
No. Every ranking on this page measures something our own rate table actually contains: price, discount against the publisher's rate, context window size, or licence family. We have not run a benchmark on any model here, and we do not publish a quality or reasoning ranking — see /press for what we do not do.
Why do some rankings include models that aren't callable?
17 of 24 catalogued models are not currently served by any configured upstream. They are priced at their intended rate, and a context-window or licence ranking is still a true fact about them even though a request would return no available channel today. Every figure that includes one says so.
How often do these numbers change?
Whenever the rate table changes. These figures were last reconciled with upstream pricing on 2026-09-02T14:00:00Z, the same date the rest of the site's pricing reflects — there is one rate table, and every page, including this one, reads it live.