Skip to content

Computed, not asserted

Open-weight model rankings

36–42% below official pricing on the rate card, 65–68% on 400m input tokens a month with a 60% repeated prefix, and 8m output. Every ranking below reads from the same rate table the API bills from — price, discount, context window and licence, not a quality score we have not measured.

Discount rankings

Rate-card discount, ranked
Rate-card discount, rankedEvery callable model beats its own publisher by a wide, comfortable margin on a plain one-shot request. Kimi K2.6 carries the deepest discount in this set at 42%, against 36% for DeepSeek V4 Flash. Figures are the rate-card discount, with no caching assumed.Kimi K2.642%GLM-5.242%Kimi K340%Qwen-Image40%Qwen3 Embedding 8B40%DeepSeek V4 Pro37%DeepSeek V4 Flash36%
Every callable model beats its own publisher by a wide, comfortable margin on a plain one-shot request. Kimi K2.6 carries the deepest discount in this set at 42%, against 36% for DeepSeek V4 Flash. Figures are the rate-card discount, with no caching assumed.Computed from our published rates against the official rate.
Agent-workload discount, ranked
Agent-workload discount, rankedThe discount widens for every model once a repeated prompt prefix is billed at the cached rate, and the ranking order is not identical to the rate-card version. Kimi K2.6 carries the deepest discount in this set at 68%, against 65% for DeepSeek V4 Flash. Figures assume 400M input tokens a month with a 60% repeated prefix, and 8M output.Kimi K2.668%GLM-5.268%DeepSeek V4 Pro66%Kimi K365%DeepSeek V4 Flash65%
The discount widens for every model once a repeated prompt prefix is billed at the cached rate, and the ranking order is not identical to the rate-card version. Kimi K2.6 carries the deepest discount in this set at 68%, against 65% for DeepSeek V4 Flash. Figures assume 400M input tokens a month with a 60% repeated prefix, and 8M output.Computed from our published rates against the official rate, on the stated workload.

Price rankings, per token

Input rate, ranked low to high
Input rate, ranked low to highInput rates cluster more tightly than output rates do across this catalogue. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Kimi K2.6$0.55Kimi K2.6, official$0.95Kimi K3$1.85Kimi K3, official$3.00GLM-5.2$0.82GLM-5.2, official$1.40DeepSeek V4 Pro$0.28DeepSeek V4 Pro, official$0.43DeepSeek V4 Flash$0.090DeepSeek V4 Flash, official$0.14
Input rates cluster more tightly than output rates do across this catalogue. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.
Output rate, ranked low to high
Output rate, ranked low to highThis is where the models actually separate, and it is why a workload's output share decides more of the bill than the model choice alone. On output, DeepSeek V4 Flash is the lowest at $0.18 per 1M tokens and Kimi K3 the highest at $9.00 per 1M tokens, a 50x spread. Every rate here runs 36–42% below the model publisher's own.Kimi K2.6$2.30Kimi K2.6, official$4.00Kimi K3$9.00Kimi K3, official$15.00GLM-5.2$2.55GLM-5.2, official$4.40DeepSeek V4 Pro$0.55DeepSeek V4 Pro, official$0.87DeepSeek V4 Flash$0.18DeepSeek V4 Flash, official$0.28
This is where the models actually separate, and it is why a workload's output share decides more of the bill than the model choice alone. On output, DeepSeek V4 Flash is the lowest at $0.18 per 1M tokens and Kimi K3 the highest at $9.00 per 1M tokens, a 50x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.
Cached input rate, ranked low to high
Cached input rate, ranked low to highThis is the rate a repeated agent or chat prefix is actually billed at, and it is a fraction of every model's standard input rate. On cached input, DeepSeek V4 Flash is the lowest at $0.018 per 1M tokens and Kimi K3 the highest at $0.37 per 1M tokens, a 20.6x spread.Kimi K2.6$0.11Kimi K3$0.37GLM-5.2$0.16DeepSeek V4 Pro$0.055DeepSeek V4 Flash$0.018
This is the rate a repeated agent or chat prefix is actually billed at, and it is a fraction of every model's standard input rate. On cached input, DeepSeek V4 Flash is the lowest at $0.018 per 1M tokens and Kimi K3 the highest at $0.37 per 1M tokens, a 20.6x spread.Our published rates, read from the same table the API bills from.

Context window ranking

Across all 24 catalogued models, callable and not — a wider window is a real property of the model, independent of whether an upstream currently serves it.

Context window, ranked widest to narrowest
Context window, ranked widest to narrowestThe widest windows in the catalogue belong to models not yet callable, which is worth knowing before architecting around one that is not being served today. Llama 4 Maverick leads at 1M, against 8K for BGE-M3.Llama 4 Maverick1MKimi K3512KKimi K2.6256KQwen3 Max Instruct256KMistral Large 3256KGLM-5.2200KMiniMax M2200KDeepSeek V4 Pro160KDeepSeek V4 Flash160KGLM-5.2 Air128KQwen3 235B A22B128KGPT-OSS 120B128KQwen3 Embedding 8B32KBGE-M38K
The widest windows in the catalogue belong to models not yet callable, which is worth knowing before architecting around one that is not being served today. Llama 4 Maverick leads at 1M, against 8K for BGE-M3.Read from the same catalogue every price on this site comes from.

Media model rankings

Honestly: none of the video models below are callable today, and only Qwen-Image is callable among the image models. The prices are real and set; the availability is not there yet. See video generation for the full picture.

Video generation, per second, ranked (none callable today)
Video generation, per second, ranked (none callable today)Every video model in the catalogue is priced and none is currently servable — the ranking is real, the availability is not. On input, CogVideoX-5B is the lowest at $0.017 per second and Wan 2.2 I2V A14B the highest at $0.032 per second, a 1.9x spread. Every rate here runs 40–43% below the model publisher's own.Wan 2.2 T2V A14B$0.029Wan 2.2 T2V A14B, official$0.050Wan 2.2 I2V A14B$0.032Wan 2.2 I2V A14B, official$0.055LTX-2.5$0.024LTX-2.5, official$0.040HunyuanVideo 1.5$0.027HunyuanVideo 1.5, official$0.045CogVideoX-5B$0.017CogVideoX-5B, official$0.030Mochi 1$0.021Mochi 1, official$0.035
Every video model in the catalogue is priced and none is currently servable — the ranking is real, the availability is not. On input, CogVideoX-5B is the lowest at $0.017 per second and Wan 2.2 I2V A14B the highest at $0.032 per second, a 1.9x spread. Every rate here runs 40–43% below the model publisher's own.Our published rates, read from the same table the API bills from.
Image generation, per image, ranked
Image generation, per image, rankedOnly one of these four is callable today; the rest are priced and catalogued ahead of an upstream serving them. On input, FLUX.2 [schnell] is the lowest at $0.0018 per image and Stable Diffusion 3.5 Large the highest at $0.0210 per image, a 11.7x spread. That is 40% below the model publisher's own rate.FLUX.2 [schnell]$0.0018FLUX.2 [schnell], official$0.0030Qwen-Image$0.0120Qwen-Image, official$0.0200Stable Diffusion 3.5 Large$0.0210Stable Diffusion 3.5 Large, official$0.0350HiDream-I1$0.0150HiDream-I1, official$0.0250
Only one of these four is callable today; the rest are priced and catalogued ahead of an upstream serving them. On input, FLUX.2 [schnell] is the lowest at $0.0018 per image and Stable Diffusion 3.5 Large the highest at $0.0210 per image, a 11.7x spread. That is 40% below the model publisher's own rate.Our published rates, read from the same table the API bills from.

Embedding model ranking

Embedding rate, ranked
Embedding rate, rankedEmbedding models have no output side to price -- the input rate is the whole bill. On input, BGE-M3 is the lowest at $0.012 per 1M tokens and Qwen3 Embedding 8B the highest at $0.030 per 1M tokens, a 2.5x spread. That is 40% below the model publisher's own rate.Qwen3 Embedding 8B$0.030Qwen3 Embedding 8B, official$0.050BGE-M3$0.012BGE-M3, official$0.020
Embedding models have no output side to price -- the input rate is the whole bill. On input, BGE-M3 is the lowest at $0.012 per 1M tokens and Qwen3 Embedding 8B the highest at $0.030 per 1M tokens, a 2.5x spread. That is 40% below the model publisher's own rate.Our published rates, read from the same table the API bills from.

Licence family distribution

Catalogue by licence family
24 models
Catalogue by licence familyApache 2.0 and MIT together account for most of the catalogue -- the two licence families with no per-model ambiguity to read.
  • Apache 2.011 models46%
  • MIT7 models29%
  • Modified MIT2 models8%
  • Llama 4 Community1 model4%
  • LTXV Open Weights1 model4%
  • Stability Community1 model4%
  • Tencent Hunyuan Community1 model4%
Apache 2.0 and MIT together account for most of the catalogue -- the two licence families with no per-model ambiguity to read.Computed from the licence field on every catalogued model, reconciled 2026-09-02T14:00:00Z. See /licences for the full reference.

Segment counts read from the licence reference at build time; they are not independently maintained here.

Further reading

The numbers above, argued out in full and cited against outside reporting.

  • What a Fixed Monthly AI Budget Actually Buys in 2026

    Three realistic budget tiers, worked by hand against this catalogue's own rate table, at one stated request shape -- how many requests and tokens $10, $50 and $250 a month actually buys on three callable models.

  • When a Closed Frontier Model Is Still the Right Call

    Closed frontier models measurably lead reasoning-heavy benchmarks as of September 2026. Where that lead and a simpler operational model are worth the higher price -- and why our catalogue is not the answer for that reader.

  • Self-Hosting vs a Managed Open-Weight API: When Each Wins

    Where the self-host breakeven actually sits, what self-hosting really costs once engineering time is priced in, and the honest cases where self-hosting wins -- this is not a blanket argument for a managed API.

  • The Real Cost of an AI Coding Agent: A Token Budget Breakdown

    Why an agent's bill doesn't look like a chat bill -- a closed-frontier full-day usage pattern reported near $594/month, agent loops burning 5-30x an equivalent chat interaction, and one study's finding that 59.4% of an agent's tokens go to review, not writing.

Questions about these rankings

Do these rankings measure model quality?
No. Every ranking on this page measures something our own rate table actually contains: price, discount against the publisher's rate, context window size, or licence family. We have not run a benchmark on any model here, and we do not publish a quality or reasoning ranking — see /press for what we do not do.
Why do some rankings include models that aren't callable?
17 of 24 catalogued models are not currently served by any configured upstream. They are priced at their intended rate, and a context-window or licence ranking is still a true fact about them even though a request would return no available channel today. Every figure that includes one says so.
How often do these numbers change?
Whenever the rate table changes. These figures were last reconciled with upstream pricing on 2026-09-02T14:00:00Z, the same date the rest of the site's pricing reflects — there is one rate table, and every page, including this one, reads it live.