Skip to content

Resources · AI Token Router

RankingsPricingOpen-weight

Open-Weight Model Pricing Ranked: The 2026 Discount Table

Ranked by discount depth against each model's own publisher, the callable open-weight catalogue here runs 36-42% below official pricing on a plain rate card, and 65-68% once a repeated prompt prefix is served from cache on a real agent workload — the two figures differ because caching changes what counts as 'the bill,' not what the model costs per token.

Put your own numbers in before you take ours on trust.

Rank these against your own volume

Two rankings, because one number hides the workload

A discount percentage means nothing without the shape of the bill it was computed on. The rate-card ranking below prices every token fresh, which is the right comparison for a one-shot request — a single classification call, a one-off summarisation job. It is the wrong comparison for anything that runs in a loop.

An agent or a chat session resends the same system prompt, the same tool schema, and a growing transcript on every turn. That repeated prefix is billed at a cached rate rather than the standard one, and industry reporting on prompt caching puts the structural saving from caching alone at 41-80% on repeated-prefix workloads, independent of which model is used [1]. The agent ranking below applies that same mechanism to our own rate table, on a stated workload rather than a vague 'up to' figure.

Rate-card discount, callable models
Rate-card discount, callable modelsEvery model here beats its own publisher by a comfortable double-digit margin, and the spread between the smallest and largest discount is narrower than the spread in raw price. Kimi K2.6 carries the deepest discount in this set at 42%, against 36% for DeepSeek V4 Flash. Figures are the rate-card discount, with no caching assumed.Kimi K2.642%GLM-5.242%Kimi K340%Qwen-Image40%Qwen3 Embedding 8B40%DeepSeek V4 Pro37%DeepSeek V4 Flash36%
Every model here beats its own publisher by a comfortable double-digit margin, and the spread between the smallest and largest discount is narrower than the spread in raw price. Kimi K2.6 carries the deepest discount in this set at 42%, against 36% for DeepSeek V4 Flash. Figures are the rate-card discount, with no caching assumed.Computed from our published rates against the official rate.

The agent-workload ranking

This is the same seven models, the same publisher's own rate as the baseline, on one stated shape of work: 400 million input tokens a month with 60% of that input a repeated prefix, and 8 million output tokens. That volume and cache share are conservative — a coding agent with a fixed system prompt and a stable tool schema typically repeats considerably more than 60% of its input.

Agent-workload discount, same seven models
Agent-workload discount, same seven modelsEvery model's discount widens once caching enters the picture, and the ranking order shifts slightly from the rate-card version — the model that wins on a one-shot request is not always the one that wins on a loop. Kimi K2.6 carries the deepest discount in this set at 68%, against 65% for DeepSeek V4 Flash. Figures assume 400M input tokens a month with a 60% repeated prefix, and 8M output.Kimi K2.668%GLM-5.268%DeepSeek V4 Pro66%Kimi K365%DeepSeek V4 Flash65%
Every model's discount widens once caching enters the picture, and the ranking order shifts slightly from the rate-card version — the model that wins on a one-shot request is not always the one that wins on a loop. Kimi K2.6 carries the deepest discount in this set at 68%, against 65% for DeepSeek V4 Flash. Figures assume 400M input tokens a month with a 60% repeated prefix, and 8M output.Computed from our published rates against the official rate, on the stated workload.

Why the market moved this direction

This is not an isolated pattern. Independent pricing surveys published in 2026 put open-weight hosted APIs at $0.07-$0.90 per million tokens against list prices as high as $15 per million input and $60 per million output at the top of the closed-frontier tier — a spread of up to 625x for comparable-class tasks [2]. The same reporting notes list pricing across the market fell roughly 80% between early 2025 and early 2026 [2], which is the general trend this table sits inside rather than an exception to it.

None of that is a claim that open-weight models are categorically better. Closed frontier models measurably lead on reasoning-heavy benchmarks as of September 2026, by a margin independent reporting puts at several percentage points [3]. The claim this table supports is narrower and checkable: for the models in it, the discount against that model's own publisher is real, it widens under caching, and both numbers are computed from the same rate table the API bills from rather than asserted.

The full table

Models marked not callable are catalogued at their intended price but not yet served by any configured upstream — a request for one returns a clear error rather than a fabricated result. Seven of twenty-four are callable today.

ModelPublisherRate-card discountAgent-workload discountCallable
Kimi K2.6Moonshot AI42%68%Yes
GLM-5.2Z.ai42%68%Yes
Kimi K3Moonshot AI40%65%Yes
Qwen-ImageAlibaba40%Yes
Qwen3 Embedding 8BAlibaba40%Yes
DeepSeek V4 ProDeepSeek37%66%Yes
DeepSeek V4 FlashDeepSeek36%65%Yes

Sources

  1. [1] Prompt Caching in 2026: Cut Your LLM API Costs by Up to 90% DevToolLab. General industry figure for caching's structural cost reduction, cited for the mechanism, not for any specific model's rate.
  2. [2] LLM Inference Cost 2026: Cost per Million Tokens packet.ai. Market-wide pricing spread and the 2025-2026 price decline, cited as market context.
  3. [3] Open Source vs Closed LLMs: Technical Comparison 2026 Hakia. Cited for the honest counterpoint: closed frontier models retain a measured benchmark lead.

Questions this raises

Why do the two rankings put models in a different order?
The rate-card ranking prices every token fresh. The agent ranking assumes 60% of input is a cached, repeated prefix — and models differ in how much their cached rate discounts their standard rate, so a model that is merely good on a one-shot request can be the best choice for a loop.
Does a bigger discount mean a better model?
No — it means a bigger gap between what we charge and what the model's own publisher charges for the same model. It says nothing about output quality, which we do not rank; see /press for what we do not claim.
Why are 17 of the 24 catalogued models not callable?
They are priced at their intended rate but no configured upstream currently serves them. We list them anyway, marked clearly, rather than removing them the moment they become servable and readding them later — the catalogue is the honest current state, not a promise.

AI Token Router is an OpenAI-compatible gateway for open-weight models, priced below each publisher’s own rate on every row.

Related