Skip to content

Use case

The best API for content generation

Content generation is the mirror image of summarisation: a short brief goes in and thousands of tokens come out, so the input rate is close to irrelevant and the comparison collapses to the output rate alone — which makes Kimi K2.6 the pick whenever the prose is published as written, since it has the lowest output rate of the strong-prose models we serve — with DeepSeek V4 Flash cheaper still for bulk copy whose drafts get edited anyway.

Read this page next to the summarisation one. Same catalogue, same arithmetic, opposite conclusion — there the input rate was the entire bill, here it is a few percent of it. Nothing about the models changed; the token shape did, and that is the whole reason a single 'cheapest model' answer is wrong.

Because output is nearly all of the cost, the expensive mistake on this workload is regeneration. Every rejected draft is paid for at the output rate and produces nothing, so the number worth tracking is cost per accepted piece — a model that needs two attempts is twice its listed rate, and a better brief is usually cheaper than a cheaper model.

What actually matters here

Output rate
The brief is small and the piece is long. This single number is close to the entire monthly bill, which is not true of any other workload here.
Draft acceptance rate
A regenerated draft is billed in full at the output rate and delivers nothing. Cost per accepted piece is the real unit, and it can differ from the listed rate by more than the models differ from each other.
Long-form coherence
Quality degrades over length in a way short benchmarks do not catch — repetition, drift, a lost thread in the final third. That is precisely the part of the output you are paying the most for.
Voice steerability
A model that ignores the style guide has to be rewritten by a human, which is the most expensive outcome available. Steerability is what turns a cheap output rate into cheap published content.

What it costs, at 60,000 pieces/month, ~2,500 tokens each

Worked from this catalogue’s published rates at 30M input and 150M output tokens a month. Your figure will differ; the arithmetic will not.

ModelInputOutputPer monthAt official rates
Kimi K2.642%$0.55/M$2.30/M$361.50$628.50
DeepSeek V4 Flash36%$0.090/M$0.18/M$29.70$46.20
GLM-5.242%$0.82/M$2.55/M$407.10$702.00

Excludes cached-input savings, which on a repeated-prefix workload typically reduce the input column substantially. Cached rates are published per model on the pricing page.

Questions

Why does the input column barely move the total here?
Because the brief and style guide are a fraction of what comes back. Look at the table: the input side of this workload is a small share of each row, which is why a model with a very cheap input rate and an ordinary output rate saves you almost nothing on generation.
Is the cheapest output rate always the right answer?
Only if the drafts are usable. Multiply the listed rate by the number of attempts a piece actually takes and compare that, because a model needing a second pass has doubled its own price while a more expensive model that lands first time has not.
Does a longer, more detailed brief cost more?
Technically yes, and it is almost always worth it. Input is the cheap side of this workload, so spending more tokens on the brief to avoid one regeneration is trading the cheap column for the expensive one.
Should I generate an outline first?
For anything long, yes. An outline pass is short output and gives you a cheap place to reject a bad direction before paying for the full piece. Rejecting at the outline stage costs a fraction of rejecting a finished draft.