You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Short answer: the spread between the cheapest and the most expensive model a working developer might reasonably pick is larger than most people budget for, and the figure that moves your bill is usually not the one on the pricing page. What follows is every figure this repo has collected on that, each kept with the sentence it was published in and the date it was read.
The figures that exist, and where each came from
$0.06–$0.2 per million — "At the low end, MiniMax M3 runs $0.06 to $0.2 per million and draws 60% to 70% of its revenue from outside its home market." (2026-08-19)
$1 per million — "Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near $1 per million tokens at inference gross margins of 10–20%." (2026-08-19)
$1.25 / $4.25 — "Meta priced Muse Spark 1.1 at $1.25 per million input and $4.25 per million output, roughly 75% and 83% below the model it was positioned against." (2026-08-19)
16x — "A spread from $0.06 to $1 per million is more than 16x, and peak pricing adds another factor of 2 on top." (2026-08-19)
60 percent — "Microsoft's evaluation of Kimi K3 landed on a number that should change how you read a pricing page: about 60 percent." (2026-08-05)
400-token — "Anthropic's 400-token SKILL.md file, through its 'two-pass workflow' and specific aesthetic guidance, has achieved results." (2026-08-07)
10,000 tokens — "OmniRoute aggregates 237 providers and advertises roughly 1.6 billion free tokens a month, and that figure is a compression ratio before it is a strategy." (2026-08-19)
What none of these are. Every per-million price is a list price on a day, not a bill. None of them accounts for the thing that actually dominates a coding agent's spend — the project context re-sent on every turn, which is charged at input rates and grows with the repository rather than with the task. A model that is 16x cheaper per token is not 16x cheaper to work with if it needs three attempts where the expensive one needs one.
Every figure quoted next to a model name, on one page:
The one thing worth replying with: what did your last full month of coding-agent usage actually cost, and roughly what share of it was context you re-sent rather than output you kept? A one-line reply with a number and a rough split is worth more than any row above, and it goes into the table with your wording kept.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Short answer: the spread between the cheapest and the most expensive model a working developer might reasonably pick is larger than most people budget for, and the figure that moves your bill is usually not the one on the pricing page. What follows is every figure this repo has collected on that, each kept with the sentence it was published in and the date it was read.
The figures that exist, and where each came from
$0.06–$0.2per million — "At the low end, MiniMax M3 runs $0.06 to $0.2 per million and draws 60% to 70% of its revenue from outside its home market." (2026-08-19)$1per million — "Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near $1 per million tokens at inference gross margins of 10–20%." (2026-08-19)$1.25/$4.25— "Meta priced Muse Spark 1.1 at $1.25 per million input and $4.25 per million output, roughly 75% and 83% below the model it was positioned against." (2026-08-19)16x— "A spread from $0.06 to $1 per million is more than 16x, and peak pricing adds another factor of 2 on top." (2026-08-19)60 percent— "Microsoft's evaluation of Kimi K3 landed on a number that should change how you read a pricing page: about 60 percent." (2026-08-05)400-token— "Anthropic's 400-token SKILL.md file, through its 'two-pass workflow' and specific aesthetic guidance, has achieved results." (2026-08-07)10,000 tokens— "OmniRoute aggregates 237 providers and advertises roughly 1.6 billion free tokens a month, and that figure is a compression ratio before it is a strategy." (2026-08-19)What none of these are. Every per-million price is a list price on a day, not a bill. None of them accounts for the thing that actually dominates a coding agent's spend — the project context re-sent on every turn, which is charged at input rates and grows with the repository rather than with the task. A model that is 16x cheaper per token is not 16x cheaper to work with if it needs three attempts where the expensive one needs one.
Every figure quoted next to a model name, on one page:
Write-ups with the full context:
The one thing worth replying with: what did your last full month of coding-agent usage actually cost, and roughly what share of it was context you re-sent rather than output you kept? A one-line reply with a number and a rough split is worth more than any row above, and it goes into the table with your wording kept.
All reactions