You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Short answer: the free part is mostly a token grant on top of a paid API that is already an order of magnitude cheaper than the US frontier models — not a permanent free tier. Below is every price figure this repo has collected on that comparison, each kept with the sentence it was published in and the date it was read.
Input price, per million tokens
$0.19 — "Chinese AI models provide a cost-effective alternative to their American counterparts, with input costs as low as $0.19 per million tokens, compared to OpenAI's $5-12." (2026-08-07)
$1 — "Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near $1 per million tokens at inference gross margins of 10% to 20%." (2026-08-08)
$0.06 – $0.2 — "At the low end, MiniMax M3 runs $0.06 to $0.2 per million and draws 60% to 70% of its revenue from outside its home market." (2026-08-08)
Where the "1.6 billion free tokens" number comes from. It is a grant total across providers, not a per-account allowance, and the write-up's own reading is that the interesting part is compression rather than the grant:
10,000 → 1,080 tokens — "This strategy revolves around using tools like OmniRoute, which aggregates 237 providers, compressing 10,000 tokens to 1080 through RTK+Caveman technology." (2026-08-07)
What this does not tell you. None of these is a rate card read off a vendor page — they are figures quoted inside write-ups, which is exactly why the sentence travels with each one. Prices move; a $0.19 from 2026-08-07 is a dated observation, not a current quote. And a low per-token price says nothing about rate limits, context window, or whether the endpoint stays up.
If what you actually want is a list of endpoints that are free with no card, that is a different list and it is maintained separately: https://xyzs996.github.io/free-llm-api/
The one thing worth replying with: which provider did you actually put a coding workload on, and what did a real day of it cost? Naming the model and the rough daily spend is enough — it goes into the table with your wording kept, and a dated real bill outranks every row above.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Short answer: the free part is mostly a token grant on top of a paid API that is already an order of magnitude cheaper than the US frontier models — not a permanent free tier. Below is every price figure this repo has collected on that comparison, each kept with the sentence it was published in and the date it was read.
Input price, per million tokens
$0.19— "Chinese AI models provide a cost-effective alternative to their American counterparts, with input costs as low as $0.19 per million tokens, compared to OpenAI's $5-12." (2026-08-07)$1— "Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near $1 per million tokens at inference gross margins of 10% to 20%." (2026-08-08)$0.06 – $0.2— "At the low end, MiniMax M3 runs $0.06 to $0.2 per million and draws 60% to 70% of its revenue from outside its home market." (2026-08-08)Where the "1.6 billion free tokens" number comes from. It is a grant total across providers, not a per-account allowance, and the write-up's own reading is that the interesting part is compression rather than the grant:
10,000 → 1,080 tokens— "This strategy revolves around using tools like OmniRoute, which aggregates 237 providers, compressing 10,000 tokens to 1080 through RTK+Caveman technology." (2026-08-07)What this does not tell you. None of these is a rate card read off a vendor page — they are figures quoted inside write-ups, which is exactly why the sentence travels with each one. Prices move; a
$0.19from 2026-08-07 is a dated observation, not a current quote. And a low per-token price says nothing about rate limits, context window, or whether the endpoint stays up.If what you actually want is a list of endpoints that are free with no card, that is a different list and it is maintained separately: https://xyzs996.github.io/free-llm-api/
Write-up with the full context: https://xyzs996.github.io/llm-api-pricing/articles/how-chinese-ai-agent-tools-leverage-1-6-billion-free-tokens.html
All 296 figures as JSON/CSV: https://xyzs996.github.io/llm-api-pricing/figures.html
The one thing worth replying with: which provider did you actually put a coding workload on, and what did a real day of it cost? Naming the model and the rough daily spend is enough — it goes into the table with your wording kept, and a dated real bill outranks every row above.
All reactions