Replies: 2 comments
|
Quick look at how other inference platforms handle the same two questions (observed 2026-08-24):
Two takeaways for the questions above:
|
Position on the six questions, with production data as of 2026-08-25Answering directly. Two framing notes first, because both change what the answers mean. No plan changes are being made right now. These are positions for review, not decisions already taken. The Token Plan continues as-is while the items below are worked through. One finding materially changes question 4. More on that below, but in short: the safeguard we already built for it currently matches zero models in production. 1. Should plans be tiered, and on which axis?Yes, eventually — but a per-subscriber value cap comes first, and it matters more. Tiering redistributes who can spend how much. It does not stop one subscriber exhausting a pool, which is the failure actually observed ($6.12 in a single day against a $6 plan). Every peer that launched a flat plan later added a cap — Chutes at 5× PAYG-equivalent value, with overflow billed as PAYG — and several also had to retire their cheapest tier afterwards. Adding tiers before a cap reproduces the same problem at four price points instead of one. When tiers do land, the axis should be concurrency and model access, not raw token volume. Token counts are what let a single subscriber concentrate spend on the most expensive model; concurrency and access bound the damage directly. The cap should be denominated in payout value, not tokens. 40M tokens means something different on an open-weight 8B than on a model paying $3.24/M output — a token cap is not a cost cap. 2. Is stake-SWAN-for-quota viable, and who funds it?Only in the DIEM shape, and never funded from the subscription pool. Venice's design is the one working reference: stake → mint a transferable unit redeemable for a fixed daily credit, no rollover, funded by emissions. It answers the "zero SWAN demand" objection honestly. The critical constraint is the second half of the question. Venice can absorb it because it has no supply side to pay. Swan does. Staked traffic still costs a provider real compute, so it needs its own funding line — capped emission or Growth Fund — declared up front. Funding it from the subscription pool would dilute providers further to create token demand, which is the same underwater dynamic with extra steps. 3. Should the platform guarantee a minimum payout ratio?Not as a standing guarantee. A floor is a subsidy with no natural ceiling: it pays out most exactly when the pool is most underwater, so it grows with the problem instead of bounding it. Fix the inflow/outflow mismatch first (cap, model exclusion, tiering). If a floor is still wanted afterwards, it should be time-boxed, capped in absolute terms, and announced as a transition measure — not an open commitment providers price into their expectations. What providers most need in the interim is not a floor but visibility: the projected ratio published during the period, not discovered at settlement. 4. Exclude high-payout models from plan coverage?Yes — and this is where the important correction is. We already built this.
Every model is It also corrects an earlier assumption of mine. I had concluded the platform pays these vendor invoices directly. It does not: the proxy path is only reached for That is not the milder problem it sounds like. The pool's design assumes a provider's cost is sunk GPU time it already owns. A provider paying a per-token invoice to a frontier vendor and receiving ~37 cents on the dollar loses money on every subscription request, and will rationally stop serving. With 4 providers online, that is a supply risk before it is an economics one. So the answer to question 4 is yes, but the mechanism is not a payout-rate threshold. It is deciding, per model, who holds the vendor relationship, and excluding relay-backed models from plan coverage regardless of how 5. Provider pricing, and the minimum provider count?Keep uniform consumer pricing. Revisit at ≥5 healthy online offerings per model — a threshold no model currently meets. Uniform per-model pricing is the category norm. The only peer with provider-set pricing is OpenRouter, whose "providers" are companies under contract, not GPU owners — and even there the consumer sees one price, with provider differences expressed through routing weight. Price-weighted routing needs liquidity to be anything other than noise. Today the top model by traffic is served by one provider; a price auction among one seller is not an auction. Adding it now would add real complexity to routing and change nothing about what consumers pay. What is worth doing before then: a provider reservation floor (a provider may decline traffic below a rate rather than quote a price to consumers), and uptime-gated routing. 6. Formalise as SIP-004?Not yet. A SIP should encode decisions that are settled and implementable. Two of the six above are not: the model-classification question in (4) is unresolved, and the value cap in (1) is unbuilt. Suggested sequence: land the value cap and the classification audit, run one full billing period, then write SIP-004 against observed numbers rather than projections. Writing it now would freeze in the projections — including the 36.8% figure, which is itself a projection from a partial month. What would change these answersDeliberately stated, since none of this is high-confidence:
Provider feedback on the first point is the most useful thing this discussion could produce. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context
This is a community-originated discussion. The questions below were raised by computing providers and community members (originally in Chinese, quoted with translation) about the Token Plan — the flat-fee subscription on Swan Inference — and how it affects providers:
To give the discussion a shared baseline, this post summarises how the plan works today and lists the pros and cons that have come up so far. Nothing here is a decision or an official position — it is a starting point for feedback. If the discussion converges, it can be written up as a SIP.
Related: SIP-003 (#21) defined the per-token catalog, the provider revenue split and Pay-with-SWAN. It did not cover the subscription pool, which was shipped afterwards and is currently described only in the provider FAQ on inference.swanchain.io.
All numbers below were pulled from the public API on 2026-08-24 (
/api/v1/subscription/plans,/api/v1/stats/subscription-pool,/api/v1/models,/api/v1/providers).1. How the Token Plan works today
Plan on offer (only one):
free+standardtiers (open-source models).premiumtier (Claude, Gemini, Gemma-4-31B…) is pay-as-you-go onlyProvider settlement: pay-as-you-go requests are paid to providers per request and are never affected by the plan. Subscription traffic is paid out of a shared monthly pool; if subscriber usage in a month costs more than the pool collected, providers' subscription earnings for that month are reduced proportionally.
Live pool data — August 2026 (24 of 31 days):
Daily cost has been $1.2–1.5/day for most of August, with a $6.12 spike on Aug 19 — a single subscriber can burn through a month of pool revenue in one day.
Where the traffic actually goes (weekly tokens, top models):
~111M of ~122M weekly tokens (91%) land on three
standardmodels that are plan-eligible, each served by 1–2 providers. Notegpt-5.4-miniis taggedstandardand therefore plan-eligible at a $3.24/M output payout — 10× the payout of the open-source models in the same pool.Provider fleet: 48 registered → 7
active, 5approved, 2under_review, 6suspended, 28pending(many are junk/test signups). 4 online right now. Of the 58 catalog models, 46 have zero online providers.2. Pros and cons of the current design (as raised so far)
What works
What doesn't
gpt-5.4-miniisstandardand thus plan-covered at a $3.24/M output payout. One subscriber routing to it can drain the pool at 10× the rate of open-source models.3. Community proposal A: multiple tiers + stake-for-quota
Tiering. A straw-man that keeps the current Pro as the middle option:
The point is not the exact numbers; it is that each tier should be sized so it can realistically fund the providers serving it, and that sizing should be published so providers can judge the pool before opting in.
Stake-for-quota. SIP-003 §7.3 rejected "stake-for-inference" as high-friction. The community suggestion is narrower and worth reconsidering as a complement: stake SWAN → receive a daily quota on
standardmodels, with the staked amount at risk of nothing but opportunity cost. Pros: creates SWAN lock-up demand, gives token holders a reason to use the product, and a staker's quota can be sized so it never exceeds what the yield on the stake would fund. Cons: someone must fund the providers serving staked traffic — either emissions (which SIP-003 just voted to sunset) or the Growth Fund. If it's emissions-funded, this is UBI by another name; it should be sized explicitly and capped.Guardrails that apply regardless of tiers (these fix the "白嫖" problem more directly than tiering does):
gpt-5.4-minican't drain the pool.4. Community proposal B: provider-set pricing vs uniform pricing
Today every model has one list price and one payout price (90%), set by admin. With most models at 0–2 online providers there is effectively no competition yet, so this is a forward-looking question.
Case for provider-set pricing
Case for uniform pricing
Middle-ground options worth discussing
5. Questions for the community
gpt-5.4-mini, Gemma-4-31B) be excluded from plan coverage?Providers currently serving plan-covered models — please share your actual settlement experience (July settled at 100%, August is projecting ~44%).
All reactions