Replies: 1 comment
|
You asked for the arithmetic, not forecasting, and for a stop the user can leave. Two of those three are answerable today, and one is honest-to-goodness not a plugin's to give. Here is what I measured, and what I shipped.
The trap is that Two other facts that decide the design:
What I shipped, so this is not a design sketch: The measured behavior, from an end-to-end proof that mounts the packed tarball into the real harness (real // phase A — provider reports is_available=false
{ "modelCalls": 0,
"turnEndReasons": [{ "kind": "blocked" }],
"eventTypes": ["agent/inbox/spliced", "turn/start", "agent/inbox/spliced",
"agent/inbox/spliced", "turn/end"],
"durablyRecordedUserMessage": false, "stepsStarted": 0,
"inboxTextAfterA": "and now spend the rest of the balance on this" }
// phase B — after a top-up, the surviving prompt runs
{ "modelCalls": 1, "firstRequestCarriedTheSurvivor": true, "inboxAfterB": 0 }
Where I am deliberately weaker than your proposal, stated plainly:
Everything else — bounded by Your companion report #7052 is real too, and its fix is smaller than it looks — I replied there with the reason (the sibling Messages path in the same package has had the 402 rule all along). |
Uh oh!
There was an error while loading. Please reload this page.
Companion to the bug report in General: #7052 — HTTP 402 (insufficient provider balance) never reaches the terminal QUOTA failure code, which covers the classification. This post is the design question behind it.
Summary
A long run started, spent ~20.8M input tokens, and died at the finish line with no deliverable. By then my balance was −$0.24. The harness had no idea, at any point, how much money was left:
GET /user/balancereturnsis_available— literally "Whether the user's balance is sufficient for API calls" — plusbalance_infos(currency,total_balance,granted_balance,topped_up_balance). The harness calls it nowhere. Grepping the checkout foruser/balance,balance_infos, oris_availablereturns nothing, and there is nomaxCost/spendLimit/ per-session budget setting in the configuration catalog.packages/workflow/workflow/README.mdlists "No journaling or resume" next to "No token-budget vocabulary — engines cap concurrency, items, and children, but neither the request nor the result accounts for model tokens across children." So during a workflow run no component sees the total spend.The core ask is arithmetic, not forecasting
If the tank reads zero, the engine should not start. Concretely, before dispatching a request:
No prediction is involved:
is_available), so the harness does not even have to interpret the number.packages/llm/token-meterprices the current surface per route (route-pricing.ts→priceSurface).The second half: make the stop resumable
A cutoff alone does not fix this failure. A task that begins with enough balance to get halfway will still stop halfway; what made it catastrophic is what happens next.
packages/session/session-persistence/README.mdstates the resume contract plainly: "Synthetic closers are the only crash story… there is no partial-turn resume that continues an interrupted turn instead of closing it." Combined with the workflow package's lack of journaling, recovering means re-deriving the work and re-running every child — so the tokens already paid for are simply gone.On a genuine exhaustion refusal, the run should enter a durable
awaiting_balancestate and resume from the log after a top-up, instead of closing the turn as a failure. For nested delegation, each child's structured result would need checkpointing so a resumed run replays completed items rather than re-burning them.Proposed behavior
is_available === false, or whentotal_balanceminus the price of the next request falls below the floor. Sample per step; show the value and its timestamp.awaiting_balanceplus resume from the durable log, rather than a closed turn.llm-retrysupports an unlimitedalwaysmode (packages/llm/llm/src/retry-policy.ts). Against an exhausted balance that mode would retry forever, billing every attempt and never succeeding; a balance gate there would turn it into genuine "wait for top-up, then continue".What I am not asking for
Not cost forecasting, and not a predicted total for a task. Step counts are model-decided and fan-out multiplies them, so an exact pre-flight estimate is not achievable, and I would not want one gating a run. A large-job estimate is at most a later, optional nicety. The request here is an arithmetic gate — can the account pay for the next request, yes or no — plus a stop that can be resumed.
Where this could live
packages/guard/already hosts small loop-hygiene guards (repeat-tool-reminder,timeout-policy), and theagent/pre-step/agent/request-errorwaterfalls are the established seams (bothllm-retryand compaction attach there).session-checkpoint-policyalready enforces fail-closed boundaries before a model request, before top-level side-effecting tools, and before the next step — the same boundary a spend check would use, which is also where a refusal could be made resumable rather than terminal.Environment
DSH
0.1.5-rc.2, DeepSeek provider; ~20.8M input tokens consumed before failure, ending balance −$0.24. Happy to supply session logs, or to run a repro against a drained test key if that helps.All reactions