Replies: 4 comments
|
This is the single largest data point yet in the billing-visibility cluster, and the way you had to find it — hand-decompressing Two details from your table worth underlining: (1) ~98.4% of the 642M was cache-read tokens — cheap per token, but nothing throttled the loop producing them and the cost still landed; (2) the usage events were in the session logs, so per-request surfacing (a cost badge / hard budget line) would have caught this in minutes, not 7 hours. If you're willing — the asks in #7812 are (1) terminable error classification, (2) a per-request cost badge, (3) a balance/budget guard, (4) usage records (incl. web_search) into session logs. Does your case argue for a fifth — a turn-loop circuit breaker (per-goal-round cap on steps/tokens before "Continuing goal" auto-fires)? Or do you think the budget guard alone would have stopped it? |
|
On the circuit-breaker vs budget-guard question: they are different controls, and a runaway like this needs both. Rate limits bound velocity; budget caps bound total spend. In our own agent stack, governance covered memory, CPU and process count only, with no token or dollar ceiling, so a firing loop could burn unbounded spend while staying inside its RAM limits. One more lesson: a cap tracked in the loop's own state is worth only as much as the loop's correctness, which is exactly what's in question during a failure. The breach behavior we ended up specifying was alert the human, pause, and wait, not retry. Disclosure: this is from a book we sell, https://book.hool.dev (free sample at /sample). |
|
Had similar issue, sometimes when dsh set goals and dispatch subagents to do the work, the main agent have nothing else to do. Then it starts to polling infinitly 18:50 no sleep between two polls. Found in both 0.1.7 and 0.2.0. Model is deepseek-v4-pro, max reasoning effort. |
|
Thanks for the report. |
Uh oh!
There was an error while loading. Please reload this page.
Environment
0.1.7-rc.2(commit477b4f4205)v24.15.0ptc(danger-full-access)z-ai/glm-5.3-flash(via OpenRouter)Summary
today i updated my DSH instance from
0.1.6-alpha.2to0.1.7-rc.2and ran my scheduled daily pipeline, and i was terrified to find that multiple sessions went into an unconstrained runaway loop, burning over 1.2 billion tokens combined across 4 sessions, with one single session hitting 642,005,901 tokens and running for over 7 hours before i noticed and manually stopped itin
0.1.6-alpha.2, all my sessions were completely normal, running 50 to 120 steps and taking around 5M to 15M tokens total, but on0.1.7-rc.2, three compounding bugs caused an uncontrollable token explosionToken usage breakdown from raw session logs
i decompressed the
session.v4.jsonl.zstdfiles from~/.dsh/sessions/to inspect the exact usage events:The 3 compounding root causes
1. Unthrottled
goal-round-driverauto-triggers "Continuing goal"in
0.1.6-alpha.2, when an agent finished Turn 1, the session stopped cleanly, but in0.1.7-rc.2, the moment Turn 1 completes,@deepseek-ai/dsh-goal-round-driverchecksgoal.roundsStarted < maxGoalRounds(which defaults to256)because the goal was still active, the driver immediately injected:
<goal_round>Objective: "..."Round: 2/256Continue working toward the objective in this same session...</goal_round>in the web UI, this showed up as cards titled "Continuing goal" (at 19:05, 19:35, 19:55) without any human clicking continue, and the agent started Turn 2 and Turn 3, repeating the entire goal over again with massive context history
as noted by @GuyueHermit in #6169, there is zero time throttling in
packages/goal/goal-round-driver/src/index.ts, so it immediately fires the next round within seconds2. Auto-compaction failed, context grew to 619,000 tokens per request
as reported by @JWIMaster in #7650, auto-compaction does not trigger in
0.1.7-rc, so in my reviewer session, by step 439 of Turn 2, every single request was sending 619,619 prompt tokens to the model:{"turn": 2, "step": 439, "inputTokens": 2955, "outputTokens": 411, "cacheReadTokens": 619619, "totalTokens": 622985}when an agent runs 1,167 steps with 600k context history,
token-metersums all the steps into 642M tokens, which is terrifying to see in the UI3. Surrogate output truncation broke tool results and trapped the agent
in
0.1.7-rc.2(PR #4827 /cordis.patch.ymlchangingmaxInlineBytes: 50000tomaxInlineTokens: 12500), file reads started returning[…]cutoffs, and bash commands insiderun_codereturned(run_code completed with no output)if not explicitly console loggedthe agent got completely confused by the missing/truncated tool output and spent 700 steps retrying bash commands in circles trying to read the files, so it never reached the end of its checklist to call
update_goal(complete), which left the goal active and triggered the runaway round driver loopSuggested fixes for upstream
dsh-goal-round-drivermust have a default hard budget (e.g. max cumulative tokens or max rounds defaulting to 3 or 5, not 256) so runaway loops cannot bankrupt users[…]in ways that confuse models into retry loopsi had to immediately downgrade my production PM2 instance back to
0.1.6-alpha.2, where everything is working normally againAll reactions