You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Some steps persisted in assistant/message.stream carry packed runs where every block has exactly one member and dt is empty. For those steps the decode window (first token → assembled message) collapses to 1–5 ms while usage.outputTokens stays in the hundreds-to-thousands. Both deriveTurnMetrics and the sessionStats projection compute throughput as outputTokens / decodeMs with nothing but a decodeMs > 0 guard, so they emit values in the 10^5–10^6 tok/s range and fold them into the session total.
Measured on one machine, same session data:
session
steps
affected steps
displayed
recomputed from healthy steps only
session-56e9f8cf… (redacted)
125
113
1906.2 tok/s
149.1 tok/s
session-3e54dd78… (redacted)
9
0
131.6 tok/s
131.6 tok/s
Impact
Turn footer and session-wide throughput are not trustworthy for any session containing such steps, and the error grows with their share.
In a scan of recent local sessions, the affected share ranged from 38% to 93% across 9 sessions.
ttftMs / ttftSteps are unaffected (separate accumulators, not derived from the decode window).
Reproduction
Have any session whose assistant/message.stream contains steps where every packed run has a single member.
Open the session and read the throughput in the turn footer or session stats.
Compare against a recomputation that only uses steps whose runs have a non-empty dt.
Decode accounting used here matches the server projection: decode window = assistant/message event time − assistantStreamFirstTokenTime(stream); throughput = that step's usage.outputTokens / decode window.
Evidence
Packed-run shape, healthy steps vs affected steps (members = total texts/args members across runs):
turn/step
runs
members
sum(dt)
1/1
3
927
924
1/2
3
670
667
2/2
3
3
0
2/4
3
3
0
The first affected step is step 2/2; every step from there on shows the same shape.
Analysis
Chunks are stamped at the consumption site, packages/core/agent-loop/src/assistant-stream.ts:
this.accumulator.push({time: Date.now(), chunk })
So “one member per block and an empty dt” is equivalent to: the whole SSE stream was consumed within the same millisecond, and each content block arrived as a single delta.
Neither accounting entry point protects against that:
packages/client/ui-chat/src/client/contract/turn-metrics.ts (deriveTurnMetrics) only checks decodeMs > 0
I did not determine the origin of the whole-block arrival. Ruled out so far:
Upstream always coalescing large requests: control requests through the same gateway and model with 20k / 80k token inputs arrived as 17 / 34 data events with 2–30 ms spacing, i.e. normally streamed.
Context size: one session with 390k input tokens had 0% affected steps.
User mid-turn injections or retries: one session with only 2 injections had 78% affected steps.
Not ruled out: the local path (client → LiteLLM gateway → upstream). The claim in this report is therefore only the missing guard in the accounting layer, which holds regardless of where the coalescing comes from.
Suggested fix
In deriveTurnMetrics and sessionStats: exclude steps where every run has an empty dt and a single member from throughput accounting; or, at minimum, add a floor for the decode window (e.g. a step with no incremental members and a sub-50 ms window does not contribute). TTFT should stay as is.
Related
Discussion 轨迹(Trajectory)面板:历史回放时首 token 延迟/生成时间/吞吐量恒显示"首 token 时间不可用" #6129 covers a different client-side read of the same packed-run shape (trajectory panel timing unavailable on historical replay). This report is the throughput accounting face of the same data shape.
Environment: dsh 0.1.7-rc.2 (npm global install), macOS. Has anyone else seen this? Happy to open a PR with the guard plus tests if that is useful.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Some steps persisted in
assistant/message.streamcarry packed runs where every block has exactly one member anddtis empty. For those steps the decode window (first token → assembled message) collapses to 1–5 ms whileusage.outputTokensstays in the hundreds-to-thousands. BothderiveTurnMetricsand thesessionStatsprojection compute throughput asoutputTokens / decodeMswith nothing but adecodeMs > 0guard, so they emit values in the 10^5–10^6 tok/s range and fold them into the session total.Measured on one machine, same session data:
session-56e9f8cf…(redacted)session-3e54dd78…(redacted)Impact
ttftMs/ttftStepsare unaffected (separate accumulators, not derived from the decode window).Reproduction
assistant/message.streamcontains steps where every packed run has a single member.dt.Decode accounting used here matches the server projection: decode window =
assistant/messageevent time −assistantStreamFirstTokenTime(stream); throughput = that step'susage.outputTokens/ decode window.Evidence
Packed-run shape, healthy steps vs affected steps (
members= totaltexts/argsmembers across runs):The first affected step is step 2/2; every step from there on shows the same shape.
Analysis
Chunks are stamped at the consumption site,
packages/core/agent-loop/src/assistant-stream.ts:So “one member per block and an empty
dt” is equivalent to: the whole SSE stream was consumed within the same millisecond, and each content block arrived as a single delta.Neither accounting entry point protects against that:
packages/client/ui-chat/src/client/contract/turn-metrics.ts(deriveTurnMetrics) only checksdecodeMs > 0packages/session/session-stats/src/projection.ts(sessionStats) accumulates unconditionallyWhat I could not establish
I did not determine the origin of the whole-block arrival. Ruled out so far:
Not ruled out: the local path (client → LiteLLM gateway → upstream). The claim in this report is therefore only the missing guard in the accounting layer, which holds regardless of where the coalescing comes from.
Suggested fix
In
deriveTurnMetricsandsessionStats: exclude steps where every run has an emptydtand a single member from throughput accounting; or, at minimum, add a floor for the decode window (e.g. a step with no incremental members and a sub-50 ms window does not contribute). TTFT should stay as is.Related
Environment:
dsh0.1.7-rc.2 (npm global install), macOS. Has anyone else seen this? Happy to open a PR with the guard plus tests if that is useful.All reactions