Replies: 2 comments
|
Thanks — the detection half of this is exactly as you describe, and I checked it against the tree. The retry half is in better shape than the report implies, which changes what a fix has to touch. The guard is keyed on content, not on accounting. But the failure path never asks whether content was streamed. So the shape you ask for already exists end to end — detect → error finish → retryable code → retry — and the retry is not blocked by the partial text. The missing half is purely the detection. And the detection is reachable from a plugin. The One caution for whoever implements the zero-usage discriminator. Zero usage is only evidence against a control: a route that never reports usage returns 0/0/0 for every clean stop, and treating all of them as failures converts a silent truncation into up to I maintain a small plugin for the sibling case you link (discussion 6743 item 2 / #6218 — a completion with no visible content reported as a successful turn); it emits the same retryable error finish at the same seam: https://github.com/argszero/cordis-plugin-empty-response-guard. Your case is the same defect one branch further out, so I plan to extend that guard rather than publish a second plugin for it, and will follow up here with the version once it lands. |
|
Follow-up: the version that ships this is published. When I replied above I said I would extend this guard rather than publish a second plugin at the same seam, and follow up here with the version once it landed. It landed:
npm install @argszero/cordis-plugin-empty-response-guard- insert:
- id: empty-response-guard
name: '@argszero/cordis-plugin-empty-response-guard'Mounting needs no configuration. For the shape you described — a stream that ends with The control is the part worth reviewing. Zero usage is only evidence if the route reports usage at all: a gateway that never sends a usage chunk would otherwise have every clean stop read as fabricated, trading one silent truncation for up to Three details that decided the implementation, in case they are useful for the upstream fix:
The earlier shape (the reasoning-only stop from #6218) is unchanged and still wins when both could apply; 45 tests pass ( |
Uh oh!
There was an error while loading. Please reload this page.
Summary
A provider can truncate a response mid-stream and still signal a clean completion. When it does, DSH records the turn as
completedand the assistant's half-finished message is treated as a finished answer. The existing guard inmapStopReason()only catches the fully-empty case, so a truncated-but-non-empty response slips through silently.Filed as a Discussion per CONTRIBUTING (issues/PRs are closed on this repo).
Observed behavior
In a real durable session, the final step of a turn ended like this:
Then:
The user-visible result: my reply stopped in the middle of a sentence (
"The spawn in runner-launch has "— literally trailing off mid-clause). The system considered the turn successfully completed, so there was no retry and no error surfaced. The only reason I could explain it afterwards was by reading the journal.Why the existing guard doesn't catch it
packages/llm/llm-pi-ai/src/stream.ts—mapStopReason()(line ~208):The guard is
content.length === 0. In this incident the step carried 1033 characters of reasoning plus 1033 characters of text (the text block was a duplicate of the reasoning), socontent.length !== 0and the response was accepted as a legitimatestop.The discriminating signal: zero usage
Across all 57 steps of that turn, this was the only step reporting
usage = 0/0/0:{"kind":"tool-calls"}{"kind":"stop"}toolUsestopSteps 1–56 all carried non-zero usage (thousands to 226k). A genuinely completed generation cannot report zero tokens in and zero out. Compare the other two completed turns in the same session, whose terminal steps had normal accounting:
Contributing factor: a flaky relay
The provider route (
deepseek/deepseek-v4.1-flashvia a third-party OpenAI-compatible relay) was returning 502 with no body earlier in the same turn:Those were retried and recovered (
SERVERis retryable by default). The terminal truncation is the same infrastructure degrading, but it arrives as a successful stream with a fabricated terminal event, so the retry policy never engages — retry covers failed requests, not truncated-but-"successful" responses.Relationship to existing discussions
This is adjacent to but distinct from #6743 (which is excellent and I read carefully before filing):
EMPTY_RESPONSE(mislabeled as model behavior){kind:"stop"}→ turncompleted#6743 asks to relabel the empty case. My case is the gap next to it: the guard is keyed on emptiness, so a truncated response that produced some content is indistinguishable from a clean finish and never even reaches
EMPTY_RESPONSE.Also related, though about
lengthtruncation rather than a fabricatedstop: #2143, #1703, #1263 (allINVALID_REPLAY_STATEafterlengthtruncation), and #6918 (client discards already-streamed body when the server drops). None of these covers "stop+ non-empty + zero usage".Proposed approach
Treat a
stopwith zero usage as a transport fault rather than a model outcome:mapStopReason(), extend thestopcase: ifcontent.length === 0or the usage is all-zero (inputTokens === 0 && outputTokens === 0), map to an error.SERVER(or a dedicated code such asTRUNCATED_RESPONSE) so the existingdsh-llm-retrypolicy retries it automatically —SERVERis already in the defaultretryableCodes."provider ended the stream with zero usage — response was likely truncated, not completed".This is deliberately narrow: it only fires on the combination (
stop+ zero usage), so it cannot affect normal completions, and since usage totals are already recorded per step it needs no new bookkeeping.A secondary, complementary idea: when a terminal step's finish is
stopand neither text nor a tool call ended on a sentence/pre-block boundary, that is suspicious — but I'd treat that as much weaker evidence than the zero-usage signal and would not gate on it.Environment
@deepseek-ai/dsh-llm-pi-ai@0.1.5-rc.2)openai-responsesAPI), modeldeepseek/deepseek-v4.1-flashretryPolicy: { mode: normal, maxRetries: 5, retryableCodes: [EMPTY_RESPONSE, RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT] },streamIdleTimeoutMs: 300000I can provide a redacted
session.v3.jsonl.zstdextract with the exact step-57 records (stream chunks,usage,finish,turn/end) if that's useful.Impact
This is worse than it first looks. A half-sentence is obvious, but the same mapping applies to any long turn: if the relay truncates while the model is mid-way through editing files or running a multi-step task, the turn is recorded as
completed, no retry happens, and the work silently stops halfway with the agent believing it finished. In this session the truncation was cosmetically obvious; on a code-editing turn it would not be.Thanks for the harness — the append-only journal is what made this diagnosable at all.
All reactions