[dsh_session_log] First upload carries the entire session log, permanently breaking long sessions with HTTP 413 #6862
Replies: 2 comments
|
Corroborating this on a different machine and a later build — same watermark mechanism, but a different gateway failure mode. dsh
Instead of your Two small additions to your write-up:
- id: session-log-deepseek
config:
enabled: falseBoth deadlocked sessions produced a normal assistant message again ~40 s after the file was written, with More detail, including the ruled-out endpoints and the per-attempt timing, is in #6847. |
|
Corroborating on A long-lived session in my local profile had 0
Four notes that may help:
Workaround confirmed here as well: Adding this as corroboration on |
Uh oh!
There was an error while loading. Please reload this page.
Environment
0.1.6-alpha.1(profile:web)deepseek-official, wire protocolmessages(the new default sinceb0641b83fc)deepseek-v4-pro/deepseek-flash(reproduces on both)session-log-deepseekenabled with defaults (2389b65246made the upload default-on)pnpm dsh webSymptom
Long-lived sessions fail every turn with:
{"message":"DeepSeek Messages request failed (413)","code":"INVALID_REQUEST","status":413}recorded in the session log as
assistant/attempt→turn/end. A brand-new sessionworks fine with the same provider, model, and settings. The failure is permanent:
the session never recovers, no matter how many times it is retried or restarted.
The
413is returned by the provider, not by the localdsh webserver. The localHTTP bridge caps buffered bodies at 300 MiB and streams responses, so it is not the
source.
Evidence
Four sessions in one profile, same provider and model:
delivery-acceptedwatermarksf7d378ca…93b06c99…4b44862f…01155af0…(new)Two further observations:
TRANSPORT, and only the retry gets aclean
413— consistent with an oversized body dying mid-upload.providerError()inpackages/llm/llm-deepseek/src/protocols/messages/transport.tstests
isContextWindowExceededError(detail)beforestatus === 413. The result wasINVALID_REQUEST, notCONTEXT_WINDOW_EXCEEDED, so the response carried nocontext-overflow wording — this is a body-size rejection, not a context overflow.
Root cause
@deepseek-ai/dsh-session-log-deepseekcontributes thedsh_session_logrequest-bodyfield. Its contract is incremental: it sends only the log suffix after the greatest
accepted watermark.
packages/session/session-log-deepseek/src/index.tscomputes the pending suffix andsets
throughSeqto the log tail, with no bound on the suffix size:For a session that has never been uploaded there is no watermark, so
afterSeq === -1and the first request carries the complete log — 91 MiB here, embedded as a JSON
member of the request body.
This is a dead end rather than a transient failure: the provider rejects the body, no
2xx follows,
accept()never runs, no watermark is recorded, and the next attemptrebuilds the same oversized body. A session whose log grows past the provider's request
limit is permanently unresumable.
README.mddocumented exactly this gap under Known Limitations:Expected
A large pending suffix should drain across successive requests instead of being sent as
one body the provider rejects.
dsh_session_logis documented as adding zeromodel-input tokens (
README.md→ Model Experience → Token effect), so bounding itcosts no model context.
Suggested fix
Bound the contributed batch to an oldest-first prefix and make
throughSeqname thelast sequence the request actually carries:
Properties:
the next request resumes at the following one. Nothing is skipped or reordered.
~23 requests instead of failing forever.
it alone exceeds the ceiling; withholding it would freeze the watermark below it.
messages, the system prompt,and tool schemas.
Full diff: 11 files, +162/−20 —
<https://github.com/bozhang1214/deepseek-harness/commit/bff8e9a7cf>Verification
still advancing the watermark); package suite 45/45 passing,
oxlintclean,verify-config-catalogandverify-translation-pairing(919 pairs) clean.bff8e9a7cf; passes the repository's ownlefthookgates — pre-commit(translation pairing, staged lint, whitespace, vendor manifest guard) and pre-push
(
typecheck).93b06c99, 91 MB log), counting structured413 records only (
assistant/attempt/turn/endwithfailure.status == 413); thestring "413" in prose gives false positives:
maxBatchBytes: 1073741824(≈ pre-fix behaviour) → 413 reproduces: yes(3 real 413s at 20:58:47, 21:00:01, 21:00:58, and the watermark did not advance at all)
4194304→ the same session recovers: yes(watermarks resume at 21:03:02 with 3.2–3.8 MiB batches, and no further 413s)
last 413 (21:00:58) and the recovery (21:03:02)
Impact
Any long-lived session whose log exceeds the provider's request-body limit becomes
permanently unusable on upgrade — and because it fails silently on first upload, users
see only a generic 413 and are likely to misdiagnose it as a context-window problem
(as in #2770) or a compaction problem (as in #2107).
Related, but not duplicate
Here the field is not model input, and a new session on the same model works.
Note
While upgrading I also saw
UNSUPPORTED_CONTENT("DeepSeek Messages cannot representuser/tool-result content reasoning") on a pre-existing session after the
protocol: messagesdefault landed — likely a separate migration issue. Happy to file thatseparately if useful.
All reactions