Replies: 3 comments
|
Independent confirmation of both defects, on different hardware and a different model. Same two failure modes, same wall-clock cost. Setup: dsh 0.1.1-rc.2 -> local relay -> LM Studio 0.4.21 (llama.cpp 2.29.1), Qwen3.8-27B-Uncensored Q4_K_M with mmproj, Mac Studio M1 Max 32 GB, ctx 32768, speculative draft Qwen3.5-0.8B. Defect 2, WebPConfirmed with a controlled A/B against the same server. Same image, two encodings: The base64 is valid in both cases. I captured the outgoing request with a proxy: Cost here matched yours: A workaround that does not require patching dsh. Anyone routing through a proxy can convert in flight. About thirty lines: walk One number worth knowing: a 31 KB WebP becomes a 667 KB PNG, so the request grows from ~42 K to ~889 K characters. Irrelevant on a local link, less so on a metered one. JPEG would be smaller but lossy, and screenshots are the main use of Defect 1, compactionAlso reproduced, with one measurement that may narrow it. My server caps output at 4096 and Raising So beyond the recovery request being larger than the one that overflowed, there is no feedback when the summary is cut a few tokens from the end. The retry replays the identical request and produces the identical truncation. A single retry with a raised budget, or simply surfacing the shortfall, would break the loop. Users who cannot fork can at least raise Happy to supply the full server logs or the relay code if either is useful. |
|
I reproduced a related but distinct non-converging compaction failure on Observed sequenceOne automatic compaction succeeded. After that checkpoint had become the oldest visible surface node, later pressure checks repeatedly selected only that checkpoint: The failed transaction correctly left the surface unchanged. However, every later This was not caused by a missing context-window declaration or by auto-compaction being disabled. The pressure listener was running and Why it does not converge
That error exits Suggested behaviorFor the specific non-shrinking-summary failure:
The important part is expanding the selected source range, not weakening the shrink check. A conceptual regression fixture would be: I can prepare a minimal test or a source-based replacement package if that would help. I am adding this here because repository Issues are disabled and this discussion already records the same error string, although the selection/retention cause differs from the overflow-request-size cause in the original post. |
|
Update: I completed a source-based replacement and a real-session validation. The initial “expand the range after a non-shrinking summary” diagnosis was correct but incomplete for this particular session. Additional root causeThe pathological session contained two tool calls from a historical step that already had Implemented solutionI published an MIT, source-based replacement derived from the official
The replacement:
The fifth rule is deliberately narrow: an orphan call from a still-open step is not synthesized, and no Real-session resultUsing Node 24 and the Web profile’s real LLM adapter, a gated repair of a copy produced: The report passed semantic validation, replace generation advanced, the output persisted and reloaded at 12,900 tokens, and a subsequent message could be appended. The source session SHA-256 was unchanged. After installing the validated output, the Web UI showed about 8% context usage and the real model completed the next turn. One deployment detail was also important: installing the package is insufficient. The profile must explicitly disable This is a user-space workaround and reproduction, not a claim that upstream has fixed the issue. I hope the closed-step orphan-call fixture and the non-shrinking range retry policy are useful as upstream regression cases. |
Uh oh!
There was an error while loading. Please reload this page.
Two defects found while running dsh against a local llama.cpp route
Running the harness against a local LM Studio (llama.cpp) endpoint, 58 deduplicated turns in one workspace ended 41% in error. Most classes were my own misconfiguration, but two look like harness defects. Both have a fix prepared on a fork branch —
CONTRIBUTING.mdsays external PRs are not accepted, so these are linked for reference rather than proposed.1. Context-overflow recovery sends a request larger than the one that overflowed
compactIfNeededselects the overflow region withretainTokens = 0, sosummarizeWithLlmreplays essentially the whole surface that just overflowed — behind the same system prompt and tool set — plus a 459-token compaction instruction, while reservingmaxTokens(default 8192) for output. The recovery request is strictly larger than the request the provider just refused.Priced in a regression test against a 4000-token window: the conversation request overflows at 4916 tokens, the recovery request costs 5617.
In the session logs this appeared as 28 of 74 compaction attempts (37.8%) ending with an error — 18 ×
summarization truncated at the token cap (incomplete checkpoint), 3 ×summarization produced no text summary content, 2 ×summary is not smaller than the shadowed content, 5 × provider errors leaking into compaction. One turn retried compaction ~28 times over 2.6 hours of wall clock.A second, smaller defect gates the first. llama.cpp reports overflow as
{"code":500,"message":"Context size has been exceeded.","type":"server_error"}. The structured classifier requiredlength/windowas the bound with no copula between bound and verb, so this wording fell through to the generic 5xx rule, was labelledSERVER, retried to exhaustion, and never reached overflow recovery at all. Fixing the classification is what makes the first defect reachable.Branch: https://github.com/nnnet/deepseek-harness/tree/fix/context-size-overflow-classification — description at nnnet#1
2. Images are sent in encodings the routed endpoint cannot decode
ImageRequestPolicycarries only pixel and byte budgets, soattachment-localpicks WebP for any transparent source and passes stored WebP through untouched, with no knowledge of what the routed endpoint decodes. LM Studio decodes via llama.cpp'sstb_image, which has no WebP support, and reports the decode failure as400 "'url' field must be a base64 encoded image"— wording that reads as a malformed field rather than an unsupported format, which sent the first investigation down the wrong path.Observed: one 176×225 WebP attachment caused 7 consecutive failed turns over 67 minutes. Every later turn, including a bare
continueand unrelated new user messages, resent the same history carrying that image and failed identically. It ended only when the provider was switched by hand.Branch: https://github.com/nnnet/deepseek-harness/tree/fix/request-image-media-type-negotiation — description at nnnet#2
Not defects, for completeness
Three classes looked like harness bugs and were not, in case the wording is worth improving anyway:
contextWindowinheritDEFAULT_CONTEXT_WINDOW = 262144while the local server's real per-slotn_ctxwas 32768. Compaction's0.8 × contextWindowthreshold then never fires and the request dies on the server's hard error. The existingdefaultContextWindowconnection field fixes it, but the default is a sharp edge for local-server routes.gpt-oss-120band the local Qwen3.8 Jinja template both reject reasoning effortoff— 8 turns lost before the config dropped that level. The harness classified these as non-retryable and failed fast, which is correct.A related observation rather than a bug: nothing detects that N consecutive turns failed with byte-identical errors. That is what turned defect 2 into 67 wasted minutes, and separately a turn burned 98 steps repeating
unknown tool "str_replace_editor"(17×) andmissing required property "description"(~97×) without changing course. Loop hygiene felt like the natural home, but it is out of scope of both branches above.All reactions