Follow-up to #4807: live hot-path evidence correlated with our current locally patched rc.2 profile #4911
Replies: 3 comments
|
The handbook's Session heap guide already covers the synchronous cold-restore/event-loop risk, separates restore duration from event-loop delay, and explicitly treats #4807 measurements as incident evidence rather than a universal benchmark. Relevant guide: https://github.com/sandbaseai/deepseek-harness-handbook/blob/main/docs/en/operations/session-heap-growth.md |
|
v0.5.265 refines the handbook's Session restore guide with this discussion's intermittent live-sampling evidence: pulsing hot paths must be sampled across multiple rounds, and projection-miss latency must remain distinct from restore cost. Release: https://github.com/sandbaseai/deepseek-harness-handbook/releases/tag/v0.5.265 |
|
Thanks for pointing us to the Point by point:
These statements are consistent with the explicit evidence limits in #4911. Cannot claim from the handbook cross-reference or #4911 alone: an independent reproduction or validation; official or upstream-canonical status within the DeepSeek Harness project; a unique JavaScript function or RPC endpoint; that projection miss occurred in the sampled window; causality for send latency; a universal benchmark or current-upstream behavior; correctness of our proposed fix; or DeepSeek Harness maintainer acceptance, adoption, or a shipped upstream fix. |
Uh oh!
There was an error while loading. Please reload this page.
Context: in discussion #4807 we reported synchronous giant-session restore/event-loop stalls and four linear-scan sites. Since then we sampled the live desktop-managed
dsh webprocess again. This adds live-process evidence correlated with the current local post-fix resolver/disk state; we did not inspect the process's ESM module cache.Method
dsh web, listener 127.0.0.1:3080 (PID/PPID/PGID frozen, launched 2026-08-28 07:24:28 +0800)./usr/bin/sampleon the live PID, 5s rounds (three rounds), plus a 1stopseries in the same window.Findings
ArraySomeframes account for 80.6% / 96.9% of main-thread samples, with native frames along HTTP → V8 →Builtins_ArraySome/Builtins_ArraySomeLoopContinuation→Runtime_GetProperty/LookupIterator→FastPackedFrozenObjectElementsAccessor. Three later 5s rounds caught the same signature in only one round (2.3%), while the others were mainly parked inkevent. A 1stopseries in the same window showed 46.1% → 99.8% → 99.1% → 100.9%. So this was a pulsing hot loop in the sampled window, not a permanent peg.dsh-host-apiproxy/lib/index.js(SHA-256a7ea5a8b…), and global/profile copies are byte-identical. This evidence strongly disfavors the old7f9…series, but is not in-process module-cache proof.session.events.some(...)call site we identified isagentPreset.select's blank check; on a miss it can scan the full frozen log. The client source contains a staged-preset path that can reach this check. The native sample did not identify the endpoint.session/event;peekStateOfdoes not lazily build; andsession/createdis published before the first live event. A freshly restored session therefore misses atsession/createduntil an explicit build or the first live event creates the cell. If the cell is still missing, that first event may trigger a synchronous prefix fold over the log, which we suspect could contribute to send latency; this is distinct from theArray.somepulse above.Our candidate fix (designed, not yet applied — happy to share a patch draft)
For the preset-select blank check: read a warm snapshot with a current watermark when available; on miss, use a bounded async blank resolver (
eventAt/seq + WeakMap prefix cache + 8ms / 1024-event yields + in-flight dedupe). Explicitly avoid astateOf()miss fallback, which would relocate the synchronous full fold.Boundaries
All reactions