Repository navigation
Bug: dsh web V8 heap OOM after long uptime — live sessions are never unloaded #7414
Replies: 3 comments
|
Independent repro on a different version and platform, same root object - plus the feed that sets the growth rate, three contributors your audit does not cover, and a caveat on stopgap #5. I hit this on Your item 3 confirmed, with the mechanism that sets the rateThe retention is the per-session in-memory log, and it is never unloaded - agreed. What determines how fast it fills is that every agent step appends a fresh full system-prompt render to it:
That also explains the deceleration you saw overnight: the rate is per-step, per-live-session, not per-wall-clock, so it is consistent with your "no new sessions opened" observation. Measured growth, one process, two snapshots 34 minutes apart:
Four aborts in this log, all the same signature ( Three contributors beyond your list
Caveat on stopgap #5: do not expect the bigger cap to buy timeRaising Happy to post the full snapshot report (retainer walks, node-group anatomy, ownership attribution) if it would help - and agreed that #1 (an unload path that actually reaches the dropped |
|
Reproduction on macOS at Setup: macOS 15.7.5 (arm64), Node v22.23.1, source checkout at tag Crash: after 2.9 days of uptime, Growth, from session logs (host-side time between At this version the per-step system-prompt append is deduplicated: |
|
Independent field reproduction on 0.2.0-rc.2 (latest stable), with growth numbers taken from systemd's own cgroup counters. Environment: dsh 0.2.0-rc.2 (global npm install, Crash, 2026-10-09, instance age 1 d 20 h 33 m: systemd recorded One measurement that may matter for triage: the default cap on this host is 2.1875 GiB, not ~4.14 GB — Growth, read out of systemd's per-unit counters (
(The first two rows are a host-side incident of ours that is now fixed — the last two rows are the harness-side growth this thread is about, and the 14.4 GB instance has the same shape as the weeks-long case in the original report.) Session volume here, for comparability with the "11 workspaces / 105 MB" figure in the OP: 3 workspaces, 834 MB of zstd JSONL across 1 185 session files, largest single session 38.9 MB compressed. Consistent with the per-opened-session accumulation model — this deployment is mostly idle overnight and still climbs. We applied stopgap #5 verbatim ( Question: has anything from the Ask list (web-session unload path / the dropped |
Uh oh!
There was an error while loading. Please reload this page.
dsh web: V8 heap OOM after long uptime — live sessions are never unloadedSuggested labels:
kind/bug,area/session,area/web.TL;DR
Every session opened or resumed in the
dsh webGUI becomes a permanently live in-memory Agent +full event log; the
AgentHandle.dispose()capability is dropped at open time, client disconnectdoes not unload anything, and no idle/refcount eviction exists. Over weeks of GUI use the heap
crosses Node's 4 GB default cap and V8 aborts.
Environment
node --import tsx/esm apps/cli/src/bin.ts web(viapnpm dsh web)Crash
heap_size_limiton this host (measured): 4144 MB.Usage profile at crash time
~/.dsh/sessions/, 105 MB on disk (zstd JSONL); largest singlesession: 23,723 events / 21.8 MB decompressed JSON (robot_client workspace, still active).
accumulation, not one allocation.
recoverable — no core was captured by apport for the V8 abort).
Observed RSS after restart (same workload resumed)
(Browser reconnects, restores session list, resumes the 23.7k-event session.)
Root cause (confirmed by code audit)
Live sessions are never unloaded in
dsh webmode — every session ever opened or resumed in theGUI stays fully in memory until process exit.
Evidence chain:
ensureSession()(packages/host/apiproxy/src/api-proxy.ts:1618-1707),which resumes/creates the agent via
ctx.agents.resume(...)/ctx.agents.create(...)andreturns only
.agent— theAgentHandle.dispose()capability is dropped(api-proxy.ts:1657-1661, 1670-1678).
AgentRegistrykeeps one entry per session in a plainMap(packages/core/agent/src/index.ts:257);SessionStorelikewise (packages/core/session/src/index.ts:792-793, plus the module-levelattachmentsMap at :932). Removal happens only viadetach()/dispose()or owner-fiberteardown — none of which web sessions ever get.
Sessionholds its full event log in memory:private log: SessionEvent[] = [](
packages/core/session/src/index.ts:425-428); "a resumed session's constructor seed is its fullstored log" (index.ts:456-458).
packages/client/connection/src/websocket-downlink.ts:107);the persistence coordinator's
retire()path is gated onsession/disposed(
packages/session/session-persistence/src/coordinator.ts:1132, 1140-1161), which never fires forweb sessions. The coordinator itself documents the posture: "State created through the public
create()/load()API has no owner" (coordinator.ts:230-234).core/agent, core/session, session-persistence). The only bounded structure is the prepared-session
LRU (
DEFAULT_PREPARED_SESSION_CACHE_SIZE = 5, coordinator.ts:27, 609-623), which boundsunpublished preparations, not live sessions.
Measured cost: parsing the 23.7k-event session log in isolation costs ~49 MB heap in plain
JSON.parse— before the agent loop, projections, surface manager, and per-client fan-out multiply it.Observed RSS after restart (same workload resumed) — update
Disk growth of the active session over the same spans: 7.49 → 7.96 MB (first hour), 8.96 MB by next
morning — server RSS grows far faster than session data itself. Growth decelerated overnight while no
new sessions were opened (813 → 974 MB over 14 idle hours), consistent with the per-opened-session
accumulation model above rather than a time-based leak: the crashed instance had weeks of GUI
browsing across 11 workspaces to reach the 4 GB cap.
Contributing factors (full audit, ranked)
packages/api/remotes/src/agent-lookup.ts:142-186resumes every cold session the GUI touchesinto a permanent
AgentRegistryentry; each entry = full in-memory event array + one Cordisagent fiber with the preset's whole plugin set + projection cells (
WeakMap<Session, UnitCell>,packages/session/session-projection/src/index.ts:150— GC-eligible only when the Session isunreachable, which never happens) + title work + telemetry
adoptedSet(
packages/session/session-telemetry/src/coordinator.ts:66, deleted only onsession/disposed).preparations.ts:33, 285-298withcoordinator.ts:244-252: up to 5 ready entries, each holding a full parsedSessionand afrozen inspection referencing the same events. Browsing the 22 MB session leaves its parsed log
pinned in the LRU after the request returns; eviction is correct but the working-set floor is
large and GC pressure matches the "Ineffective mark-compacts" signature.
FrameQueueon the all-sessions firehose —api-proxy.ts:413-437(
private buffer: F[] = [], push with no high-water mark); the per-connectionsession/eventlistener (
api-proxy.ts:3475-3489) pushes every committed event of every session to everyconnected client — no per-session filter, no backpressure. Clean aborts are handled
(
:3530-3533), but a slow/stalled client (half-open TCP, paused reader) accumulates withoutbound while its stream stalls. Grows with (events × clients).
Audited and cleared: subagent children (disposed after run,
packages/subagent/subagent-in-process-driver/src/index.ts:195-199), workflow workers(
worker.terminate(),packages/workflow/workflow-worker-thread/src/host.ts:200,244), todo/plan/compaction state (WeakMap on Session lifetime), projection-cache (disk-backed), history replay
(re-reads per request, no cache), apiproxy dedupe maps (deleted after settle).
Ask
the currently dropped
AgentHandle.dispose()(packages/core/agent/src/index.ts:159-174) so thesession/disposed→retire()chain (coordinator.ts:1132-1161) actually runs.FrameQueuewith a high-water mark + drop/backpressure policy(
api-proxy.ts:413-437), and consider per-session subscription filters.long-running
dsh webdeployments.NODE_OPTIONS=--max-old-space-size=...andrestart on RSS growth.
All reactions