Repository navigation
[bug] Subagent catalog cold reads replay full session logs and cause sustained high CPU #2606
aphelioussss-gif
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
With a large number of persisted sessions and durable subagents, opening the DSH Web UI can keep the
dsh webNode process near one full CPU core for an extended period. On my machine this caused visible thermal load even while the browser UI itself was relatively quiet.The hot path appears when
dsh-subagentresolves cold child identities. If the fastcachedSnapshot()check is not usable,resolveColdIdentity()falls directly back tosessionPersistence.inspect(), which reads/decompresses the full JSONL/Zstd log and rebuilds every registered projection.A local canary that tries
sessionProjectionCache.coldSnapshot()beforepersistence.inspect()removed the sustained thermal/CPU problem in my environment.Environment
0.1.0-rc.6dsh webon127.0.0.1:3080Symptoms
node .../dsh websustained roughly 74% to 136% CPUskill-filesystemwatching did not resolve the high CPUProfiling evidence
An 8-second Node Inspector CPU profile attributed inclusive sample cost approximately as follows:
Representative call path:
A breakpoint on
session-persistence-jsonl.loadStored()confirmed calls throughprepareCore()frominspect()while resolving cold subagent identities.Relevant code
Current behavior in
dsh-subagent:There is already a projection-cache API designed for this case:
sessionProjectionCache.coldSnapshot(id, signal). It can reuse a durable checkpoint, replay only the missing suffix, and write the refreshed checkpoint back.Local canary
I inserted the following path before the full
persistence.inspect()fallback:After restarting
dsh web, the previously sustained heating/high-CPU behavior disappeared in normal use.Expected behavior
Cold subagent listing should use the durable projection cache and tail replay path. Full log inspection and full projection reconstruction should be the final fallback for an unavailable or corrupt projection cache.
The default prepared-session cache size is only 5, so increasing that value may reduce churn temporarily but will not scale well to 100+ durable subagents. A projection-first listing path should keep catalog work proportional to lightweight metadata and missing event tails rather than total historical log size.
Suggested change
In
resolveColdIdentity():cachedSnapshot(header)fast path.cache.coldSnapshot(childId, signal).persistence.inspect()plus a full projection restore only when the projection cache is unavailable, fails, or cannot produce a subagent identity.I can provide a focused patch or additional CPU profiles if helpful.
All reactions