You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
session-query: a "bounded" readEvent deep-clones the whole session log — synchronous main-thread stall on large sessions
#8277
The bound constrains only the returned value, not the work performed. Reading one event costs a deep copy of the entire session log.
structuredClone is synchronous. On a large session this is a multi-second-to-minute main-thread block, during which the host serves no HTTP at all — the whole Web UI (Settings, other sessions, send/stop buttons) stops responding, with no error output anywhere.
Environment
DSH 0.1.5-rc.2 (workspace install), Node v24.9.0, Windows
13,820 ms structuredClone @ node:internal/worker/js_transferable:112 <- 92%
893 ms (garbage collector)
788 ms snapshotLive @ dsh-session-query/lib/index.js:299
217 ms structuredClone @ (no url)
The main thread was 100% saturated for the whole sampling window, 92% of it inside structuredClone.
Root cause
readEvent is supposed to produce one event plus a small window, but the path it takes is:
asyncreadEvent(request,signal){
...
returnthis._readEvent(sessionId,seq,before,after,signal);}async_readEvent(sessionId,seq,before,after,signal){constloaded=awaitthis._corpus.load(sessionId,signal);// whole log + full deep cloneconsttarget=loaded.events[seq];// only this one is needed
...
}
Note that the same file already contains a non-copying counterpart:
// sourceLive() in the same file: no clone at allfunctionsourceLive(session: Session): LogicalSessionSource{return{header: session.header,events: session.snapshotEvents()}}
and the persistence layer already uses a freeze-once / share-without-copying model:
The backend deep-freezes each event graph once before memoization, so later handle reads reuse it without copying.
The package documents the cost itself:
Exact reads replay whole logs — readSession, readSurface, filterEvents, and event traces load and validate the complete logical log, so very large histories pay full inspection per call
and this comment shows the migration is known and deferred:
// oxlint-disable-next-line typescript/no-deprecated -- Existing Session history read; migration deferred.
In short: the machinery for avoiding a whole-log copy already exists in this codebase (sourceLive, plus the persistence layer's freeze-and-share). snapshotLive / readEvent is the path that was missed.
Scope beyond a single consumer
The incident was triggered by a third-party plugin that called readEvent once per repaired sequence on agent/pre-step — O(items × logSize), i.e. 1,311 × 31.2 MB. That plugin switched to the session/event feed in its own 2.2.1 and is no longer affected.
But the amplifier described here lives in core and remains reachable: any plugin, tool, or UI path that consumes readEvent / readSession / readSurface / filterEvents against a long session pays a full synchronous deep clone for what is documented as a bounded read.
Suggested changes (by cost/benefit)
Share instead of copying.session.snapshotEvents() is already the shareable path sourceLive() uses, and the persistence layer already freezes each event graph once for copy-free reuse. A "detached" result need not be a fresh deep clone — sharing frozen events is enough, or fall back to a lazy / copy-on-write snapshot.
Make the bound real.readEvent should not route through load(). It needs one event plus a small window: a slice read on the persistence handle, or the observation path the docs already describe as never copying the log for header/cursor/projection-only consumers. At minimum, a batch of repairs should take one snapshot, not one per item.
Bound the synchronous work. Even with 1 and 2, a pathological log can still stall the loop; consider chunking with yields, or a cap per operation.
Reproduction
No private data is required. Any session with a few thousand events and a few hundred tool/result events reproduces this as soon as a consumer calls readEvent once per event. Our instance was 6,732 events / 31.2 MB / 1,311 tool/result events.
Notes
This post contains no session content. The attached CPU profile carries only function names, module paths (user name and workspace path replaced with placeholders), and sample timing.
We are not reporting the third-party plugin — it has been fixed on its own side. We are reporting the core cost model it tripped over.
Full CPU profile: attached below (sanitized — the user name and workspace path are replaced with <user> / <workspace>; it contains only function names, module paths, and sample timing).
中文版(点击展开)
摘要
sessionQuery.readEvent() 的文档是:"one full event plus a bounded raw-log context window"。
13,820 ms structuredClone @ node:internal/worker/js_transferable:112 ← 92%
893 ms (garbage collector)
788 ms snapshotLive @ dsh-session-query/lib/index.js:299
217 ms structuredClone @ (no url)
The backend deep-freezes each event graph once before memoization, so later handle reads reuse it without copying.
包文档本身也承认这个代价:
Exact reads replay whole logs — readSession, readSurface, filterEvents, and event traces load and validate the complete logical log, so very large histories pay full inspection per call
而源码里这条注释说明该迁移是已知且被推迟的:
// oxlint-disable-next-line typescript/no-deprecated -- Existing Session history read; migration deferred.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
sessionQuery.readEvent()is documented as "one full event plus a bounded raw-log context window".Internally it calls
load(), whichstructuredClones every event in the log:The bound constrains only the returned value, not the work performed. Reading one event costs a deep copy of the entire session log.
structuredCloneis synchronous. On a large session this is a multi-second-to-minute main-thread block, during which the host serves no HTTP at all — the whole Web UI (Settings, other sessions, send/stop buttons) stops responding, with no error output anywhere.Environment
0.1.5-rc.2(workspace install), Nodev24.9.0, Windowstool/resulteventsWhat we observed
After sending one message in that session:
turn/end { reason: { kind: "interrupted" } }Evidence
(1) Call stack at the moment of the stall (captured with
Debugger.pause, 10 frames):(2) Self time over a 15-second CPU sample:
The main thread was 100% saturated for the whole sampling window, 92% of it inside
structuredClone.Root cause
readEventis supposed to produce one event plus a small window, but the path it takes is:Note that the same file already contains a non-copying counterpart:
and the persistence layer already uses a freeze-once / share-without-copying model:
The package documents the cost itself:
and this comment shows the migration is known and deferred:
// oxlint-disable-next-line typescript/no-deprecated -- Existing Session history read; migration deferred.In short: the machinery for avoiding a whole-log copy already exists in this codebase (
sourceLive, plus the persistence layer's freeze-and-share).snapshotLive/readEventis the path that was missed.Scope beyond a single consumer
The incident was triggered by a third-party plugin that called
readEventonce per repaired sequence onagent/pre-step—O(items × logSize), i.e. 1,311 × 31.2 MB. That plugin switched to thesession/eventfeed in its own 2.2.1 and is no longer affected.But the amplifier described here lives in core and remains reachable: any plugin, tool, or UI path that consumes
readEvent/readSession/readSurface/filterEventsagainst a long session pays a full synchronous deep clone for what is documented as a bounded read.Suggested changes (by cost/benefit)
session.snapshotEvents()is already the shareable pathsourceLive()uses, and the persistence layer already freezes each event graph once for copy-free reuse. A "detached" result need not be a fresh deep clone — sharing frozen events is enough, or fall back to a lazy / copy-on-write snapshot.readEventshould not route throughload(). It needs one event plus a small window: a slice read on the persistence handle, or the observation path the docs already describe as never copying the log for header/cursor/projection-only consumers. At minimum, a batch of repairs should take one snapshot, not one per item.Reproduction
No private data is required. Any session with a few thousand events and a few hundred
tool/resultevents reproduces this as soon as a consumer callsreadEventonce per event. Our instance was 6,732 events / 31.2 MB / 1,311tool/resultevents.Notes
Full CPU profile: attached below (sanitized — the user name and workspace path are replaced with
<user>/<workspace>; it contains only function names, module paths, and sample timing).中文版(点击展开)
摘要
sessionQuery.readEvent()的文档是:"one full event plus a bounded raw-log context window"。但它内部调用
load(),而load()会逐个事件structuredClone整个日志:也就是说:
bounded只约束了返回给调用方的那一段,没有约束它实际做的工作量。 要一个事件,代价就是一整个会话日志的深拷贝。structuredClone是同步的。在大会话上,这是一次几十秒到几分钟的主线程阻塞 —— 期间宿主无法响应任何 HTTP 请求,整个 Web 界面(设置页、其他会话、发送/停止按钮)一起失去响应,且没有任何错误输出。环境
0.1.5-rc.2(工作区安装),Nodev24.9.0,Windowstool/result事件现象
在该会话中发送一条消息后:
turn/end { reason: { kind: "interrupted" } }证据
① 主线程暂停时的调用栈(用
Debugger.pause抓取,10 帧):② CPU 采样 15 秒的自身耗时分布:
主线程在采样窗口内 100% 饱和,92% 花在
structuredClone。根因
readEvent的职责是「一个事件 + 一小段窗口」,但它走的路径是:值得注意的是,同一个文件里已经有一条不拷贝的对照路径:
而且持久化层已经采用「冻结一次、之后共享」的模型:
包文档本身也承认这个代价:
而源码里这条注释说明该迁移是已知且被推迟的:
// oxlint-disable-next-line typescript/no-deprecated -- Existing Session history read; migration deferred.简单说:避免整日志拷贝的机制在这份代码里已经存在(
sourceLive+ 持久化层的冻结共享),snapshotLive/readEvent是漏网的那一处。影响范围不止一个消费者
触发我们这次故障的是一个第三方插件(在
agent/pre-step上逐条调用readEvent,复杂度O(待修条目数 × 日志大小),即 1,311 × 31.2 MB)。该插件已在自己的 2.2.1 版本改用session/event事件流,那一侧已经闭环。但这里描述的放大器在核心里,且仍可达:任何消费
readEvent/readSession/readSurface/filterEvents的插件、工具或 UI 路径,只要面对长会话,都会以「一次 bounded 读取」的代价触发一次整日志同步深拷贝。建议(按性价比排序)
session.snapshotEvents()已经是sourceLive()使用的可共享路径,持久化层也已经把每个事件图冻结一次供后续无拷贝复用。「detached 结果」未必需要一次全新的深拷贝 —— 共享已冻结的事件即可,或者退一步做惰性/写时复制。readEvent不应该经过load()。它只需要一个事件加一小段窗口:可以走持久化句柄的切片读,也可以走那条文档里写明「header / cursor / 投影类消费者永不拷贝日志」的 observation 路径。最起码,批量修复应当只取一份快照,而不是每条一次。复现
不需要任何私有数据。 任何具备「几千个事件 + 几百个
tool/result」的会话,只要有消费者对每个事件各调一次readEvent,即可复现。我们这次的具体规模是 6,732 个事件 / 31.2 MB / 1,311 个tool/result。说明
freeze.sanitized.cpuprofile
All reactions