Repository navigation
单个损坏的会话日志会导致整个 dsh 无法启动(0.1.6-alpha.2 实测) #7161
Replies: 2 comments
|
Your chain is right, and the matrix is the most useful part of the report — six states, measured, with source attribution. I re-checked all six against Two refinements, then the part where I'd argue against your proposal. The tolerance set is four cases, expressed in two different vocabulariesThe skip decisions are not all made in
So rows 2, 4 and 5 fail not because the listing "forgot a case" but because they throw a bare
The asymmetry is therefore vocabulary-shaped: skip is spelled Why I would not ship "treat a failed read as a skip"The The distinction worth preserving is:
That makes your "minimal change" slightly larger than a catch, but it is the version that does not trade a loud failure for a quiet one. Row 4 is predictable before the bootWorth highlighting because it is the row with the cheapest fix and the one your matrix proves the most about: the identity refusal is path recomputation, and What I built on thisI maintain Exit code 1, so it can gate a launch. npm: https://www.npmjs.com/package/@argszero/cordis-plugin-session-audit, source: https://github.com/argszero/cordis-plugin-session-audit. The path derivation is verified against the harness's own One caveat, since the tool exists precisely because of it: On
|
|
Thanks — and thanks for re-checking all six rows against You are also right about the proposal, and I will not defend it: a blanket Row 4 being predictable pre-boot is the direction I am building in, from the plugin side. I maintain dsh-plugin-ops — startup-lifecycle protection for dsh plugins: pre-boot static checks (patch-row resolution through the runtime generation, dependency drift, peer gaps), boot-failure attribution via the CLI's saved startup diagnostics, and reversible repairs. Session containers are the surface we do not audit; The On |
Uh oh!
There was an error while loading. Please reload this page.
环境
$DSH_HOME/sessions/<workspace>/<session-id>/session*.jsonl.zstd问题怎么暴露的
一次文件恢复工具的事故把一个会话日志写成了全 null(恢复工具的常见产物)。之后
dsh web完全无法启动;唯一的办法是从报错信息里找到文件、手动移出sessions/目录。实测行为矩阵(隔离 DSH_HOME,dsh 0.1.6-alpha.2)
中止发生在哪里(源码链路)
listArtifacts(dsh-session-persistence-jsonl)→readGenerationHeader→readFirstZstdLine在帧魔数损坏时抛错;路径/header 不匹配在assertStoredIdentity抛错。错误经WorkspaceRegistry.listStoredHeaders直接向上传播(workspace 的sessionKnown注释明确了这一设计:"a failingsessionPersistence.list()propagates so storage faults never masquerade as an unknown session"),在 workspace 插件的初始化(cordis.init 生命周期)阶段失败,导致必需的 bootstrap Include 失败,进而整树中止。为什么这看起来像一处不一致
listArtifacts其实已经容忍读不出的会话:外部格式跳过、首行 JSON 损坏返回 undefined、torn/空文件正常启动。代码注释写明了设计意图:"Listing skips a foreign format while opening its id still refuses with the selected physical location."(列出时跳过、显式打开时严格拒绝)。帧魔数损坏与 identity 不匹配属于同一类「这个列表项读不出来」,却直接中止了列出——于是单个坏文件拖垮整个应用,而列出路径本已有既定的跳过机制。建议的修改(最小)
在
listArtifacts的 catch 中,把单个会话的读取失败当作跳过项处理(记录一条带文件名的警告),显式 open/attach 保持严格(它已经会用物理位置拒绝)。这样对用户操作的 fail-loud 不变,同时让启动对单个坏文件有韧性。另外建议加一个 CLI 修复入口(如dsh sessions prune --corrupt),让遇到这个问题的用户有内置出路。背景
在做生态的插件运维工具(dsh-plugin-ops)时踩到的;如果方向认可,我们很乐意帮忙测补丁。
All reactions