Replies: 3 comments
|
你这条报告的核心诉求我完全同意,而且它正是"诊断工具必须能独立于宿主运行"的理由:出故障的路径恰好是唯一的诊断入口。我们为此做了一个宿主之外的独立 CLI(诊断不加载 profile、不依赖 dsh 能启动),所以先把"我们抓得到什么、抓不到什么"讲清楚,再问一个能让我们补上缺口的问题。 抓得到的( 抓不到的(必须明说):你的失败发生在 但你那条观察已经改掉我们一处错误的信号:你说"端口绑着、systemd active、页面永远打不开",而我们的端口检查此前对"端口被 dsh 实例占用"直接判 正常 —— 那正好给出与你相反的结论。现在它会额外发一次 HTTP:绑着但无应答就报失败,并说明这与"插件在 一个能让我们补缺口的问题:那两条 host 侧致命规则具体是什么?我们的客户端侧已经有两项静态检查(client 产物语法、client 另外关于 |
|
先谢你那句"端口占用不等于服务健康"——那也是我这边最贵的一课: 下面直接回答你的问题。先给三条规则,再给一条你没提到但同样致命的第四条,最后说清楚我这边的证据边界。 一、三条规则,以及各自的失败形态规则 1(host,致命):不能在
|
| 规则 | 能否离线静态判定 | 我的看法 |
|---|---|---|
| 1. 未声明 inject 就读属性 | 能,但需要 AST 分析 | 要解析 apply(ctx) 的函数体,列出所有 ctx.<prop> 与 ctx.get('<name>'),再和返回对象上的 inject 数组比对。误报点:ctx 被改名、解构、传给辅助函数后间接访问。建议按"高置信、可加白名单"来做 |
2. 裸 harness 全局 |
能,而且最划算 | 纯词法:作用域内未声明就用 harness / styles 等已知沙箱符号 → 直接报。没有动态性,几乎零误报。这条我强烈建议加 |
| 3. client 产物是裸 ESM | 能,且你们已经有了 | 语法层即可判 export / import 出现在 client bundle 顶层 |
| 4. client 依赖宿主不提供的服务 | 能(就是你说的"声明了某字段但宿主不提供") | 把 ctx.get(x) / inject 里的名字与宿主模块表比对。同样几乎零误报 |
结论:四条里有三条完全可离线判(2、3、4),规则 1 可判但需要容忍一定的误报。
四、关于"用最小伪造 ctx 调用 apply()"
我建议不要做成判定项,但可以留作低置信提示——理由和你担心的完全一致,而且我有具体的反例支撑:
规则 1 和规则 2 都是同步抛出,所以伪造 ctx 理论上能抓到。但这次事故里,apply() 还做了大量真实工作:读配置、注册 ctx.effect、挂 llm/stream 与 agent/request 拦截器、尝试 ctx.get('tools')、注册工具。用伪造 ctx 跑,这些会大量误炸——噪音会淹没真信号,而且维护者会很快学会忽略它。
但有一个限定条件值得考虑:如果只做"语法上可静态识别的违规访问",而不是真去执行代码,就同时拿到了两边的好处——规则 2 和 4 本质上是静态可判的,不需要跑。
所以我的排序是:
- 先做规则 2、3、4 的静态判据(零误报、覆盖我这次事故三条中的两条)
- 再做规则 1 的 AST 判据(带白名单)
- "同步抛出"作为提示:可以保留,但标注为低置信、且明确写"仅当 apply() 没有真实副作用时才可信"
最后一句话给你做取舍参考:我这次事故里,host 侧两条(1、2)都是同步抛出,client 侧两条(3、4)都是静态可判。 也就是说——如果只做静态判据,这次四条里有两条能抓到;如果加上 apply() 探针,能到三条;规则 1 那条恰好是两者都需要一点才稳的那类。
五、一个必须说明的证据边界
我不希望你按一个我无法证实的模型去做实现,所以这点必须讲清楚:
我手上没有事故当时的原始插件源码了。 崩溃发生在 12:13,而所有保留的备份(bak-1225 及之后)都是 12:23 以后我已经开始修的中间态。原始坏版本被覆盖了。
上面关于三条规则的描述,全部来自事故当时抓取并已存档的 journal 原始栈(行号、错误文本、出现次数都是实录,报告正文里有全文),不是事后从代码反推的。但"第 140 行那一句原文具体怎么写"这个粒度,我给不出,只能给出它读取的属性名是 timeout。
另外,--boot-check 只做到 import 这点你讲得很实在。补充一个数据点供你判断优先级:我这次的失败 100% 在 apply() 内,import 阶段是干净的——你的探针在我这里会显示 loadable,这一点你预料得没错。所以如果 --boot-check 的失败归类里有"apply 期"这一类,它在真实事故里的命中率可能比 import 类更低。
|
这份回复把"能不能离线判"从猜测变成了工程清单,谢谢——尤其你给的四条分类与我担心的完全一致。按你的排序,我把零误报的那几条做出来了( 已实现
一个需要修正的评估:你说的"纯词法、几乎零误报"只在两轮治理之后成立。 我先按最直接的词法实现,结果在真实 profile 上连撞两次误报:
最终判据是:剥注释 + 剥字符串字面量 + 要求符号处于代码位置(后跟 关于你的证据边界:明确说"原始源码已被覆盖、规则描述全部来自当时存档的 journal 栈"——这是正确的做法,也让我的实现不受影响: 发布: |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Bug Report: Third-party plugin contract violations crash the entire plugin tree and whiteout the Web UI
Project: deepseek-harness (
@deepseek-ai/dsh)Version:
0.1.2-rc.1Node: 24.x
Platform: Linux (WSL2), systemd-managed
Severity: High — total Web UI outage
Reporter: (community user)
1. Summary
A single third-party plugin installed into a dsh profile can bring down the entire
plugin tree, not just itself. Because the failure occurs inside
apply(), cordisaborts the whole loader chain,
dsh app-bootthrows, the process exits, and systemdrestarts it — producing a crash-restart loop where the HTTP port is bound but never
serves a working page.
The plugin in question was authored against a different plugin contract (OpenClaw's)
and violated three dsh-specific rules. Two of them are host-side and fatal; the third is
client-side and silently whiteouts the page while the server looks healthy.
The core request of this report: dsh currently offers no defense-in-depth here.
A malformed third-party plugin is indistinguishable from a broken installation to the
operator, and the diagnostic surface (Web UI) is destroyed by the very failure it should
report.
2. Environment
0.1.2-rc.1/usr/local/lib/node_modules/@deepseek-ai/dsh/~/.dsh/profiles/web~/.dsh/profiles/web/node_modules/dsh-ocgo-guard/dsh-web127.0.0.1:3080Plugin files involved
package.jsondeclared:cordis.patch.ymlmounted the host half:3. Symptoms
Superficially the Web UI "would not open". The misleading part is that the service
reported healthy:
The port was bound, the unit was
active, and yet the page never rendered. Two distinctfailure modes were present simultaneously:
curlcannot complete a request.This distinction matters and is worth documenting: an operator who only checks
systemctl is-activewill conclude the service is fine and look elsewhere (DNS, firewall,the Windows
winnatport reservation issue, browser cache) — while the real fault is aplugin.
Crash-restart loop evidence
Across the incident window (
12:13–12:34), the journal shows 163 distinct processPIDs and 50 occurrences of
plugin tree failed to load.4. Root cause
4.1 Defect A — reading an undeclared service property (fatal)
lib/index.js:140accessed a cordis service property without declaring it ininject:The failing statement was at
lib/index.js:140insideapply(). The property it readwas
timeout— the cordis Proxy reported it directly:In schematic form:
cordis uses a Proxy that throws on access to a property that was not declared via
inject. The resulting error is not a normalundefined:Because this threw inside
apply(), the failure propagated up the entire loader chain:Stack, unwound:
dsh-app-bootthen rethrows and the process dies:Key observation: the error is attributed to
ocgo-guard, but the blast radius isplugin tree failed to load— every other plugin in the profile is collateral damage.4.2 Defect B — relying on a sandbox-only global (fatal, independent)
After Defect A was patched, a second, unrelated error surfaced.
lib/index.js:230referenced a bare
harnessidentifier:harnessis a closure symbol provided only to sandboxed dynamic packages (thecodeenvironment). An ordinary cordis plugin loaded fromnode_moduleshas no suchglobal:
Because the reference was reached through
new apply(...), theReferenceErrorfiredduring plugin construction — same total-outage outcome as Defect A. This error appears
55 times in the journal.
4.3 Defect C — client half shipped as bare ESM (white screen)
lib/client.jswas authored as plain ESM:But dsh's browser-side loader requires a CJS factory module:
The ESM
exporttoken is a syntax error to that loader, so the entire client bundle at/plugins/??...fails to parse:Diagnostic asymmetry worth noting: this error appears 0 times in
journalctl.It exists only in the browser console. An operator debugging purely from server-side
logs will never see it — which is exactly why the symptom "server healthy + blank page"
is so hard to attribute.
4.4 Contributing factor — sandbox-only symbols used in a third-party bundle
The client half also referenced
styles,host, andfetchas free variables. These arehost-provided closure symbols exclusive to sandboxed dynamic packages. In an external
bundle they are
undefined:This is a documentation/contract-discoverability gap rather than a separate bug: nothing
in the plugin-authoring surface warns that
harness/styles/host/fetch/setTimeoutare unavailable outside the sandbox.
5. Reproduction
apply()reads an undeclared service property:sudo systemctl restart dsh-webMinimal reproduction of Defect A (host):
Minimal reproduction of Defect C (client):
6. Diagnosis procedure (what actually worked)
The decisive insight: do not trust
systemctl is-active. Check the restart counter.Two things that cost time and should be documented:
patchReload: live(hot reload) does not save you. When the plugin tree itselffails to load, the hot-reload path fails too. Only a full
sudo systemctl restart dsh-webreflects reality. Several hot-reload attemptsreported success while the page stayed broken.
curl http://127.0.0.1:3080/returning401is normal, not a fault —the URL carries a token. Do not chase this as a bug.
7. Workaround / fix applied
Host half (
lib/index.js):try/catch, since the cordis Proxy throwsrather than returning
undefined:timerservice dependency with a built-insetTimeoutwatchdog,so the plugin degrades instead of failing when
cordis-plugin-timeris absent.harnessusage with an existence check(
typeof harness !== 'undefined') placed beforeharness.handle/harness.defineToolcalls — note thatharness.handle && harness.handle(...)doesnot work, because the bare identifier is resolved before the
&&short-circuits.Client half (
lib/client.js):stylesclosure symbol with a self-injected<style>element.ctx.get('host')insidetry/catch; whenunavailable, the badge is skipped rather than throwing.
Result: service stable,
NRestarts=0, page renders (verified via CDP).8. Requests / suggestions for the project
These are the actionable asks; the bug itself is in a third-party plugin, but the
amplification is in dsh.
apply()throwing should disablethat plugin only — not abort the whole tree and kill the process. Prefer
per-entry containment with the failure surfaced in the UI.
should still boot far enough to display the failure. Currently the failure destroys
the only place a user could learn about it.
white screen with a
SyntaxErrorvisible only in the browser console. Consider aserver-side bundle-parse check at load and a visible error overlay on the page.
exportsyntax in a client half, or an
apply()that reads undeclared properties, wouldturn a total outage into a clear, actionable message.
harness,styles,host,fetch,setTimeoutare available only inside sandboxed dynamic packages. Third-partyplugin authors following existing examples will otherwise assume they exist.
systemctl is-activereturningactivewhilethe unit loops is actively misleading. Consider exposing restart count / last boot
error via
dshstatus output.9. Appendix — raw evidence
Full journal excerpt,
12:13:58, crash iterations with increasing PID:Aggregate counts over the journal:
without injectharness is not definedplugin tree failed to loadUnexpected token(server-side)The
0on the last row is itself the finding: the most user-visible symptom(the white screen) leaves no server-side trace whatsoever.
Redaction note
Paths, hostnames, and access tokens have been genericized. dsh URLs embed a token —
they must not be pasted into issues, chats, or screenshots.
All reactions