write/edit reject a repeated same-mode escalation with a blank justification while bash accepts it
#8848
Replies: 5 comments
这是跨工具不一致——而且"重复当前模式"本就不该报错1. 你指出的两点我认同,其中第二点更容易被忽略
建议把诉求写成:
2. 为什么第 ② 条是根本"再次请求当前已有的模式"在语义上是幂等的:会话已经是 3. 你已声明"已在 upstream
|
|
补充数据(原发帖人):你要的三项,外加一处必须更正的表述 0. 先更正:不存在"空白 justification 文本"没有任何代码路径会渲染出空白的 justification。拒绝时的报错是固定字面量,它从不回显取值;空白的是入参
⇒ 结论:这不是"缺值 bug",也不是"读错字段",而是文案有意省略取值。所以「the error text with the blank justification」这个对象不存在;可以引用的是入参为空这一事实。 1. 你要的 (i):上游 file + symbol + line(并附 0.2.0-rc.2 已安装 bundle 行号)
两边的具体差别就一句话:bash 在校验之前因为"同模式"直接 附带一条你要的旁证: 测试覆盖差异也成立:
2. 你要的 (ii):原始调用与返回(逐字)先更正一处:全盘 102 个 journal 里, 真实存在的原始记录(session 注意日志是多帧 zstd(该文件 893 帧): 我独立跨全部 journal 做的聚合(同一脚本、同一口径):
跨 102 个 journal / 8,618 次 原报告自报的 42 / 21 / 19 / 2 / 0 / 1 是从一份部分导出里统计的,方向正确但已过时(同一会话现在 3. 你要的 (iii):
|
| 包 | rc.2 → alpha.1 | 结论 |
|---|---|---|
dsh-sandbox |
空 diff(逐字节相同) | 未改 |
dsh-tool-bash |
只有参数顺序/描述文案;gate 仍在 lib/index.js:238 |
仍存在 |
dsh-tool-fs |
只有描述文案;lib/index.js:1123 仍无条件先校验 |
仍存在 |
dsh-tool-pwsh |
参数顺序/文案 | 仍存在 |
dsh-tools(run_code) |
ptc.js:297、index.js:1187 未变 |
仍存在 |
四条 escalation 路径行为零变化,alpha 只调整了参数顺序与提示文案。另附口径更正:0.2.1-alpha.1 只挂在 npm 的 alpha tag(latest/next 仍是 0.2.0-rc.2)。
4. 额外一条:schema 层的不对称(模型为什么会这么调)
四个工具的 sandbox_permissions / justification 描述:
fs(dsh-tool-fs/lib/index.js:1101、:1105):sandboxPermissionsDescription("operation")+ "Required with sandbox_permissions: one sentence for the user explaining why this exact file operation needs the wider access. …"bash(:517、:521):同一 helper、换成"command",句子逐字相同;pwsh(:487、:491)同样。run_code(dsh-tools/lib/types/ptc.js:61-62):措辞不同,但也写 "requires justification and approval"。
不对称在于"缺了一条":三处 shell/fs 描述逐字相同、都声明理由 "Required"——而对 bash 来说这句是假的(同模式空理由会通过);四个工具没有一个提到"重复当前模式无需理由"。这解释了模型为什么会习惯性带空理由。
5. 语义澄清(供你判断改动边界)
- 拒绝需要同时满足"同模式"且"理由为空";同模式 + 非空理由在 write/edit 上是通过的(7 + 22 次、0 个相关错误)⇒ 所以
tool-fs/src/sandbox.ts:88那一行是唯一障碍。 sandbox_permissions与justification恰好只给一个 → 按设计拒绝(正确,别改)。- 严格更宽的模式 + 空理由 → 继续拒绝(正确,别改;
tool-bash/tests/tools.spec.ts:741已覆盖)。 - fs 自己的文档注释(
dsh-tool-fs/lib/index.js:1112,"Repeating the standing mode requires no approval")被它自己的:1123否掉了 —— 这条注释与escalation.ts的文档、以及那个"两个 enforcing family 校验一致"的注释都是过期表述。
6. 边界(未能确认的部分)
- 本机没有做 live 复现:当前会话禁用 escalation 参数(审批提示全局关闭),所以 §1/§2 的结论基于静态代码 + 已存日志,没有现场重放。
edit的行为是推断(0 条观测)。- 报告里"已在
master验证"这一点:我核到的 5 个 master 行号全部正确,但那是"行号与源码一致",不是"我确认你当时跑过"。 app.asar是单个 121 MB 归档文件(不是目录),但可按字节搜索,缺陷三行在其中各出现一次;其runtime.json的desktopVersion=0.2.0-rc.2。0.2.1-alpha.1是从哪个 commit 构建的未确认(我只比对了已发布的 alpha 包与 master 源码中引用的表达式)。
附:复核命令(可自行跑)
# 共享校验器与报错文案
grep -n 'invalid justification' <install>/node_modules/@deepseek-ai/dsh-sandbox/lib/index.js
# bash 的同模式豁免
grep -n 'effectiveMode' <install>/node_modules/@deepseek-ai/dsh-tool-bash/lib/index.js
# fs 的无条件校验
grep -n 'validateEscalationArgs' <install>/node_modules/@deepseek-ai/dsh-tool-fs/lib/index.js
# 版本一致性
npm pack @deepseek-ai/dsh-sandbox@0.2.1-alpha.1 && diff <(tar xzOf *.tgz package/lib/index.js) <install>/…/dsh-sandbox/lib/index.js
Supplementary data (original reporter) — English versionEnglish version of my earlier comment. One framing correction first, because it changes what the reviewer should look for. 0. There is no "blank justification text"No code path can render a blank justification. The rejection is a fixed literal that never echoes the value; what is blank is the input argument.
⇒ Verdict: not a missing-value bug, not a wrong-field read — the message deliberately omits the value. The object "the error text with the blank justification" does not exist. 1. Reviewer ask (i): upstream file + symbol + line (with the shipped-bundle line for each)
The concrete difference, in one sentence: bash returns before validation when the request repeats the current mode; the fs family validates unconditionally, so Corroboration: commit Test-coverage gap also holds:
2. Reviewer ask (ii): the raw calls and responsesOne correction first: an exhaustive scan of all 102 session journals finds zero The calls that do exist (session Aggregated by me, independently, across every journal:
Across 8,618 The report's own figures (42 / 21 / 19 / 2 / 0 / 1) come from a partial export of that session; the direction is right but the numbers are stale (bash same-mode+blank is now 65 in that session, 72 overall). Journal format note: these files are multi-frame zstd (893 frames in that one). 3. Reviewer ask (iii):
|
| Package | rc.2 → alpha.1 | Status |
|---|---|---|
dsh-sandbox |
empty diff (byte-identical) | unchanged |
dsh-tool-bash |
parameter order / description text only; gate still at lib/index.js:238 |
present |
dsh-tool-fs |
descriptions only; lib/index.js:1123 still validates first |
present |
dsh-tool-pwsh |
parameter order / text | present |
dsh-tools (run_code) |
ptc.js:297, index.js:1187 unchanged |
present |
All four escalation paths are behaviourally unchanged in the alpha; only parameter ordering and prompt text moved. Tag note: 0.2.1-alpha.1 is on the npm alpha tag (latest/next are 0.2.0-rc.2).
4. Extra: the schema-level asymmetry (why the model calls it this way)
fs(dsh-tool-fs/lib/index.js:1101,:1105):sandboxPermissionsDescription("operation")+ "Required with sandbox_permissions: one sentence for the user explaining why this exact file operation needs the wider access. …"bash(:517,:521): the same helper with"command", the same sentence verbatim;pwsh(:487,:491) likewise.run_code(dsh-tools/lib/types/ptc.js:61-62): different wording, but also says the reason is required.
The asymmetry is an absence: the three shell/fs descriptions are word-for-word identical modulo the noun and all declare the reason "Required" — which is false for bash — and none of the four mentions the same-mode no-op rule.
5. Semantics (useful for scoping the fix)
- Rejection requires both same-mode and a blank reason. Same-mode with a non-empty reason is accepted on write/edit (7 + 22 calls, 0 related errors) ⇒
tool-fs/src/sandbox.ts:88is the only obstacle. - Exactly one of
sandbox_permissions/justificationsupplied → reject (correct today, do not change). - A strictly wider mode with a blank reason → keep rejecting (correct today;
tool-bash/tests/tools.spec.ts:741covers it). - The fs file's own doc comment (
dsh-tool-fs/lib/index.js:1112, "Repeating the standing mode requires no approval") is contradicted by its own:1123; the same goes for the comment claiming both enforcing families validate identically. Both are stale.
6. Bounds / what I could not confirm
- No live re-run. Approval prompts are disabled in the environment where I gathered this and escalation arguments are refused, so §1/§2 rest on static code plus stored transcripts.
editbehaviour is inference (0 observations).- The report's "verified on
master": all five cited upstream line numbers are exact, but that is "the lines match the source", not "I confirmed the run happened". app.asaris a single 121 MB archive file, not a directory (the pathapp.asar/dsh/does not exist as a directory); it is byte-searchable and contains each defect line exactly once. Itsruntime.jsonreportsdesktopVersion: 0.2.0-rc.2.- Which commit
0.2.1-alpha.1was built from is not established (I compared the published alpha packages against the expressions cited inmaster).
Re-verification commands
grep -n 'invalid justification' <install>/node_modules/@deepseek-ai/dsh-sandbox/lib/index.js
grep -n 'effectiveMode' <install>/node_modules/@deepseek-ai/dsh-tool-bash/lib/index.js
grep -n 'validateEscalationArgs' <install>/node_modules/@deepseek-ai/dsh-tool-fs/lib/index.js
npm pack @deepseek-ai/dsh-sandbox@0.2.1-alpha.1 && diff <(tar xzOf *.tgz package/lib/index.js) <install>/…/dsh-sandbox/lib/index.js
你的更正我接受——"空白 justification"这句是我说错了1. 撤回我上一轮写"空 justification 的拒绝比有理由的拒绝更糟",隐含"错误信息是空白的"。你核到的恰恰相反: // 抛错点(0.2.0-rc.2 bundle)
// dsh-sandbox/lib/index.js:54
// 上游 master:packages/sandbox/sandbox/src/escalation.ts:58-60
if (justification !== void 0 && justification.trim().length === 0)
throw new Error("invalid justification: expected a non-empty sentence")⇒ 错误信息是固定字面量、且它说清了缺什么("expected a non-empty sentence");空白的是入参 2. ⇒ 修正后,这条真正的问题是两个(都不在文案上)
3. 建议的诉求(改后版本,三条)
第 1 条仍是核心,但理由换了:不是"文案不好",而是"这个要求本身不必要"。 4. 你补的三项我逐条收下
5. 请补一样三次调用的入参原文( 6. 版本你写 一条边界我确认的是你的更正方向成立(错误文本是固定字面量、空白的是入参)这一结论与给定代码一致。 |
回复:你要的三次调用入参(其实只有两条)+ 版本确认1. 入参原文:
|
Uh oh!
There was an error while loading. Please reload this page.
Summary
In a sandboxed composition, a call that repeats the session's current mode (
sandbox_permissionsequal to theeffective mode) with an empty
justificationproduces opposite outcomes in two tool families:bash{"sandbox_permissions":"danger-full-access","justification":""}write{"sandbox_permissions":"danger-full-access","justification":""}Error: invalid justification: expected a non-empty sentenceThe escalation-argument check is ordered differently in the two packages.
bash'svalidateBashArgs()first tests whether the request merely repeats the current mode — if so it skips validation entirely —
and it also normalizes an empty justification to
undefined. The filesystem family'sFsSandboxController.resolvePolicy()callsvalidateEscalationArgs()unconditionally first and onlyafterwards resolves the standing policy.
""is notundefined, so it is treated as "a reason was supplied butit is blank" and rejected before anything else can happen.
This contradicts three existing pieces of documentation:
approveEscalation()inpackages/sandbox/sandbox/src/escalation.ts— "Repeating the call's effective mode returns it without approval." Its implementation isif (mode === effectiveMode) return effectiveMode, i.e. no approval prompt is ever raised for this case.resolvePolicy()'s own doc comment inpackages/fs/tool-fs/src/sandbox.ts— "Repeating the standing mode requires no approval."In other words, the "repeat current mode" case was explicitly exempted on the
bashside (commit8cf9c0ed,see below), while the filesystem tools are stopped at the validation step and never reach the
"no approval needed" branch.
Symptom (observable criteria)
All three hold:
M(journal headersandboxMode: "M";danger-full-accessin the observed case).sandbox_permissions: "M"(equal to the effective mode) andjustification: ""(empty or whitespace-only).write/edit→Error: invalid justification: expected a non-empty sentencebash→ normal result, and noapproval/requestevent at all.Both branches of (3) must occur for the divergence; a blank reason being rejected is correct when the requested
mode is genuinely wider (see Notes / non-bugs).
Reproduction
1. Observed session journal (the actual trigger)
In one session, using the identical escalation arguments:
Why this proves the mode was identical:
bashcan only have passed because it hitargs.sandbox_permissions === effectiveModeand returned early — without that early return it would have thrownthe same
invalid justificationerror. HenceeffectiveMode === "danger-full-access", so the rejectedwritewas a no-op repeat of the current mode, which by rule needs no justification at all.
2. Control within the same session
Every
bashcall withsandbox_permissions+ an emptyjustificationwas counted and checked:writecalls with a non-empty reason: 4/4 succeeded;editwith a non-empty reason: 1/1 succeeded.So the failing input is exactly
justification: "", and only on the filesystem family.3. Live check on the same build
The second line shows the filesystem family also lacks the blank-string normalization that
bashperforms(
justification?.trim() === '' → undefined): the empty string is treated as a supplied-but-blank reason ratherthan as "no reason given".
4. Upstream
On
master,packages/fs/tool-fs/src/sandbox.ts:88still performs the unconditional check whilepackages/shell/tool-bash/src/index.ts:83keeps the exemption — see Root cause.Root cause
① The shared validator treats a blank string as a supplied reason
packages/sandbox/sandbox/src/escalation.ts:51The function behaves as documented. The problem is that callers do not remove the no-op repeat-current-mode
case before invoking it.
②
bashhas the exemption and the normalizationpackages/shell/tool-bash/src/index.ts:83This came from commit
8cf9c0ed(2026-09-20, "fix(bash): allow blank justification without escalation" /"permit empty reasons when repeating the current mode"). Its file list is limited to
packages/shell/tool-bash/**(source, three-language README,tests/tools.spec.ts, and the snapshotsnapshots/session/bash-same-mode-empty-justification).③ The filesystem family validates before resolving the standing policy
packages/fs/tool-fs/src/sandbox.ts:87Both
write(packages/fs/tool-fs/src/write.ts) andedit(edit.ts) go through this single path(
resolvePolicy("write", …)/resolvePolicy("edit", …)). Even with a non-empty reason there is noargs.sandbox_permissions === standingPolicy.modeshort-circuit, so the call continues intoapproveEscalation(), whereif (mode === effectiveMode) return effectiveModereturns immediately withoutprompting. In other words, line 88 is the only thing standing in the way.
④ Same gap in
pwshand the PTCrun_codepackages/shell/tool-pwsh/src/index.ts:118validatePwshArgs()does not even receiveeffectiveMode, so it structurally cannot apply the same-modeexemption (Windows hosts are affected too). The comment claiming both enforcing families validate identically is stale.
packages/core/tools/src/ptc.ts:385:run_codehas the same validate-then-resolve ordering.⑤ Why tests did not catch it
packages/shell/tool-bash/tests/tools.spec.ts:783-800undefined/''/' \t\n'justifications for a repeated same-mode request, assertingisError === falseand no approval prompt — this is the intended specsnapshots/session/bash-same-mode-empty-justification/bashsame-mode blank reasonpackages/fs/tool-fs/tests/tools.spec.ts:967'use the current permissions'), so the blank branch is never exercisedThe spec exists only on the
bashside; the fs same-mode test happens to sidestep the blank reason.Scope: four tools are affected
write,edit(packages/fs/tool-fs/src/sandbox.ts:88),pwsh(packages/shell/tool-pwsh/src/index.ts:118),and PTC
run_code(packages/core/tools/src/ptc.ts:385). Onlybashcarries both the same-mode exemption andthe blank-string normalization.
Impact
case the rejected call was the first draft of a ~23 KB generated script; the retry wrote a different script
(~22.9 KB, different docstring and structure). The draft was never written to disk and exists only in the session
journal, so the artifact the user ends up with is not the one the model intended to write.
current mode, which needs no reason". The model learns "never leave the reason empty" instead of
"repeating the current mode needs no reason".
default to an empty
justification(the triggering model did: 42 calls carriedsandbox_permissions, 21 of themwith a blank reason — 19 on
bash, all accepted; 2 onwrite, both rejected), this is not sporadic: the firstwrite of every file fails.
Suggested fix
P0 — port the
bashexemption into the filesystem family (minimal)packages/fs/tool-fs/src/sandbox.ts:This covers
writeandeditat once (singleresolvePolicy).P1 — close the rest of the family
validatePwshArgs(args, effectiveMode): add the parameter and mirrorbash's two rules; update pwsh's"both enforcing families validate identically" comment.
packages/core/tools/src/ptc.ts(run_code): resolve the standing policy first, then judge.P2 — pin the semantics with tests
packages/fs/tool-fs/tests/tools.spec.ts: combineit.each([undefined, '', ' \t\n'])withsandbox_permissions: mode, assertingisError === falseand that noapproval/requestfires (i.e. porttool-bash/tests/tools.spec.ts:783to fs).snapshots/session/fs-same-mode-empty-justification/, symmetric to the existing bash snapshot.(
tool-bash/tests/tools.spec.ts:741already covers this correct behavior).P3 — optional: make the message self-explanatory
Once the same-mode decision is moved ahead of validation,
invalid justificationcan only arise for a genuinewidening. Consider appending the way out, e.g.
invalid justification: expected a non-empty sentence (repeating the current mode needs no justification).Expected behavior
sandbox_permissionsabsent + blankjustificationbashtoday; fs should match)sandbox_permissions == current mode, reasonundefined/''/' \t\n'bashtoday; fs should match)sandbox_permissions == current mode, non-empty reasonsandbox_permissionsstrictly wider + blank reasoninvalid justification: expected a non-empty sentence(correct today — do not change)sandbox_permissions/justificationpresentNotes / non-bugs
bash'stests/tools.spec.ts:741tests exactly that. This report only argues that the no-op repeat of thecurrent mode must not be stopped.
operation; a retry with a reason succeeds.
bashalready exhibits (repeating thecurrent mode needs no reason). The fix target is tool-family consistency, not teaching models which tool
secretly requires a reason.
Evidence appendix
Inspecting the shipped bundle (no source checkout needed)
app.asaris plain concatenation plus a JSON header, so it can be searched by bytes:The error text
invalid justificationis likewise greppable inside the bundle.Decoding the multi-frame zstd session journal
Session journals are multi-frame zstd JSONL. Node's
zlib.zstdDecompressSyncdecodes only the first frame(223 bytes, measured), so use Python's
zstandard:Facts and how each was verified
bashran,writewas rejectedtool/call+tool/resultrecords in the session journaldanger-full-accessbashonly skips validation when the requested mode equals the effective mode; its success implies equality (plus the journal header'ssandboxMode)write2/2 rejected,bash19/19 acceptedsandbox_permissions+justificationwrite+justification: ""→justification is only valid together with sandbox_permissionsmaster:packages/fs/tool-fs/src/sandbox.ts:88unconditional;packages/shell/tool-bash/src/index.ts:83exemptbash8cf9c0ed(2026-09-20), file list limited topackages/shell/tool-bash/**bashsidetool-bash/tests/tools.spec.ts:783(blank reason + same mode must pass) vstool-fs/tests/tools.spec.ts:967(same mode tested only with a non-empty reason)All reactions