You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
bash tool advertises sandbox_permissions that can never succeed under a danger-full-access session — GPT-6 burns a whole turn on 11 identical failing calls
#6701
Under a session whose standing sandbox policy is danger-full-access (approval policy never), the bash tool's schema still advertises sandbox_permissions / justification — parameters that can never succeed in that state. Models with a strong prior for harnesses where the shell tool takes a sandbox/elevation parameter pattern-match on the advertised schema (not the prose) and burn an entire turn in a deterministic failure loop.
We reproduced this with GPT-6 (third-party model via an openai-completions provider): 11 consecutive identical failing bash calls, interleaving two errors neither of which names the actual fix (removing the field), and 6 more identical calls after the harness's repeat-detection warning. The session's bash channel was effectively dead; the model only recovered by switching to tools whose schemas don't carry the fields.
Environment
dsh 0.1.5-rc.1 (Homebrew), macOS arm64; verified the same logic on current master 0d1f500.
Model: GPT-6 Astra via a third-party openai-completions provider (JoyGate adapter) — any schema-prior-heavy model should behave similarly.
Session: permission/preset: danger-full-access, sandbox/mode: danger-full-access, approval/policy: never (the shipped composition for DSH_PERMISSION_MODE=danger-full-access).
Prompt: a neutral one-liner — "检查一下joygate代码结构" (check the JoyGate code structure). No mention of sandbox/permissions anywhere in the prompt, so the field usage is entirely spontaneous.
Evidence (from the session log)
Timeline of the bash tool calls (17 tool calls total in the turn):
invalid justification: expected a non-empty sentence
2
same, justification still empty
same error
3–5
justification="只读查看目录与 Git 状态。"
sandbox escalation to "danger-full-access" is not strictly wider than this call's current "danger-full-access" mode
6–7
justification="只读检查"
same "not strictly wider" error
8–10
justification="" again
back to "invalid justification"
—
harness injects the repeat-detection warning here ("You are repeating the exact same tool call with identical arguments…")
11
identical call, justification=""
"invalid justification"
Additional facts:
All 11 commands were the same trivial pwd && ls -la && git status --short --branch.
Zero assistant text between any of the calls — no narrative, no self-correction.
Recovery came only from abandoning bash for read / glob (whose schemas don't advertise the fields), after which the task completed normally.
The runtime context and the tool description both already forbid the field in this state ("approval prompts are disabled … do not set sandbox_permissions"; the bash description's escalation paragraph says the same). Prose did not stop this model class; the model only ever respected schema shape.
Root cause
Advertisement is gated on executor capability only, not on reachability.packages/shell/tool-bash/src/index.ts:192:
A confining executor is mounted ⇒ all escalation targets are advertised, even when the session's effective mode is danger-full-access, where WIDER_MODES admits nothing. packages/fs/tool-fs/src/sandbox.ts:45 mirrors the same pattern for the fs tools.
This is a documented trade-off, but it has a blind spot.packages/sandbox/sandbox/src/escalation.ts:24-26:
Checked at EXECUTION, never baked into a tool schema — the schema's enum is ESCALATION_TARGETS, because schemas are registry-global while the effective mode is per-call truth.
The registry-global choice is reasonable, but it means a danger-full-access session advertises two dead parameters, and the per-call check then rejects every use — deterministically, with no retry shape that can ever succeed.
Neither error names the fix. The model sees invalid justification: expected a non-empty sentence (escalation.ts:59) when it omits the sentence, then sandbox escalation to "…" is not strictly wider than this call's current "…" mode (escalation.ts:163) once it fills one in. Oscillating between the two is exactly what GPT-6 did, 11 times. An error that is a terminal dead end should say what to do instead: remove sandbox_permissions and rerun the command as-is.
The repeat-detection warning is generic. It fired mid-loop and was followed by 6 more identical calls. Quoting the last error verbatim (with its remediation) in that warning would give schema-driven models one more targeted chance to break out.
Suggested fixes (increasing scope)
Minimal, execution-time — respects the documented design: make the dead-end errors self-remediating. In approveEscalation, when effectiveMode admits no strictly-wider target (or equals the requested mode), throw e.g.:
sandbox_permissions cannot succeed: this call already runs in "danger-full-access" (nothing wider exists; approval policy: never) — remove sandbox_permissions and rerun the command without it
Optionally surface this before the pairing validation (empty justification first sent the model down the wrong path — fixing the sentence felt like progress, then hit the wall).
Deeper, if the team ever revisits the schema boundary: per-request schema pruning against the session's effective mode would deterministically stop schema-driven loops — the strongest evidence in this repro is that the model immediately behaved once it landed on tools whose schemas lack the fields.
Cheap addition: have the repeat-detection warning include the last error text verbatim.
Option 1 is a small, well-scoped change; I'm happy to prepare a PR for it (noting CONTRIBUTING.md says external PRs are not accepted at the moment, so raising it here first). The full session log (session.v3.jsonl, 17 tool calls) is available on request.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Under a session whose standing sandbox policy is
danger-full-access(approval policynever), thebashtool's schema still advertisessandbox_permissions/justification— parameters that can never succeed in that state. Models with a strong prior for harnesses where the shell tool takes a sandbox/elevation parameter pattern-match on the advertised schema (not the prose) and burn an entire turn in a deterministic failure loop.We reproduced this with GPT-6 (third-party model via an
openai-completionsprovider): 11 consecutive identical failingbashcalls, interleaving two errors neither of which names the actual fix (removing the field), and 6 more identical calls after the harness's repeat-detection warning. The session'sbashchannel was effectively dead; the model only recovered by switching to tools whose schemas don't carry the fields.Environment
0.1.5-rc.1(Homebrew), macOS arm64; verified the same logic on current master0d1f500.GPT-6 Astravia a third-partyopenai-completionsprovider (JoyGate adapter) — any schema-prior-heavy model should behave similarly.permission/preset: danger-full-access,sandbox/mode: danger-full-access,approval/policy: never(the shipped composition forDSH_PERMISSION_MODE=danger-full-access).Evidence (from the session log)
Timeline of the
bashtool calls (17 tool calls total in the turn):sandbox_permissions="danger-full-access",justification=""invalid justification: expected a non-empty sentencejustificationstill emptyjustification="只读查看目录与 Git 状态。"sandbox escalation to "danger-full-access" is not strictly wider than this call's current "danger-full-access" modejustification="只读检查"justification=""againjustification=""Additional facts:
pwd && ls -la && git status --short --branch.bashforread/glob(whose schemas don't advertise the fields), after which the task completed normally.sandbox_permissions"; thebashdescription's escalation paragraph says the same). Prose did not stop this model class; the model only ever respected schema shape.Root cause
Advertisement is gated on executor capability only, not on reachability.
packages/shell/tool-bash/src/index.ts:192:A confining executor is mounted ⇒ all escalation targets are advertised, even when the session's effective mode is
danger-full-access, whereWIDER_MODESadmits nothing.packages/fs/tool-fs/src/sandbox.ts:45mirrors the same pattern for the fs tools.This is a documented trade-off, but it has a blind spot.
packages/sandbox/sandbox/src/escalation.ts:24-26:The registry-global choice is reasonable, but it means a
danger-full-accesssession advertises two dead parameters, and the per-call check then rejects every use — deterministically, with no retry shape that can ever succeed.Neither error names the fix. The model sees
invalid justification: expected a non-empty sentence(escalation.ts:59) when it omits the sentence, thensandbox escalation to "…" is not strictly wider than this call's current "…" mode(escalation.ts:163) once it fills one in. Oscillating between the two is exactly what GPT-6 did, 11 times. An error that is a terminal dead end should say what to do instead: removesandbox_permissionsand rerun the command as-is.The repeat-detection warning is generic. It fired mid-loop and was followed by 6 more identical calls. Quoting the last error verbatim (with its remediation) in that warning would give schema-driven models one more targeted chance to break out.
Suggested fixes (increasing scope)
Minimal, execution-time — respects the documented design: make the dead-end errors self-remediating. In
approveEscalation, wheneffectiveModeadmits no strictly-wider target (or equals the requested mode), throw e.g.:Optionally surface this before the pairing validation (empty
justificationfirst sent the model down the wrong path — fixing the sentence felt like progress, then hit the wall).Deeper, if the team ever revisits the schema boundary: per-request schema pruning against the session's effective mode would deterministically stop schema-driven loops — the strongest evidence in this repro is that the model immediately behaved once it landed on tools whose schemas lack the fields.
Cheap addition: have the repeat-detection warning include the last error text verbatim.
Option 1 is a small, well-scoped change; I'm happy to prepare a PR for it (noting CONTRIBUTING.md says external PRs are not accepted at the moment, so raising it here first). The full session log (
session.v3.jsonl, 17 tool calls) is available on request.All reactions