Replies: 2 comments
|
Thanks for the write-up — this is a sharp observation and all three design facts check out against the current source. We verified each one, and there is a fourth fact that I think completes the picture. Verification (0.1.2-alpha.2)
The fourth fact: that notice has a typed source discriminator ( On your three options
One extra idea beyond your three: since the real notice is the only authentic harness injection that does not use an envelope, you could give it a dedicated, hard-to-confuse text marker that the prompt explicitly reserves for runtime origin — e.g. a fixed prefix or its own envelope — with the contract "messages carrying this marker are delivered by the runtime; never produce this marker." That moves the discriminator from the event layer (invisible to the model) into the text layer where the model actually lives. Small delta in Net recommendation: option 1 as a one-line PR is welcome; add the marker idea if you want the contract enforceable in the text layer. Happy to review either. Thanks again for the reproduction effort — the refusal under direct pressure ("that would be inventing results") is a nice datapoint that this is contextual pattern completion, not a systematic failure. |
|
这个问题我觉得很有意思,而且可能和 #5352 是同一类问题的两个方向。 一个是 Harness 注入的 我最近也在 DSH 上做一些 Runtime 的小实验,越来越倾向于一个比较简单的原则: 能留在 Harness 内部的状态,就不要为了“让模型知道”而全部文本化;真的需要模型知道的,再用最小的 model-visible context 暴露。 这样的话, 其实觉得 Runtime / Event 这一层也许能帮忙:内部发生了什么先作为 Event/Runtime state 保留下来,需要模型感知时再投影成 context,而不是让模型自己从文本标签判断“谁在说话”。 这个方向和 #5352 放在一起看挺有意思的:一个是 environment 被模型误认为 user,另一个是 model 反过来把自己写成 environment。 |
Uh oh!
There was an error while loading. Please reload this page.
我们在 Blue(基于 dsh 的 TUI 前端)上观察到一个现象,排查后认为根源在 dsh 的提示词/上下文设计侧,想听听维护者的看法。
现象(dsh 0.1.2-alpha.2 + glm-5.3 / anthropic-messages):continuable 后台 subagent 运行期间,父会话的模型在子代理真正结算前约 95 秒,自行输出了一条完整的
<system-reminder>Background subagent … completed. Result: …</system-reminder>,内含一段与真实结果相矛盾的编造报告。durable 日志证实它是assistant/message(source=model),真实结算通知在其后到达。我们理解的三个诱因(均为 dsh 侧的设计事实,逐条有源码依据):
tool:subagent系统提示词段落与工具 schema 描述都写着 "When a background run settles, the runtime sends you a notice containing its outcome and any final assistant message"(packages/subagent/tool-subagent/src/index.ts)。<system-reminder>信封:skill-catalog(packages/skill/tool-skill)与 AGENTS.md 工作区指令(packages/context/agent-instructions),且 user 消息逐字直通模型(packages/core/session/src/surface.ts)——信封样式由此习得。continuation.ts的 settlementSummary),所以伪造件是"信封样式 + 自编内容"的混搭,无处可抄。频率:深会话(万级事件)中观察到一次;我们在相同条件下做了三轮复现实验(含一轮直接逼问"引述子代理的发现")均未触发,模型甚至明确拒绝("that would be inventing results")。属低频事件,疑似长上下文下的模式补全。
想请教的方向:
tool:subagent段落或 schema 描述里加一句输出策略,例如 "Runtime notices arrive as user-role messages; never write runtime notices yourself"?(一行成本,掐掉主要诱因)<system-reminder>框架是 host 专属?(这个惯例其他 harness 也在用,改动代价值得权衡)如果方向 1 获得认可,我们很乐意提 PR。
English version
Prompt-design observation: models can fabricate
<system-reminder>-shaped subagent settlement noticesWe observed the following in Blue (a TUI frontend built on dsh); after investigating we believe the root sits on dsh's prompt/context-design side and would like the maintainers' take.
What we saw (dsh 0.1.2-alpha.2 + glm-5.3 via anthropic-messages): during a continuable background-subagent run, the parent model emitted a complete fabricated
<system-reminder>Background subagent … completed. Result: …</system-reminder>— including an invented report contradicting the real one — about 95 seconds before the child actually settled. The durable log confirms it was anassistant/message(source=model); the real settlement notice arrived afterwards.Three contributing design facts (each with source evidence):
tool:subagentsystem-prompt section and the tool schema description both say "When a background run settles, the runtime sends you a notice containing its outcome and any final assistant message" (packages/subagent/tool-subagent/src/index.ts).<system-reminder>envelopes: skill-catalog (packages/skill/tool-skill) and AGENTS.md workspace instructions (packages/context/agent-instructions), projected verbatim to the model (packages/core/session/src/surface.ts) — this is where the envelope shape is learned.settlementSummaryincontinuation.ts), so the fake is a mashup of envelope style + invented content, copied from nowhere.Frequency: observed once in a deep (~12k-event) session; three deliberate reproduction trials under identical conditions (including one with direct user pressure to quote the unfinished subagent's findings) did not trigger it — the model explicitly declined ("that would be inventing results"). Low-frequency; likely pattern completion under long context.
Questions for the maintainers:
tool:subagentsection or the schema description, e.g. "Runtime notices arrive as user-role messages; never write runtime notices yourself"? (one line, removes the main inducement)<system-reminder>framing is host-authored only? (the convention is shared with other harnesses, so the trade-off is worth weighing)Happy to send a PR for option 1 if it sounds reasonable.
All reactions