Repository navigation
Replies: 2 comments
|
Thank you for a report whose diagnosis survives contact with the source. I checked every leg against a fresh tree: your mechanism is right, and two facts make it sharper — then the seam that lets this be prevented rather than only described. 1. The asymmetry is where you say, and the write path is why nothing landsThe throw is The "the failing event never lands" half is mechanical, and it is worth stating because it is what makes the wedge permanent: 2. A response-stream seam exists, and suppressing a block needs no index renumbering
Dropping a block is index-safe: 3. What the adapter already tells you, before the arguments start
What I am doing with itI am preparing a guard on that seam that suppresses an empty-identity tool-call block before it is advertised — no |
|
The diagnosis above holds, and it has a preventive half that does not need the write path. I packaged it as a plugin and published it:
npm install @argszero/dsh-phantom-tool-call-guard@0.1.0It mounts itself; the package ships a bundle patch, so a deployment using the bundle convention adds one entry: - insert:
- id: phantom-tool-call-guard
name: '@argszero/dsh-phantom-tool-call-guard'Why the response side of
|
| evidence about the block's identity | the assembler will | verdict |
|---|---|---|
block-end.block.id present |
use it verbatim (a closed block is authoritative) | keep |
block-end.block.id empty or absent |
write id: '' — this is the poison |
answer: write call-phantom-<index> |
last delta's id is '' |
write id: '' — the poison |
answer: write call-phantom-<index> |
last delta's id is undefined |
fall back to call-<index> |
keep, id-never-arrived |
deltas stated an id, block-end arrived without one |
lose the id the stream already declared | answer: restore it, id-lost-at-close |
Three deliberate boundaries:
- Only the identity is repaired. An empty tool name is left exactly as the model produced it — inventing a name would dispatch a tool the model never asked for, while a nameless call fails honestly at the registry and its durable
tool/resultcloses the transaction either way. - The written id is namespaced (
call-phantom-<index>) so a reader can tell at a glance that the provider did not write it. A value that looked like the provider's own would be a lie about provenance. - Mode-tunable, and the typo is safe.
answer(default) writes the id;suppressdrops the block instead (the posture the harness already takes for tool calls it cannot execute — max-token truncation);observedecides and counts exactly asanswerwould and emits the stream unchanged;offinstalls no listener. An unrecognisedmodespelling resolves toobserve, because the failure mode of a typo has to be the one that changes nothing.phantomToolCallGuard.snapshot()returns the counters, the bounded history and the plugin's own declared blind spots.
What it does not do, and why your patch is still needed
It is preventive, not curative. A session whose tool/call already landed with callId: "" cannot be repaired from here — the refusal happens while the encoder builds the row, before the append, and reaching a log that already contains the dangling call needs the write path, which is core's. The local patch in your report and this plugin are the two halves, not alternatives: the patch rescues an already-poisoned log, the plugin stops the next one from being produced.
Evidence
- The integration suite builds a real Cordis context, a real
LlmRuntimeand the realllm/streamwaterfall; it feeds the repaired block to the shippedBlockAssemblerand then to the shippedassertV4RowAdmission(session-format-v3-to-v4/src/codec.ts) — the assertion is "the row this boundary refused is the row it now admits", not a paraphrase. The control arm runs the same chunks with the plugin absent and reproduces your refusal,format v4 tool/result at seq N requires toolCallId matching its tool source, verbatim. - The listener position is pinned in both directions: a block contributed by a listener mounted inside the guard is covered; one contributed by a listener mounted outside it is not (and is reported as not covered rather than silently repaired).
- 33 defect-injection arms — one distinction removed per arm — are all caught, none silent; the suite is green on the
0.1.7-rc.2and0.2.0-rc.2lines with every harness package pinned to the line. - Zero runtime dependencies: every bare import is
import type, erased at emit, so the shippedlib/reaches nothing but its own modules. That is deliberate — a plugin's own dependency can materialize a second copy of a harness package in aprofiledeployment, a failure this repository has already shipped more than once.
Happy to adjust the seam or the default if you would rather see the identity minted somewhere else; the gate itself is a pure chunk transformer and it is the only place the decision lives.
Uh oh!
There was an error while loading. Please reload this page.
TL;DR
模型(
mimo-v2.6-flash)在回合尾发出空 id/name 的伪 tool-call(id:""、name:""、arguments:"{}")时,DSH 的持久化校验存在不对称:tool/call事件可以落盘,配对的tool/result事件永远过不了校验,报format v4 tool/result at seq N requires toolCallId matching its tool source。悬空 open transaction 每轮都要补 result → 该会话从这个 seq 起永久卡死,每轮重试报同一个 seq,无法自愈。0.1.7-rc.2;已核对 npm 最新发布0.2.0-rc.2(latest)与0.2.1-alpha.1,均未修复(见下"版本核对")。症状
format v4 tool/result at seq 788 requires toolCallId matching its tool source(seq 取决于中毒点,本例 788、复现第二次为 825)。tool/call上,session/list冷读正常、打开会话后一发消息就炸。根因:持久化校验不对称
中毒事件(真实会话导出,脱敏后):
两处校验的宽严不一致(
dsh-session-persistence-jsonl):于是
tool/call成为永远无法闭合的 open transaction:turn 恢复/续跑时第一件事就是补这个 result → 同一 seq 永久失败。复现
cline-pass/mimo-v2.6-flash,让模型围绕"反引号 / 伪 tool_call"话题写回复(或要求大量 markdown 代码块)。本机 2 小时内复现 2 次;模型自己的推理也预告了触发条件("每次我讲到伪 tool_call、要写出 `` …")。assistant/message(content 含{"type":"tool-call","id":"","name":"","arguments":"{}"})及其tool/call(callId:""),重启后发任意消息即触发。注:这是模型侧已知毛刺(见关联链接),harness 无法要求上游不发,只能保证自己不被一击致死。
版本核对(是否已修)
拉取了最新发布包逐条比对:
dsh-session-persistence-jsonl校验断言集合requires toolCallId matching→ 三个版本的
SessionFormatError消息集合逐条一致,该问题在最新发布版依旧存在。修复建议
方案 A(推荐,根治入口):在
dsh-llm的BlockAssembler.assembled()这一"唯一 keep/drop 决策点"把空 id/name 的 tool-call 块与 max-token 截断走同一个 mask 丢弃,消息侧与 replay 侧同源派生、不会不一致。本地补丁(已验证):效果:空伪调用在组装层被丢弃 →
toolCalls.length === 0→ 按既有分支completed收尾,不再产生悬空 call。冒烟测试 6 项全过(毒 delta / 毒 block-end / 合法调用保留 / max-tokens 回归 / replay 同步过滤 / 全正常路径)。方案 B(持久化层兜底):对空 id 的悬空
tool/call允许落一个TOOL_NOT_STARTED修复 result——0.2.x 已有isExactToolNotStartedRepair机制,可直接复用到该场景,保证 turn 总能闭合。方案 C(对称校验,止血):
tool/call落盘时同样拒绝空callId/name,让失败发生在广告阶段(turn 内可重试),而不是把会话变成永久毒态。三个方案不互斥;A 已在本地端到端验证,可直接合入。
关联 / 上游参考
unknown tool ""(空名工具调用,流分块机制,症状同类、不卡会话)环境
dsh web)cline,modelcline-pass/mimo-v2.6-flash(openai-completions 兼容层)附带发现(供参考,可另开帖)
exec,但 mimo 系模型带强exec先验(MiMo Code code-mode 单 exec 设计),会持续尝试调用并产生unknown tool "exec"噪音。本地已在 dispatch 层做别名(exec → mcp__fastctx__run,历史事件保留原名),需要的话可以贴完整补丁。finish=stop截断(流记录:文本碎片后 finish kind=stop,无 tool_calls)——属上游生成问题,harness 侧无法根治,截断后发"继续"可续跑。All reactions