-
Notifications
You must be signed in to change notification settings - Fork 2
plat 024
github-actions[bot] edited this page Sep 20, 2026
·
1 revision
| Coordination | Value |
|---|---|
| Assigned agent | Claude Code |
| Follow-up reviewer/fixer | Codex |
| Ticket state | implemented |
| Execution order | C — after A, B, and E |
| Last synchronized | 2026-08-04 |
Claim this ticket as
in_progressbefore implementation. Update this fragment during active work; synchronize the shared index at handoff.
- Priority: P2
- Owner: mcpagent/multi-llm-provider tool-error attribution and logging
-
Evidence: 35 of 90 sampled
[TOOL_ERROR] CLI tool payload failuremarkers have an emptytool=field; attribution requires regexing the inner JSON envelope. - Source: what_the_runtime_tells_an_agent_about_itself.md
- Problem: runtime behavior is unchanged, but failure counts and ownership cannot be measured reliably from the canonical marker.
- Implementation boundary: populate the tool name from the structured call identity before formatting the marker, with a narrow envelope fallback only where the adapter genuinely omitted it.
-
Implementation (2026-08-04, mcpagent
d1eca1f): two-layer recovery, in priority order. (1)streamingManagerrecords eachToolCallStart's name against itsToolCallID; aToolCallEndwith an empty name looks itself up by the same ID — authoritative, since it is the same call reporting its own name. (2)toolerr.ToolNameFromResult: a narrow regex matching only the one harness-generated phrase (tool execution failed: layer=\S+ tool=(\S+)) present in every unattributed marker sampled this session, tried only when no start event was recorded. Deliberately this narrow rather than a generaltool=scanner: a broader pattern risks lifting an unrelated nested tool name and misattributing the marker to it. Whenchunk.ToolNameis already non-empty, the transport wrapper's own name always wins over anything mentioned inside its own result. If neither source proves a name, the marker now saysunknownexplicitly rather than going out blank. -
Verification: four table-driven streaming tests (structured recovery,
envelope fallback, explicit unknown, and a dedicated misattribution guard
using the exact nested-failure shape from the existing
TestStreamingManagerPromotesNestedCLIToolFailureToErrorEvent), plus five directtoolerr.ToolNameFromResultcases including two proving the regex does not match ordinary prose or an unrelatedtool=assignment. Verified three of the four streaming cases fail without the fix. -
Regression tests:
TestStreamingManagerRecoversToolNameForUnattributedErrorEventinagent/llm_generation_streaming_test.go;TestToolNameFromResultintoolerr/toolerr_test.go. -
Independent follow-up (Codex, 2026-08-04): the first implementation
recovered the name only inside the failure branch, after calling
CanonicalFailureForToolwith the empty end-chunk name. That bypassed the intentional suppression for reporting tools such asquery_workflow_dband could promote legitimate domain rows into false tool errors. The effective name is now resolved before classification and reused for logging, events, CLI history, and callbacks.TestStreamingManagerRecoversToolNameBeforeFailureClassificationproves an unnamed end chunk forquery_workflow_dbremains a successful, correctly named result when its payload merely describes a failed record. -
Acceptance: table-driven adapter/envelope fixtures always emit a stable
tool name when one exists, explicitly emit
unknownwhen none can be proven, and never misattribute a nested tool error to its transport wrapper. Met by direct fixture test; a real cross-workflow log scan confirming the emptytool=rate has dropped has not been run, so this staysimplementedrather thandone.
Auto-synced from docs/ on main. Edit there, not here.