Support tool_choice for OpenAI-compatible adapters and detect text-format tool calls as a fallback
#3609
Replies: 1 comment
|
Ground-truth measurement supporting your premise, so this feature request doesn't have to argue about what the current behavior is: We fault-injected the real pi-ai accumulator with a completion whose content is literal tool-call text (Anthropic-style Related family datapoint: even when the model DOES use native |
Uh oh!
There was an error while loading. Please reload this page.
Local OpenAI-compatible models (e.g. Qwen3.8-27B-8bit served via MLX DSpark) sometimes emit tool calls as literal
<tool_call>text instead of native structuredtool_calls. The harness currently has no fallback: the text is not parsed,tool_choiceis not supported, and the agent loop treats a text-only reply as a normal completed turn — so the calls silently never execute and the user just sees raw XML in the transcript.Background, proposed changes, and acceptance criteria
Background (reproduced in a real session)
Session 「商业情报国内外现状扫描」 (
session-46125110-12f6-4646-97ec-6d98d7b063c4, providermlxdspark/Qwen3.8-27B-8bit,api: openai-completions):Step 1: the model correctly emitted one native structured call (
ask_user_question) → executed fine.Later steps: the model wrote 6
subagentcalls plus atodo_writecall as literal text blocks:The harness executed zero of them (
tool/callevent count = 1 for the whole session), and the turn ended ascompleted/max-tokenswith the raw text displayed to the user.Current state verified in the codebase:
<tool_call>text parsing anywhere in the repo.GenerateOptionshas notool_choice— documented as an explicit MVP cut inpackages/llm/llm-deepseek/README.mdandpackages/llm/llm/README.md("tool_choice is not mapped — not part of the core vocabulary").response_formatconstrained decoding (GenerateOptionsistemperature/maxTokens/stoponly).completedwhen the assistant message contains notool-callblocks (packages/core/agent-loop/src/agent.ts), andllm-retryonly retries transport-level failures.Proposed changes (can be split)
tool_choicesupport (smallest viable, preferred) — addtool_choicetoGenerateOptionsand map it through the DeepSeek and pi-ai (OpenAI-compatible) adapters. The official DeepSeek API and most OpenAI-compatible local servers (vLLM, llama.cpp, etc.) accept"required"/"auto". When tools are available, the agent loop can sendtool_choice: "required"(configurable), forcing structured tool calls and eliminating text-format calls at the source. This MVP cut now has a real-world trigger.Text tool-call fallback detection (enhancement) — when a step finishes
completedwith text that contains<tool_call>/<function=markers while tools were available, inject a corrective user message ("tool calls must be issued through the native structured protocol, never as text") and re-run the step instead of ending the turn. Executing parsed text calls should be at most an opt-in degraded mode, with a tool-name allowlist, JSON argument validation, and prompt-injection protection.Acceptance criteria
mlxdspark/Qwen3.8-27B-8bit; with change 1 enabled the model returns nativetool_callson every step (verifiable viatool_choice: "required").tool_choiceis not mapped.User- / model-visible change
Test evidence
session.jsonl.zstd(assistant messages containing<tool_call>text blocks;tool/callcount = 1) as reproduction evidence.All reactions