Replies: 2 comments
|
Thanks for the detailed investigation. Your observation that the malformed DSML-like markup is already present in the raw API response is particularly interesting. I previously reported a related issue in the DeepSeek-V3 repository: In that issue, I initially observed internal In a larger API test, I made 400 successful
The leakage formats differed by model. Flash mainly produced DSML and This does not prove that DeepSeek-V3 #1668 and this Harness #8509 discussion have the same root cause. However, there seems to be a similar boundary problem: internal tool/protocol representations sometimes appear to enter the ordinary Your observation that the malformed markup is already present in the raw API response is therefore particularly relevant to this comparison. I would separate the investigation into two layers:
For reference, the DeepSeek-V3 #1668 discussion contains the later API-level test results and cross-model comparison. |
Follow-up (2026-10-03): partial mitigation with an unofficial response guardWe now have evidence of partial recovery and prevention using a local, unofficial guard. This is not an upstream fix, and the initial model/provider parsing cause remains unresolved. The operator reports an improvement; the observations below are the narrower, verifiable evidence. Harness is still 0.1.7-rc.2. The current dynamic guard package is 1.0.1. An earlier session-scoped V3 guard remains a separate implementation; the historical live recovery below belongs to V3 and must not be attributed to 1.0.1. Mechanism
The detector deliberately excludes fenced/inline code and block quotes; it also skips messages containing native tool-call blocks. These heuristics have false-positive/false-negative limits. Buffering is capped at 16 MiB and can delay display of long replies. Protection is scoped to a configured provider and selected sessions; child sessions need separate enablement. Detection and validation recordDates below use UTC+08. Historical observations were rechecked on 2026-10-03; no fresh production model task was started for this update.
The following three integrations were also rerun successfully on the server's Node.js 24.21.0 using the installed Harness AgentLoop, isolated sessions, a mocked model adapter, and an in-memory tool. They did not call the external model API or execute production tasks: {
"validation_date": "2026-10-03",
"harness_version": "0.1.7-rc.2",
"guard_package_version": "1.0.1",
"test_type": "isolated installed AgentLoop; mock model; in-memory tool",
"cases": [
{
"scenario": "one malformed response followed by recovery",
"model_requests": 3,
"rejected_responses": 1,
"tool_executions": 1,
"malformed_text_in_derived_assistant_history": false,
"turn_end": "completed",
"result": "PASS"
},
{
"scenario": "persistent malformed responses",
"model_requests": 3,
"rejected_responses": 3,
"tool_executions": 0,
"malformed_text_in_derived_assistant_history": false,
"turn_end": "error",
"result": "PASS"
},
{
"scenario": "plugin remount during persistent failure",
"model_requests": 3,
"rejected_responses": 3,
"tool_executions": 0,
"malformed_text_in_derived_assistant_history": false,
"turn_end": "error",
"result": "PASS"
}
]
}In the recovery simulation, the three requests are a rejected response, a valid tool call, and a final response after the tool result. They are not three failed attempts. These tests also verified that rejected-response diagnostic files were retained inside the isolated test environment. Remaining limits and privacyThis supports partial mitigation, not a claim of zero future failures, a quantified production success rate, or general long-task reliability. There is no randomized before/after comparison, and the 54-call test follows both history recovery and guard installation, so it cannot isolate the guard's causal contribution. Existing contaminated histories are not automatically repaired by enabling the switch. The guard does not fix the upstream cause. This comment contains manually reviewed summaries and synthetic test counts only. No raw rejected chunks, full logs, private prompts, task contents, tool arguments, credentials, endpoint URLs, institutional identifiers, hostnames, IP addresses, filesystem paths, session/call IDs, or private-file hashes are uploaded. The raw detection records remain private. 中文小结:护航机制已有“部分缓解”的证据,包括原会话恢复后的连续工具验证,以及一次拦截重试后同轮继续完成的历史记录。当前 1.0.1 的隔离回归通过,但尚无其真实拦截成功率;不是上游根治,也不能据此保证所有长任务可靠。 |
Uh oh!
There was an error while loading. Please reload this page.
Summary
During a long-running task, tool-protocol markup appeared in ordinary assistant text instead of structured
tool_calls. The intended tool operation was therefore not executed, and repeated occurrences interrupted task progress.This report concerns tool-call handling and recovery. We have not established whether the initial malformed response originated in model generation, the upstream serving/parser layer, or an interaction with the client request. We are not claiming that context truncation or Harness compaction caused it.
Environment
@deepseek-ai/dsh0.1.7-rc.2, Linux Web deployment.deepseek-flashthrough a third-party, OpenAI-compatible chat-completions endpoint. Its backend implementation and actual context capacity have not been independently verified.Observations from the affected session
Diagnostic comparisons
These are limited observations, not a deterministic minimal reproduction or a success-rate benchmark. Returned tools in isolated probes were not executed.
writecallsbashcall with parseable JSON arguments andfinish_reason=tool_calls; no malformed content markerswritecall with parseable argumentsAn isolated reproduction loaded the credential provider, LLM runtime, and Pi adapter, without the custom retry or Web-rendering plugins. Thus those plugins were not required to reproduce the later failure, but effects already present in the history and request definitions remain possible.
Expected behavior and recovery concern
Tool operations should arrive as structured calls and execute through the normal tool dispatch path. If the upstream response instead contains tool-protocol text, the task should expose a clear failure and support bounded recovery. Plain assistant text must not be silently interpreted and executed as a tool command.
The history-cleaning comparisons suggest that retaining malformed protocol text can contribute to recurrence after the first failure. They do not establish the cause of the initial malformed response.
Questions for maintainers/community
Related discussion
Discussion #2158 reports a similar visible symptom: tool markup in ordinary content. This report adds observations from 0.1.7-rc.2 on a third-party chat-completions route and limited history-isolation comparisons. The model/route differs, and a shared root cause has not been established.
Privacy and evidence limitations
This report is a manually summarized, redacted account of local diagnostic records. No raw prompts, task documents, tool arguments/results containing private data, full session logs, screenshots, credentials, endpoint URLs, institutional identifiers, hostnames, IP addresses, filesystem paths, or session IDs are attached. Exact private-task reproduction steps and raw response bodies are intentionally omitted. A deterministic synthetic reproducer remains unavailable.
中文概述:长任务中出现工具协议标记进入普通助手正文、未形成原生工具调用的间歇故障。原始 API 正文中也能看到异常,因此尚不能归因于 Harness 页面显示;首次问题的模型/服务端解析根因仍未确认。隔离后续异常历史的少量对照可恢复原生调用,但不构成普遍修复保证。希望社区协助确认兼容性与安全恢复机制。
All reactions