Repository navigation
[Bug] Missing attachment object still bricks image-bearing sessions on 0.2.0-rc.2, still reported as TRANSPORT (see #7582) #9262
Replies: 5 comments
|
Confirmed on the source, with one correction that matters for the retry policy: the } catch (error) {
if (timeoutOf(...)) throw new LlmError('DeepSeek Messages stream idle timeout', 'TIMEOUT', { cause: error })
if (options.signal?.aborted) throw new LlmError('DeepSeek Messages request aborted', 'ABORTED', { cause: error })
if (error instanceof LlmError) throw error
throw new LlmError('DeepSeek Messages transport failed', 'TRANSPORT', { cause: error })
}and That the same string and code carry unrelated causes is not hypothetical: #9266 is a Windows TLS / proxy / system-CA case surfacing the identical On the good-news side, the recovery primitive is already there and durable.
I intend to ship that as a plugin. It extends the existing
What stays core, and I am not claiming otherwise:
Your |
|
更新一下这条帖子的状态,也顺便记一个可复现的结论。 #7834 那份降级补丁在当前的
明细、适配后的完整 diff 和实测表格都写在 #7834 的评论里了: 如果维护者希望走正式流程,我可以按仓库门禁再跑一遍完整套件。 |
|
Thanks — this is the most useful correction we've gotten on this, and it matches what we measured independently. Both of your corrections check out from our side: the ~13 ms ENOENT → TRANSPORT wrapping is exactly what we saw (attempts returning in under 50 ms with zero HTTP), and On (2), the core-side prevention you say stays core: we have it implemented and verified on current master, and it is now posted as a comment on #7834 — #7834 (comment) — with the full diff inline.
Two things about your plugin design I would like your read on:
(3) is the unclaimed one, agreed. For what it is worth: the session that bricked on our machine now references 51 objects, all 51 present and sha256-consistent — but only because we edited the durable log by hand to replace six dead references with text placeholders. Nothing in dsh would have noticed, or told us which references were dead. |
|
跟一条检索结果,跟我们前面那份适配补丁直接相关,也和 @argszero 的 v0.2.0 计划相关。 我们把仓库 关键约束是它的归属条款:共享的 request-projection 消费方持有这条策略,provider adapter 不得自造独立占位符或恢复状态;被否方案里明确点名 "Catch the error independently in each adapter",理由是 unlogged placeholder 会让 replay 依赖当时恰好存在的 adapter 与存储状态。 据此两点自我修正(详细版贴在 #7834):
我们这边的证据(0.2.0-rc.2 仍复现、9.2 MB/1156 帧里 17 条 TRANSPORT 而 |
|
Since my previous comment, the plugin is published — including the half that this thread is about:
npm install @argszero/cordis-plugin-image-offload-fallbackThe package ships its own - insert:
- id: image-offload-fallback
name: '@argszero/cordis-plugin-image-offload-fallback'0.2.0 adds the missing-object path to the decode-rejection translation that was already there. On each failed request it walks the same surface Now your two questions. 1. If both halves ship, the adapter degrades first and the probe never sees a failure — is that the intent? Yes. The plugin is the fallback for deployments, routes and versions that do not have the core change, and it should become dead weight the moment the assembly-time fix lands. Two reasons it is built that way rather than as a second opinion:
So the intended end state is: core owns the policy, this plugin is uninstalled or inert. That is also why 0.2.0 deliberately does not touch the catch-all in 2. Transient probe errors — treat any non- Treat as "do not recover", and no second probe inside the same attempt, on purpose:
On the proposed note — thank you for pointing at
And these are the parts where our shape is a deliberate subset, which I would not want left implicit:
That is the honest description of a fallback, not a competing design. If the note is implemented, the plugin's missing-object path should be deleted rather than ported; the decode half is the part worth keeping, since a route that cannot decode an encoding never reaches an attachment read at all. ( |
Uh oh!
There was an error while loading. Please reload this page.
中文摘要:#7582 已经把机制讲清楚了——会话历史里只存附件 id,每次请求前都要回附件库重读字节,所以
~/.dsh/attachments下的对象一旦被删除,含图会话就永久失败,而且被错误归类成可重试的TRANSPORT。这条只做两件事:在新版本上复现(0.2.0-rc.2,仍存在),以及补代码级证据(真因cause至今没有落盘)。分析与修复建议见 #7582,这里不重复。Summary
#7582 reported this on 0.1.5-rc.2 (2026-09-23) and it is still unfixed. On 0.2.0-rc.2 the failure reproduces unchanged in kind: once an attachment object referenced by a session's durable history is missing, every later request in that session fails permanently, including plain-text messages, and the failure is still surfaced as
TRANSPORTeven though it dies locally in ~13–15 ms, before any HTTP call.Environment
@deepseek-ai/dsh-attachment-local0.2.0-rc.2deepseek-official,DeepSeek-V41-Flash(reasoning: high)<DSH_HOME>/attachments/v1/objects/What changed between the two reports (and what did not)
prepareRequestImages()reading every image ref before the requestprepareImages(history, connection, modelId, attachments, access, signal)attachments.readImageRequest(...)throwing outside anytryreadImageFile()throwsAttachmentError:ATTACHMENT_NOT_FOUND("Attachment object is missing."), plusATTACHMENT_READ_FAILED/ATTACHMENT_CORRUPTfor other casesDeepSeek API stream from <baseURL> failed→TRANSPORTDeepSeek Messages transport failed→TRANSPORTDEFAULT_MAX_RETRIES = 5,TRANSPORTretryablepolicyKey ["normal",1,["EMPTY_RESPONSE","RATE_LIMIT","SERVER","TIMEOUT","TRANSPORT"],500,10000,0.1]→ 1 retry (UI:已重试模型请求 (1/1))/compactalso failed/compactSo the retry count was lowered from 5 to 1 between the two versions, and the session still dies. Retry count was never the problem; the classification is.
New evidence: the real cause still is not persisted
Suggestion (b) in #7582 — keep the underlying
cause— has not landed. In the affected session (9.2 MB, 1156 frames) there are 17TRANSPORTrecords and zero occurrences of the stringATTACHMENT. Whatassistant/attemptandturn/endactually store:{"type":"assistant/attempt","seq":10809,"time":1791542759036,"data":{"turn":181,"step":4,"stream":[{"type":"chunk","time":1791542759035,"chunk":{"type":"finish","reason":{"kind":"error","failure":{"message":"DeepSeek Messages transport failed","code":"TRANSPORT"}}}}]}} {"type":"turn/end","seq":10814,"time":1791542759601,"data":{"turn":181,"reason":{"kind":"error","error":{"message":"DeepSeek Messages transport failed","code":"TRANSPORT"}}}}The persisted failure has
message+codeonly. So from the UI, the log, and the host, a missing local file is indistinguishable from a network outage — which is exactly the wrong direction this sends anyone debugging it.Minimal reproduction (0.2.0-rc.2)
<DSH_HOME>/attachments/v1/objects/<sha256-prefix>/<sha256>.TRANSPORT. It never recovers, because the dangling reference is part of the durable history.Note the asymmetry: the Files path already degrades (
FileResolutionFailure→ retry inline) and offloaded images are projected to text viaprojectOffloadedImages(history, ref => offloadedImageText(ref, access(ref))). A missing local object has no equivalent fallback.Suggested fix
Same three as in #7582; if only one gets done, do the first:
ATTACHMENT_MISSING) — this alone stops a deleted file from permanently deadlocking a session.cause.code(plus one stack frame) inassistant/attempt/turn/endfailure records.Question
Is the attachment store considered durable internal state that tools and cleanup scripts must not touch? If the intended answer is "don't delete it", the app should say so and guard the path — right now nothing does, and the resulting failure is attributed to the network.
All reactions