Replies: 6 comments 1 reply
|
Confirmed at HEAD, and the mechanism is exactly the one your measurements point at. Three lines do it:
So the two halves lock each other: the only thing that could move the watermark is the callback that a 413 prevents from running. Hence your observation that it is permanent, Immediate unblock (no code change): mount If you want to keep the field: the registry merge happens in
npm install @argszero/cordis-plugin-session-log-budget# cordis.patch.yml of the profile
- insert:
- id: session-log-budget
name: '@argszero/cordis-plugin-session-log-budget'
config:
mode: enforce # or: report, which measures and logs and changes nothing
maxFieldBytes: 8000000What it does: it wraps the registry's Boundaries, stated rather than left to be discovered:
The real fix belongs in core, and it is small: the owner should re-anchor when replay is undeliverable (write a current-generation acceptance record covering what is known to be held, e.g. from the last pre-migration watermark, instead of only ever moving forward after a 2xx). The plugin is a seam fix that works without that change — not a substitute for it. Tested against |
|
Same field, same lock — but a second way to reach We hit this independently on 0.1.7-alpha.2 and ended up at the same three lines. One thing your report and @argszero's confirmation leave open is which way
Measured on the session we recovered (all 122 frames decoded, 206,405,771 B, 40,833 events): The session lived twelve days and 40,717 events with no accepted watermark at all, because the add-on was not enabled for the first part of that life: That changes where the fix has to go. In our case nothing was invalidated — there was nothing to invalidate. So "do not skip other-generation watermarks" does not cover it, while @argszero's What we measured, in case it is useful alongside your numbers: the rejected body was 206,429,335 B, of which Two things we took from your report specifically:
We also confirmed your workaround on our session. With the field dropped at the adapter, the session's first ever Our byte-level breakdown, and the local gate + degradation-ladder implementation with three rounds of self-review, are in #7699. |
|
Independent data point on the open question @txxt114514 raised — which path reaches
Workaround verified end-to-end (home patch, applies to every profile): - id: session-log-deepseek
config:
enabled: falseAfter a restart no new Generalised guardrails for this class (bounded first upload / auxiliary-field isolation / upgrade-path checks): #7737. |
|
补充一下我们把同一根因的新证据整理成了一篇补充帖:👉 #7753 那里补了 4 条本贴尚未提到的事实:
另外那篇帖子里附了一份 非程序员也能照做的 3 步修复步骤(含一键 PowerShell 脚本与回滚方式),方便普通用户先自救。 核心修复仍然期待官方:迁移/恢复时重新锚定水位 + 给该字段加字节/token 上限 + 让 413 直接可见(与 @argszero 的结论一致)。 |
|
Independent data point on the "watermark was never established" path (not the migration path), on Symptom matched the report exactly: after upgrading, the first message in an old long session and it repeated on every subsequent turn with no Local evidence (all frames of the multi-frame zstd log decoded,
So Two things I could add to the thread:
Confirming the rc.2 fix path: Thanks for the source-level confirmations — they made this trivial to verify locally. |
|
The watermark is not gone; it is filtered — and that changes what a re-anchor has to carry. Versions: Where the invalidation actually happensAfter a V3→V4 migration the old So the invalidation is not a deletion. It is one equality on the read side: The records carry their generation, the reader compares it to the header's, and every pre-migration watermark leaves the accepted set. "The watermark is unusable" and "the watermark is absent" are different states — and this is the first of them, which is why it is the one a re-anchor could act on. What a re-anchor needs that does not exist yetTaking the last accepted sequence from a pre-migration marker looks like reading a number out of the log. The migration does something adjacent that shows why it is not. That stage reassigns sequence numbers: And it knows records may hold sequence references.
That is the part a re-anchor has to carry. Adopting A form that might recurFor a migration, every field that survives across generations and contains a sequence number is in one of two states: it is in the remap table, or it is not. Records of the first kind are safe because a table entry was written for them. Records of the second kind are quiet while nothing reads them, and misplaced the moment something does — and the change that starts reading them is exactly the change that would be blamed for being wrong. A cheap check, if it is usefulThe mapping is built in the open — one Version surfaceOn the published Authorship note: code locations were read by me against the published tarballs of the versions named above; drafting assisted by AI; verification and publication by me. I have no engineering background — if any technical claim reads wrong, please call it out; I will re-verify against the toolchain and correct. Reported by the OfferKuai Team — Founder: Zhaofeng (Yaming). Website: https://www.offerkuai.com/ | Contact: contact@offerkuai.com 中文版接受标记不是没了,它是被过滤掉了 —— 而这一点改变了"重新锚定"必须携带什么。 版本: "作废"实际发生在哪V3→V4 迁移之后,旧的 所以这次作废不是删除,而是读侧的一个等式: 记录带着自己的代际,读者把它和表头的代际比一下,于是一切迁移前的接受标记离开那个已接受集合。"接受标记不可用"与"接受标记不存在"是两种状态 —— 而这里是第一种,这也正是"重新锚定"能作用的那一种。 "重新锚定"需要一个现在还不存在的东西从迁移前的标记里取出"最后一个被接受的序号",看上去只是从日志里读一个数字。而迁移做的一件相邻的事,正好说明为什么不是。 那个阶段会重新分配序号: 它也知道记录里可能带有序号引用:
这正是"重新锚定"必须携带的部分:把迁移前标记里的 一种可能会重复出现的形态对一次迁移而言,每个"跨代际存活且含序号"的字段只有两种状态:进了变换表,或没进。前一类是安全的,因为有人为它写过表项;后一类在无人读取时安静,在有人读取的那一刻错位 —— 而那个开始读取它的改动,恰好就是出问题后会被归咎的那个改动。 一个便宜的核对(若有帮助)那张映射表是当众建起来的 —— 每个源事件一次 版本面在已发布的 声明:上列代码位点由我本人对照上述版本的已发布产物重新通读;文稿撰写由 AI 辅助;核验与发布由我本人负责。我没有工程背景 —— 若任何技术表述有误,请直接指出,我会对照工具链重新核实并更正。我们是插件作者、不是维护者,本文不陈述任何官方政策。 本报告由 OfferKuai(Offer快)团队提交 —— 创始人:Zhaofeng(Yaming)。官网:https://www.offerkuai.com/ | 联系:contact@offerkuai.com |
Uh oh!
There was an error while loading. Please reload this page.
Summary
会话格式迁移(V3 → V4)后,
session-log-deepseek不再认可迁移前写入的交付水位,导致每次请求都附带整份会话日志(dsh_session_log)。日志较大的会话因此超出 endpoint 的请求体上限并收到 HTTP 413;而失败又导致无法写入新的delivery-accepted,于是该会话永久卡死,重试无效。English summary: after a session-format migration the
dsh_session_logwatermarks written under the previous format generation are all skipped, soafterSeqfalls back to-1and every request re-sends the entire session log. For long sessions the body exceeds the provider's request-body limit → HTTP 413, and because no new acceptance can be written the session is permanently stuck.Reproduction
SESSION_FORMAT_VERSION = 4的构建后打开该会话(日志按设计迁移为 V4,原代际文件保留)。稳定复现:
DeepSeek Messages request failed (413),code: INVALID_REQUEST,status: 413。与消息内容、大小无关("ping" 同样失败)。Current behavior
每个请求都携带完整日志。用本地转发代理抓到的真实请求体(dsh 0.1.7-alpha.2;会话 29,290 事件,解压后日志 84.9 MB):
把该附加项关闭后,同一会话、同一条消息:请求体 1.53 MB → HTTP 200,正常作答。
endpoint 体积上限实测(
https://api.deepseek.com/anthropic/v1/messages,以无效 body 探测:超限 → 413,未超限 → 422,两种情况都不消耗 token):8 / 16 / 24 / 32 MB 通过,48 MB → 413。即任何日志接近 ~48 MB 的会话在格式迁移后都会永久不可用;日志较小的会话会先付出一次超大上传,之后恢复增量。上下文本身是健康的,可排除"压缩失效"方向:该请求的
messages只有 1.48 MB,压缩标记在迁移前后一一保留(compaction/start 11、summary 11、end 11、prune 15),contextPressure.surfaceTokens279,257 /contextWindow1,000,000。Expected behavior
迁移后应能继续在该会话中对话:要么水位被重新锚定(不重发整份日志),要么该附加项按字节/token 上限截断,使请求体不可能超过 provider 限制。至少不应表现为无法重试的永久失败。
Environment
deepseek-official(默认 baseURLhttps://api.deepseek.com/anthropic),modeldeepseek-flashRoot cause
packages/session/session-log-deepseek/src/index.tsacceptedThrough()会跳过session-log-deepseek/delivery-accepted中data.sessionFormatVersion !== session.header.version的事件(约第 136 行)。prepare()取afterSeq = acceptedThrough(session),并以session.snapshotEvents(afterSeq + 1)组成后缀;afterSeq === -1即"从头开始发"。格式迁移后
session.header.version为 4,而所有历史确认事件都记录 3 → 全部被跳过 →afterSeq回落-1→ 附带整份日志。新水位只在accept()(请求成功后的回调)写入,而请求先 413,永不写入,形成自锁。按代际限定水位本身是刻意设计("Highest confirmed sequence for this exact Session format generation"),但迁移路径缺少"重发不可交付"时的重新锚定,而请求体上限(provider 侧)与设计假设的"重发总是可交付"相冲突。
Suggested directions
Workaround
在 profile patch 中关闭该附加项即可立即恢复(已验证请求体 96.5 MB → 1.53 MB,同会话正常作答):
代价是模型失去按需读取原始会话日志的能力;新会话(日志较小)即使保持开启也只在首次请求付出一次大上传。
All reactions