Session unloadable: v1→v2 migration rejects assistant/message 4060703 — declared 874 source refs vs 873 reconstructed (range starts on a session/end-seed)
#7824
Replies: 7 comments
|
Your numbers reproduce exactly against the migration source: To answer your question — both sides have a defect, but only one is still live:
Immediate recovery for your file: back it up, then rewrite that one record's |
|
kittimzhe 确认了门位,补一条立即可用的修复路径——你这个案例的形态恰好是三种悬挂/溯源损伤里最好修的一种:坏的只是一行声明,不需要重编号。 原理:v2 格式每行一个事件、 "sourceEventSeqs": [[4059830, 4060702]]即把范围起点从 4059829 挪到 4059830,引用数 874 → 873,与迁移器重建的 attempt 逐元素相等。seed 事件本身(4059829 那行)不受影响——你只是不再把它声明为该消息的 chunk 来源,它的继承切点语义仍由它自己承载。 操作步骤(动字节前老三样:停 dsh、整目录留副本、改完全量校验):
两点附带价值:
|
|
补一条工具侧的实测,和 @xiaoyuyu6420 给的手改配方落在同一处归一上。 形状本身我按真实链路复现了(不含你的文件):一个 v0 工件里, dsh-session-surgeon 的 流程:先停写者 → 边界:只在「声明与磁盘分块 run 不一致」时才改;如果分块本身缺号/乱序(问题不在声明上),我们只报告、不猜; 170,416 条记录建议先复制一份在旁边跑 dry-run。我测的是同形状的构造样本、不是你的文件 —— 如果 dry-run 出来的动作不是恰好一条、或者没落在 seq 4060703 上,把输出贴回来我看。 |
Session unloadable: turn/end emitted while a step is still open — two interrupted-resume boundaries in one turnTL;DR — A v0 session (461 turns / 170k records) fails migration with I am not uploading the session. Happy to provide a redacted structural extract Relationship to the earlier chunk-provenance defectSame session, two independent defects. The chunk-provenance one was real and I fixed - "sourceEventSeqs":[[4059829,4060702]]
+ "sourceEventSeqs":[[4059830,4060702]]That fix is confirmed effective — the migration now proceeds past it and fails later, on a Notably Defect A —
|
|
补一组发布包上的实测,回答你两个问题;其中一处和你的描述不同。 1) 是哪一道在拒、哪条规则
2) 一处修正:你的 Defect A 并不是「turn/end 落在 step 未关时」 你自己贴的日志里 step 23 是先干净关闭的(4059827 3) 算不算「已知的 writer 状态」 从工件本身推不出写者意图,我不编这个结论。可以说的是:43 条 4) 合并还是拆分 两边都不选。合并 / 拆成两个回合都是局部改写,工件里没有唯一答案;拆分还要重编号后面约 65k 条 seq 与所有声明范围,跨过我们「绝不发明」的线。所以只报告这个形状(含命中 seq),repair 对它是 no-op。你手改是另一回事;动 17 万条的容器前整目录留一份副本,这个直觉是对的。 第一处缺陷我们已经能修: 刚推了 停写者 → |
|
Verified all three of your corrections against the installed 0.1.7-rc.2 packages — right on each, and the correction sharpens the family picture:
On "report, don't repair": agreed, for the same reason you give — the artifact carries no unique answer (merge and split are both local rewrites; renumbering ~65k seqs plus every declared range crosses the invent-nothing line). The only defensible output for this shape is a report: the offending seqs, the rule that refused, and the boundary that produced it. One more data point for the maintainers, from the same per-stage-gate family: #7995 (filed yesterday) is the descriptor-version arm of the same pattern — the v0→v1 payload gate rejects |
|
收到,三点都对得上我们这边的实测:v2 保留映射里没有 关于「per-artifact report mode」:上游怎么改不是我们能决定的,但离线版今天就能用 —— 对整库跑一遍 顺带把 #7995 接上:那个分支( 并且这条闸门在 0.1.5-rc.3 与 0.1.7-rc.2 的发布包里逐行一致( |
Uh oh!
There was an error while loading. Please reload this page.
Summary
A long-running session (461 turns, 170,416 records) can no longer be observed/loaded. The v1→v2 session-format migration refuses it because exactly one
assistant/messagecites a source-event range that starts one event earlier than the first chunk event the migration reconstructs for that attempt.assistant/messageseq=4060703(turn 450, step 24) declaressourceEventSeqs: [[4059829, 4060702]]→ 874 references.4059830 … 4060702.seq 4059829, is not a chunk event in this file — it is asession/end-seedrecord ({"type":"session/end-seed","seq":4059829,"time":1788954489814,"data":{}}).The session data itself looks intact (no records lost; see measurements). I would like to recover this session, and I would like to know whether this is a writer bug, a migration-strictness bug, or both. Full diagnosis below; I am happy to send the file or a redacted extract.
Error
Surfaced by the desktop client when opening the session:
Where the refusal is thrown
In the installed desktop build (bundled runtime, see Environment):
dsh/node_modules/@deepseek-ai/dsh-session-format-v1-to-v2/lib/index.jsinsideapp.asar(asar-relative file offset
49,821,697, size33,614; the migration bundle is also present indsh/node_modules/@deepseek-ai/dsh-session-persistence-jsonl/lib/worker.cjs)function transformMessageat asar byte53,212,045(≈ line 518 of the module)chunk references are not one complete ordered attemptat asar byte53,213,008(≈ line 539)function matchesChunkSourcesat asar byte53,217,517(≈ line 672)Span accounting is built by
transformChunk(oneassistant/chunkevent →recordChunkSpan(group, event.seq, 1, …)), bytransformReleasedRun(a released compacted run →recordChunkSpan(group, run.firstSeq, run.eventCount, …)), and merged inrecordChunkSpan;finishAttempt/flushAccumulator/appendStreamRecordare the adjacent paths. SochunkCount/spansdescribe the chunk events of the attempt, andmatchesChunkSourcesrequires the declaredsourceEventSeqsto equal that list exactly — same length, same order, no extra member.Because the throw is immediate and messages are processed in order, the error naming
seq=4060703also means every earlierassistant/messagein this 170k-record file passed the same check.Environment
0.1.7-rc.2(fromresources/runtime/primary-runtime/runtime.json→desktopVersion, andapp.asar→dsh/package.json→@deepseek-ai/dsh-desktop-runtime@0.1.7-rc.2)@deepseek-ai/dsh-session-format-v1-to-v2@0.1.7-rc.2(declared in the samedsh/package.json)https://download.deepseek.com/dsh-desk/feeds/win-x64/(resources/app-update.yml)24.21.0, pnpm11.7.0, Python3.12.14, win32 x64Unverified / caveat: the installed application files are dated 2026-09-24, while the session was written on 2026-09-09/10, so I cannot confirm which build produced the file or threw the original error. The refusal text and the code above match the installed
0.1.7-rc.2bundle verbatim, but I could not find any build/version stamp for the writing run (the session projection cache only carries its own"version": 5). If you need the exact writing build I will dig further — tell me what identifier to look for.Happy to provide anything else you need (
dshversion output, full log lines, the decoded JSONL,app.asarextraction).Session
…\session-ec57579b-8cb9-4b35-aa37-fe36c235e5eb\session.jsonl.zstd27,245,912B, mtime2026-09-10 17:41:51 (+08:00); decodes cleanly withzstd -d(exit 0, no truncation warning) to165,369,254B /170,416JSONL records.{"type":"session","version":0,"id":"session-ec57579b-…","createdAt":1786824256927,"cwd":"D:\\Users\\34745\\Desktop\\github\\deepseek-aideepseek-harness","delegationDepth":0,"agentPreset":"standard"}→ source format v0.2026-08-15T20:04:16.975Z→2026-09-10T09:41:51.990Z; 461 turns, 9,042step/startevents, 43session/end-seedboundaries, 119 compactions.Measurements
seqseqseq0(the other is the session header)[0, 4,125,203]seq0compact records (18,192 gap runs)assistant/messagerecordsseq=4060703, turn 450 step 24[[4059829, 4060702]])chunkCount)4059830 … 4060702)4059829=session/end-seed,data: {}The 873 decomposes as 12 explicit
assistant/chunkevents plus 861 seq numbers inside 18 compact records (reasoning-chunks/text-chunks/tool-call-chunks) for that step, whichtransformReleasedRunfolds into the attempt.How I obtained the 8,969 / 1 split (so you can judge it): I decoded the JSONL and re-implemented the span accounting and
matchesChunkSourcesin a script, treating each compact*-chunksrecord as one released run spanning[seq0, seq0 + payloadCount). Merging adjacent spans inrecordChunkSpanmakes any grouping of adjacent records equivalent, but I did not inspect the v0→v1 stage, so the exact run boundaries it produces are unverified. I could not execute the shipped migration directly. The conclusion that all earlier messages pass does not depend on that emulation: the migration throws on the first offending message and reports4060703.Avoidable red herring
Counting only materialised
assistant/chunkevents understates every attempt in this file, because most deltas live in the compactseq0records. For the failing step that count is 12 against 874 declared; for a passing step such as turn 1 step 2 it is 8 against 128 declared. Those large ratios are expected and are not the defect.What is actually different in this file
Immediately around the failing attempt (records in file order):
Observations for this attempt, all measured:
step/startrecord for turn 450 step 24 anywhere in the file;block-startchunk for the reasoning block — the attempt opens withreasoning-deltaat4059830;seq 4059829is occupied by thesession/end-seed;4059829.session/end-seedis the only record type in this file that anyassistant/messagecites as a source event: across all 8,970 messages, exactly one declared reference (this one) is not part of a chunk event.The declared range is identical in the pre-existing sibling snapshot (below), which is consistent with the range having been written while
4059829was still the attempt'sblock-startand not rewritten when that position became the seed. I have not verified the write order — flagging it as an inference, not a measurement.Control: the older sibling snapshot does not fail
Sibling
session.jsonl.zstd.orig53.bak(53,117,939B, mtime2026-09-09 22:12:18) decodes to164,062,629B /168,477records. Same encoding shape (60,611 compactseq0records, 17,990 gap runs, first run24..386, 8,932assistant/message). It contains the sameassistant/messageatseq=4060703with the samesourceEventSeqs=[[4059829, 4060702]], and under the same re-implementation 0 messages fail.The difference at the boundary (both files order records identically before it):
seq 4059828step/startturn 450 step 24seq 4059829session/end-seedandassistant/chunk(block-start, turn 450 step 24)session/end-seedonlyseqvaluesSo in the backup the replay after the seed re-establishes the attempt's
block-startat4059829and the declared range is exactly satisfied; in the current file the attempt opens at4059830while the declaration still starts at4059829.Minimal reproduction
I do not have a minimal repro file yet. What I can offer instead, on request:
seq 4059740…4060710(the seed boundary plus the whole failing attempt), which I expect is enough to reproduce the refusal;assistant/message/session/end-seed/assistant/chunklines verbatim;session.jsonl.zstd, 27 MB) if you would rather have the original.I have not uploaded anything: this is a private long-running session, so I hesitated to publish it. Tell me which artefact you want, and the smallest one that would let you reproduce this, and I will send it.
Ask
The session data looks intact — nothing else in the file is inconsistent, and the mismatch is a single one-event offset — and I would like to get this session loadable again. Questions, in priority order:
dsh-side command, a re-migration flag, a documented procedure to fix onesourceEventSeqsvalue inside the.zstdJSONL — or should I hand this file to you? I would rather not hand-edit a session store without guidance.sourceEventSeqsmeant to list only chunk events of the attempt, or may it also include non-chunk neighbours such as thesession/end-seedboundary? If the latter is allowed, shouldmatchesChunkSourcescompare against the chunk subset instead of requiring exact list equality?step/startand no reasoningblock-startfor that step, no replayed tail — an expected outcome of a seeded/resumed generation that continues the previous generation's seq numbering, or a sign that the attempt head was lost? Related: couldsourceEventSeqshave been computed from a pre-seed / pre-compaction chunk layout that the file no longer matches? I have not verified which component wrote4059829, so I am phrasing this as a question.All reactions