Replies: 2 comments
|
这个问题与 SandBase Harness 的事件恢复目标相同,但存储实现不同:SandBase v0.3.8 的 canonical session log 是 SQLite 中的 append-only events 表,不是 DSH 的 zstd JSONL 文件,因此不能直接作为 zstd reader 修复方案。 可参考的恢复原则是:
实现与测试:
对 zstd 文件格式本身,我赞同你提出的“完整 frame 内的截断尾记录按可恢复尾处理”:保留 scanner 已确认的完整记录,截断到受影响 frame 起点,并把 repair 结果写入审计信息;不要用一次尾部格式错误否定此前所有完整事件。 |
|
Follow-up to my own report above, after a deeper read of the sources. Three corrections and one honest narrowing; the overall conclusion (demote this class to the existing torn-tail contract) stands. 1. The damage mechanism in the repro needs correcting. As written, step 2 implies a non-graceful kill can directly leave the last complete frame with an unterminated final record. The writer's behavior does not support that:
A plain mid-write cut therefore always lands in the tolerated class. Class (b) — a complete, checksummed frame whose plaintext tail record lacks its newline — requires a crash/persistence-timing window around a frame's write+fsync. We have no deterministic reproducer for that window; the check's existence shows the class is anticipated, and our field data (180 of 267 session logs failing with this error) shows it is the common outcome in practice. Validate with a unit test that constructs the state directly: compress a batch lacking the final newline into a complete frame, append it after a valid header frame, and assert the demotion contract (no throw; truncate at that frame's start; complete records recovered). One precision: the demotion preserves every complete record before and inside the offending frame; content after a known-bad frame is dropped by design as untrustworthy — the loss bound is per damage site, not an unconditional zero. 2. The artifact path is missing a directory level. I wrote The project-key level groups sessions by working directory; the session id is path-encoded. 3. Two scope claims were overstated.
Environment: Thanks to the maintainers and the community. This follow-up is the result of a deeper source read before submitting the patch proposal — better exactly right than fast. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
@deepseek-ai/dsh-session-persistence-jsonlstores each session as a concatenation of zstd frames, where every frame carries a batch of JSONL records (the first frame is a header). When the writer is killed mid-append, the last record of the final complete frame can end up inside the frame without its trailing newline. In that casereadZstdPrefix()throwscorrupt Zstandard session log: complete frame contains a torn JSONL record,readPrefix()fails, and the entire session history becomes unreadable — even though every complete record before the damaged tail is perfectly valid. Structurally torn final frames already have a tolerance/recovery path; this adjacent damage class does not, and one missing newline byte currently poisons a whole session.Background & Evidence
<workspace>/<session-id>/session.jsonl.zstdis a concatenation of zstd frames (magic28 b5 2f fd); each frame holds a batch of JSONL records; the first frame holds a single header line.SessionLogScanner, the reader comparescheckpoint()values and hard-throws whencommittedBytes !== inputBytes, i.e. when the plaintext of the last complete frame ends with an unterminated JSONL record (readZstdPrefix()indsh-session-persistence-jsonl/lib/index.js, ~L983-1011 in 0.1.1-rc.2).session.jsonl.zstdfiles failed to load with exactly this error. A local repair that truncates the damaged tail and re-appends the complete records restored all of them with zero event loss — i.e. the data was recoverable and the hard rejection was over-strict.Failure Mode or Reproduction
taskkill /Fon Windows, OOM kill, power loss) while a frame append is in flight, such that the last record's bytes reach the frame but its trailing newline does not.Observed: the session fails to load with
corrupt Zstandard session log: complete frame contains a torn JSONL record; there is no recovery path in the UI and the history is effectively lost from the reader's point of view.Note: whether a given kill lands in the tolerated "structurally torn final frame" path or in this zero-tolerance branch depends on exact flush timing, so the same termination can produce either outcome.
Proposed Fix
Demote this case to the same contract already used for a structurally torn final frame, instead of throwing:
tornMarker.truncateTo= physical start of that frame;recoveredEvents= the complete events inside that frame (the scanner'sfinish()already ignores an unterminated tail record).commitRepair()path truncate the artifact and re-append the recovered events plus closers. Net event loss is zero; truncation granularity is one whole frame (zstd frames are checksummed, so a mid-frame cut is not possible).We have a working local patch against 0.1.1-rc.2 implementing exactly thishappy to share the diff in this thread if that would help.
Additional Context
@deepseek-ai/dsh0.1.1-rc.2 (developer preview)All reactions