Replies: 6 comments 2 replies
|
如果你要在 Agent Sessions 里读 DSH 的 dsh-session-surgeon 是离线扫描/修复这条格式的实现(decode + pack 往返、默认 dry-run)。可以当对照,不要直接当生产 decoder。真实语料请用户自己机器上跑 |
|
A useful implementation boundary for the proposed DSH source: treat the on-disk artifact as a storage format, not as one JSONL event per line. For rc.8, the default path is concatenated checksummed Zstandard frames; the first logical record is a Session header, and streamed assistant deltas may be stored as text-chunks, reasoning-chunks, or tool-call-chunks rows. A reader should decode every frame, expand packed rows through the shipped @deepseek-ai/dsh-session codec, then validate contiguous logical seq values. A parser that only accepts row.seq or only known SessionEvent tags will silently lose output. The handbook format map turns those boundaries into a read-only local validation checklist (copy/hash first, reject unsupported versions, preserve the only source, and test packed/unpacked/mixed layouts): https://sandbaseai.github.io/deepseek-harness-handbook/deepseek-harness-session-log-format.html This is documentation-based guidance, not a claim that I have a real DSH corpus or have completed the Agent Sessions steward step. The official rc.8 sources and verification date are linked in the guide; synthetic fixtures and a volunteer own local corpus remain the right way to validate the integration. |
|
补充一条今天验证过的:格式已经动了,而且 版本线: 我 diff 了 rc.8 和 alpha.2 的 也就是说,外部 reader 现在需要两个 codec,不是一个:packed chunk rows,加上 seq-range provenance。 最关键的一点: @xiaoshenming 你那份 这恰好说明我要的为什么是 steward,而不是一次性实现:写一次、之后不再复查的 source 会静默腐烂,因为版本号不会提醒任何人。所以还是那两个问题:
One thing I verified today, worth adding: the format has already moved, and the The line runs Diffing So an external reader now needs two codecs rather than one: packed chunk rows, plus seq-range provenance. The part that matters most: @xiaoshenming your Which is exactly why I'm asking for a steward rather than a one-off implementation: a source written once and never re-checked will rot silently, because the version number won't warn anybody. So, the same two questions:
Guide: https://github.com/jazzyalex/agent-sessions/blob/main/docs/adding-a-session-source.md |
|
Closing the loop here: Agent Sessions 5.5 shipped local DeepSeek Harness history browsing and search for macOS, alongside 15 other supported coding-agent histories. Thanks to @xiaoshenming and @liyangbing for the format guidance. The current download is 5.5.1. The integration is read-only and its test evidence is still limited: sanitized fixtures, a pinned event catalog, and a local UI smoke test, without broad real-world corpus coverage. If you run DSH on a Mac, I'd appreciate a report on sessions that are missing or misrendered; no private transcript needs to be shared. I posted the full announcement and scope here: new bilingual announcement. This is an independent third-party integration, not endorsed by DeepSeek. 补充一下:Agent Sessions 5.5 已支持在 Mac 本地浏览、搜索 DeepSeek Harness 历史记录,并可与另外 15 种编程 Agent 的会话一起检索。感谢 @xiaoshenming 和 @liyangbing 提供的格式信息。具体功能边界和下载链接见上面的新讨论;请不要公开私人会话内容。 |
|
Congrats on 5.5 — good to see DSH history land in Agent Sessions, and thanks for the citation. If your importer ever runs into a session that refuses to load (seq gap / torn frame / v0-migration refusals), the repair path plus the generation notes (v0 … v4, |
|
Happy to compare notes. Everything below is measured on Generations on disk. The v3→v4 edge changes message sources, not the envelope. A v3 message could carry
Native v4 admission then refuses any source whose A v4 The current generation validates relationships while it is read, and the refusal is whole-file. What we would test an importer against. Five entry points, one fixture: fixtures/probes/run.mjs (two migration stages accept, two restores refuse, a stored current-generation read refuses; it exits non-zero if any row drifts). It is the smallest thing we have that pins "acceptance" to a specific path instead of to a version number. One caveat if you build before/after checks in your importer: the official |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I maintain an open source (811 GitHub stars) Mac app that reads local session history off disk for 13 coding agents — Codex, Claude Code, OpenCode, Cursor, Copilot CLI, Pi, Kimi, Grok, Qwen and a few more, with Devin CLI and fx landing in the next release. Reading other people's transcript formats is basically the whole job, so I spent some time going through
packages/session/session-persistence-jsonland the architecture notes behind it. DSH would be the next one supported, and I'd like someone here to help me do it.The ask
The app is Agent Sessions — MIT, macOS, local-only, one indie dev plus a few harness contributors. It browses and searches transcripts across those agents in one window, with resume where the CLI supports it.
I can write the DSH source myself, but it's better done by someone who actually runs DSH daily. Format work is only as good as the sessions you can test against, and I don't have a real corpus of DSH sessions.
Two sizes of help, and the small one is genuinely small:
docs/adding-a-session-source.md. It's a deliberately closed list — if the compiler surprises you with something not in that doc, that's a bug in the doc. Issue template: new agent source.Either way you get named credit, and DSH sessions stop being the ones nobody can search after the terminal scrolls away.
What DSH gets right, compared to the field
Worth saying, since I read the whole persistence layer and most harnesses don't come out of that reading well:
assistant/chunk. The note in2026-06-14-session-persistence.mdexplicitly turns down Codex's chunk-filtered rollout shape becauseseq = log.lengthandevents[i].seq === ineed a contiguous log. That's the correct call and it's rare. Several agents I read throw away the reasoning or the streaming deltas before they ever hit disk, and once that's gone no reader can bring it back.SessionHeadercarriesversion,cwd,createdAt,parentSession,agentPreset. Because it's a dedicated first frame, listing a directory of huge sessions reads one frame per file. Most JSONL agents make you scan into the body — or the whole file — just to learn a session's cwd and title.runPersistenceContractrunning over both the JSONL bytes and the SQLite rows is a nice trick — it means "which backend" never becomes "which semantics."For context on the bar: I run a small benchmark (Session-Bench, 20 pass/fail gates on how well a harness preserves its own record) across these agents, and losslessness plus cheap listing is where most of them lose points. DSH isn't scored yet — I don't score anything off documentation, only off a real store on disk — but on paper it would be competitive at the top.
One practical note for anyone else writing a reader:
compressiondefaults to'zstd', so the artifact issession.jsonl.zstdand there's nogrepping a session. Standard RFC 8878 frames, so it decodes fine, andpackChunksrows unpack from the codec indsh-session. Just budget for it.All reactions