Skip to content

Output Format en

ClarusIubar edited this page Aug 17, 2026 · 4 revisions

Output Format

한국어 | English

One markdown note is generated per session (conversation).

  • Frontmatter carries title/session_id/url/date/turns_count/ content_hash/tags.
  • The body lists turns as > [!question]- User (...) / > [!tip]- <Vendor> (...) callouts (preserving Obsidian's display format as-is).
  • Image/file attachments are copied into Attachments/ and embedded with ![[...]] where possible.

Turn-metadata HTML comment

Immediately before each callout there's an HTML comment shaped like this:

<!-- turn: {"turn_index": 0, "role": "user", "parent_turn_index": null, "has_attachment": false} -->

It's invisible in Obsidian's preview, but lets a RAG chunking pipeline read QA pairs, session boundaries, and attachment context directly, without needing to interpret callout syntax ([!question] vs [!tip]) or rely on ordering heuristics like "up to the next question."

  • turn_index: this turn's position within the note.
  • role: user, or the vendor's (assistant) role.
  • parent_turn_index: which question (turn_index) this answer belongs to. Question turns are always null (the start of a new turn window). Even when one question's answer is split across multiple turns (this does happen — a long response getting split into several messages), they all point to the same parent_turn_index.
  • has_attachment: whether an attachment block immediately follows this turn.

This is computed in a single place — common/session_markdown.py::build_session_markdown() — by walking the turn list in role order, vendor-agnostically (vendors/chatgpt.py and vendors/gemini.py aren't involved).

content_hash

content_hash is a SHA-256 hash of the entire body (excluding frontmatter). Since the turn-metadata comments are part of the body, they're automatically included in the hash. This value is what drives the upsert decision (created/updated/unchanged) described in Configuration.

Note: filenames are based on session_id, not title

ChatGPT's conversations.json provides a real title field for every conversation (the same title you see in the chat list; it falls back to the first user message's first sentence only when that field is empty). Gemini's My Activity.html, on the other hand, has no per-conversation title at all — the only field that looks like one (class="... title") isn't a per-session value, it's just the fixed product name "Gemini Apps", identical in every block. So Gemini always uses the first question's first sentence as the title instead (first_sentence(...)).

Because of this difference between vendors, both vendors use session_id, not title, as the filename (sanitize_filename(cid/sid, ...) in vendors/chatgpt.py/ vendors/gemini.py). If title were used as the filename instead:

  • Gemini's title is derived from the first message, so its stability differs from ChatGPT's; two different sessions could also collide on the same first sentence.
  • Upsert assumes "same session = same file path" (see Configuration). Using title as the filename means a change to the title-deriving logic, or even a slightly different first message, could create a new file and sever the link to the existing note.

session_id comes straight from the source (ChatGPT's conversation_id, Gemini's session URL) as a stable, unique identifier, so none of that happens — title only ever shows up for display in the frontmatter; the actual file path and upsert decision are always based on session_id.

Claude's conversations.json (non-project chats) usually has a real name, but design_chats/*.json (project chats) often just says "Chat" if the user never renamed it — for the same reason as Gemini, Claude also uses the conversation's uuid (session_id), not title, as the filename.

Claude project-attached conversations get their own subfolder

Other vendors write flat into result/<vendor>/*.md, but Claude writes project-attached conversations into result/claude/<project name>/*.md subfolders (non-project conversations are written flat, same as the other vendors). The project name comes from project.name, embedded inline in design_chats/*.json; characters that aren't valid in filesystem paths are cleaned up with sanitize_filename(). --publish reproduces this subfolder structure as-is in the vault (see Configuration).

Out of scope for Claude

memories.json (memory feature summaries), login_history.json, users.json (account info), and the docs field inside projects/*.json (project knowledge files) aren't conversations, so none of them get converted. This export never contains actual attachment bytes (only reference filenames), so attachments always show up as a "missing" notice.

Related page: Architecture

Clone this wiki locally