Repository navigation
Output Format en
한국어 | English
One markdown note is generated per session (conversation).
- Frontmatter carries
title/session_id/url/date/turns_count/content_hash/tags. - The body lists turns as
> [!question]- User (...)/> [!tip]- <Vendor> (...)callouts (preserving Obsidian's display format as-is). - Image/file attachments are copied into
Attachments/and embedded with![[...]]where possible.
Immediately before each callout there's an HTML comment shaped like this:
<!-- turn: {"turn_index": 0, "role": "user", "parent_turn_index": null, "has_attachment": false} -->It's invisible in Obsidian's preview, but lets a RAG chunking pipeline read
QA pairs, session boundaries, and attachment context directly, without
needing to interpret callout syntax ([!question] vs [!tip]) or rely on
ordering heuristics like "up to the next question."
-
turn_index: this turn's position within the note. -
role:user, or the vendor's (assistant) role. -
parent_turn_index: which question (turn_index) this answer belongs to. Question turns are alwaysnull(the start of a new turn window). Even when one question's answer is split across multiple turns (this does happen — a long response getting split into several messages), they all point to the sameparent_turn_index. -
has_attachment: whether an attachment block immediately follows this turn.
This is computed in a single place —
common/session_markdown.py::build_session_markdown() — by walking the turn
list in role order, vendor-agnostically (vendors/chatgpt.py and
vendors/gemini.py aren't involved).
content_hash is a SHA-256 hash of the entire body (excluding frontmatter).
Since the turn-metadata comments are part of the body, they're automatically
included in the hash. This value is what drives the upsert decision
(created/updated/unchanged) described in Configuration.
ChatGPT's conversations.json provides a real title field for every
conversation (the same title you see in the chat list; it falls back to the
first user message's first sentence only when that field is empty).
Gemini's My Activity.html, on the other hand, has no per-conversation
title at all — the only field that looks like one (class="... title")
isn't a per-session value, it's just the fixed product name "Gemini
Apps", identical in every block. So Gemini always uses the first
question's first sentence as the title instead (first_sentence(...)).
Because of this difference between vendors, both vendors use
session_id, not title, as the filename
(sanitize_filename(cid/sid, ...) in vendors/chatgpt.py/
vendors/gemini.py). If title were used as the filename instead:
- Gemini's title is derived from the first message, so its stability differs from ChatGPT's; two different sessions could also collide on the same first sentence.
- Upsert assumes "same session = same file path" (see Configuration). Using title as the filename means a change to the title-deriving logic, or even a slightly different first message, could create a new file and sever the link to the existing note.
session_id comes straight from the source (ChatGPT's conversation_id,
Gemini's session URL) as a stable, unique identifier, so none of that
happens — title only ever shows up for display in the frontmatter; the
actual file path and upsert decision are always based on session_id.
Claude's conversations.json (non-project chats) usually has a real name,
but design_chats/*.json (project chats) often just says "Chat" if the user
never renamed it — for the same reason as Gemini, Claude also uses the
conversation's uuid (session_id), not title, as the filename.
Other vendors write flat into result/<vendor>/*.md, but Claude writes
project-attached conversations into
result/claude/<project name>/*.md subfolders (non-project conversations
are written flat, same as the other vendors). The project name comes from
project.name, embedded inline in design_chats/*.json; characters that
aren't valid in filesystem paths are cleaned up with sanitize_filename().
--publish reproduces this subfolder structure as-is in the vault (see
Configuration).
memories.json (memory feature summaries), login_history.json,
users.json (account info), and the docs field inside projects/*.json
(project knowledge files) aren't conversations, so none of them get
converted. This export never contains actual attachment bytes (only
reference filenames), so attachments always show up as a "missing" notice.
Related page: Architecture