Repository navigation
Architecture en
ClarusIubar edited this page Aug 17, 2026
·
3 revisions
한국어 | English
common/ # Logic shared across vendors
├── markdown_safety.py # code-fence safety net
├── text.py # first_sentence / yaml_quote / sanitize_filename / format_callout
├── session_markdown.py # frontmatter + callout markdown assembly, content_hash compute/extract,
│ # turn-metadata HTML comment insertion
├── attachment_cache.py # shared attachment resolver skeleton (caching, dry-run copy, tallying)
├── attachment_types.py # attachment extension classification (embeddability etc, shared across vendors)
├── zip_extract.py # extracts *.zip in data/<vendor>/ in place (includes zip-slip defense)
├── fs_discovery.py # filters archive-tool junk paths like __MACOSX, handles candidate ambiguity
├── upsert.py # content_hash-comparison-based upsert writes (for result/)
├── publish.py # result/ -> real vault mirroring (for --publish, reuses upsert,
│ # recursively mirrors result_dir subfolders)
└── config.py # config.json loader (creates with defaults if missing)
vendors/
├── base.py # vendor module interface contract (Protocol) + runtime validation + auto-discovery
├── chatgpt.py # conversations*.json tree parsing + .dat attachment recovery
├── gemini.py # "My Activity.html" block parsing + local attachment matching
└── claude.py # parses conversations.json (non-project chats) + design_chats/*.json
# (project-attached agentic chats, a different schema)
run.py # CLI: config loading + path priority resolution + vendor execution + publishing
config.example.json # example config.json shape (the real config.json is .gitignore'd)
tests/ # pytest -- pure functions in common/ + vendor parsing logic (tree branch
# selection, KST parsing, etc) + config/publish unit tests
-
Shared logic in
common/, vendor-specific parsing invendors/. ChatGPT's, Gemini's, and Claude's raw export formats are all completely different, but markdown assembly, attachment handling, upsert, and config loading are all vendor-agnostic pure logic, so they're shared in one place. Claude takes this a step further internally — non-project conversations (conversations.json) and project-attached agentic conversations (design_chats/*.json) are two genuinely different schemas with different field names, sovendors/claude.pykeeps two separate loaders and merges them into the same common turn model ({role, text, time_str}). -
Registry (dispatch-table) pattern.
PART_RENDERERSinvendors/chatgpt.pyandTAG_RENDERERSinvendors/gemini.pyare dicts that map tag/content-type to a render function. This was chosen over an if/elif chain because each case is a stateless, independently-handleable lookup-and-dispatch shape. Conversely,gemini.py's_parse_block(a prompt → post_marker → response state transition) is intentionally left as if/elif since it's a genuinely sequential state machine — the registry pattern wasn't applied uniformly, it was chosen based on each piece of code's actual shape.vendors/claude.pykeeps a separate registry per schema (STANDALONE_BLOCK_RENDERERS/PROJECT_BLOCK_RENDERERS), so if Anthropic adds a new content-block type later, anything missing from the registry is silently skipped instead of crashing the pipeline. -
Content-addressable upsert. Rather than deciding purely from whether a
session_idexists, actual change is determined by comparing the rendered body'scontent_hash(see Output Format for details). -
Zip-slip / path-traversal defense.
_is_safe_member()incommon/zip_extract.pyrejects paths that escape the parent directory during extraction, and_is_junk_member()filters out archive-tool junk files like__MACOSX. -
CLI > config > default priority, lazy config loading.
run.pykeepsCONFIG = Noneat module level and only callsload_config()insidemain()— this is what prevents merelyimport run(e.g. during test collection) from creating a realconfig.jsonas a side effect.
Related page: Development