-
Notifications
You must be signed in to change notification settings - Fork 1
Data and Backup
All data is stored centrally in ~/.knowledge-graph/. The files are plain JSON, so any file backup tool works:
| Data | Path | Scope |
|---|---|---|
| User graph | ~/.knowledge-graph/user.json |
Global, all projects |
| Project graphs | ~/.knowledge-graph/projects/<slug>/graph.json |
Per-project |
| Sessions | ~/.knowledge-graph/sessions.json |
Active session tracking |
| Server logs | /tmp/mcp_server.log |
Ephemeral |
| Server PID | ...server/.mcp_server.pid |
Runtime |
The <slug> is derived from the project directory name (last path component). For example, /home/user/DevProj/my-app → my-app.
Both user and project graphs use the same JSON structure:
{
"nodes": {
"node-id": {
"id": "node-id",
"gist": "Short description",
"notes": ["additional context"],
"touches": ["related/file.py"],
"_archived": true,
"_orphaned_ts": 1706000000.0
}
},
"edges": {
"source->target:relationship": {
"from": "source",
"to": "target",
"rel": "relationship",
"notes": ["edge context"]
}
},
"_meta": {
"versions": {
"node:node-id": {"v": 3, "ts": 1706000000.0, "session": "abc12345"}
},
"progress": {
"scout": {"last_ts": 1706000000, "sessions_reviewed": ["xyz"]}
}
}
}-
_archivedand_orphaned_tsare optional flags on nodes -
_meta.versionstracks change history for sync -
_meta.progressstores persistent task state (scout, extract)
Every mutation (kg_put_node, kg_put_edge, kg_delete_*, kg_read with archived node promotion) saves to disk immediately via atomic write. No data loss on crash or unexpected termination.
All saves use atomic writes to prevent corruption:
- Write to
<file>.tmp -
fsyncto ensure data hits disk -
renametemp to final path (POSIX atomic operation)
If the process crashes mid-write, the temp file is left behind and cleaned up on next save attempt.
Every save keeps one rolling copy of the previous good state as <file>.prev. This covers crash/corruption recovery but not accidental deletion or longer-term history.
# Restore previous state (user graph)
cp ~/.knowledge-graph/user.json.prev ~/.knowledge-graph/user.json
# Restore previous state (project graph)
cp ~/.knowledge-graph/projects/<slug>/graph.json.prev \
~/.knowledge-graph/projects/<slug>/graph.jsonAfter restoring, restart the MCP server (or start a new Claude Code session) to reload from disk.
If graph files are lost entirely, you can reconstruct knowledge from Claude Code conversation history using the kg-scout skill. Scout scans ~/.claude/projects/ JSONL session files for past decisions, preferences, and patterns and rebuilds knowledge graph entries from them.
/skill kg-scout
Scout uses a tension-driven approach — it scans session metadata first and only deep-dives into sessions that show signals of useful knowledge (decisions, corrections, recurring patterns). See Skills Reference for details.
Claude Code stores full conversation transcripts as JSONL files in ~/.claude/projects/. These files are the raw material that kg-scout mines for knowledge recovery — and they're also the archive that lets you trace back any decision or insight from past sessions.
By default, Claude Code deletes session files older than 30 days at startup, controlled by the cleanupPeriodDays setting.
Why this matters: If you rely on kg-scout to recover or fill gaps in your knowledge graph, a 30-day window limits how far back it can reach. Extending this gives you a longer recovery window and richer history for mining.
Recommended: set it to 90 days (or whatever matches your working style):
// ~/.claude/settings.json
{
"cleanupPeriodDays": 90
}You can also use /config in Claude Code's interactive REPL to set this via the Settings UI.
Note: Session files can grow large over time. At 90 days with active use, expect a few hundred MB in
~/.claude/projects/. The knowledge graph itself stays compact — scout extracts the signal and discards the noise.
The plugin does not include a backup scheduler. For versioned history or off-machine copies, set one up externally. Two options:
Simple, no extra tools. Good for occasional manual snapshots.
cd ~/.knowledge-graph
git init
echo "*.prev" >> .gitignore
echo "*.tmp" >> .gitignore
git add -A && git commit -m "initial"Commit on demand, or add a cron job to do it periodically. Not ideal for high-frequency snapshots — every tool call mutates the JSON, so automated hourly commits produce noisy diffs and large history quickly.
Better fit for frequently-changing data. Deduplication means hourly snapshots of mostly-unchanged JSON files cost almost nothing.
# Initialize archive (once)
borg init --encryption=none ~/.knowledge-graph-borgAdd to crontab (crontab -e):
0 * * * * borg create --stats ~/.knowledge-graph-borg::'{now}' ~/.knowledge-graph
0 2 * * * borg prune ~/.knowledge-graph-borg --keep-hourly=24 --keep-daily=7 --keep-weekly=4
Point-in-time restore:
# List available archives
borg list ~/.knowledge-graph-borg
# Extract a specific archive (adjust strip count to your path depth)
cd / && borg extract ~/.knowledge-graph-borg::2026-05-17T03:00Graph files are plain JSON — you can edit them with any text editor. This is intentional.
Safe edits:
- Delete a node: remove its entry from
nodesand any edges referencing it - Edit a gist or notes: modify the text directly
- Unarchive: delete the
_archivedkey from a node
Note: With write-through persistence, the server saves on every mutation. If you edit files while the server is running, use the visual editor or restart the server after manual edits to reload from disk.
Typical graph sizes:
- Small project: 5-20 nodes, 10-30 edges, ~2-5 KB on disk
- Medium project: 20-50 nodes, 30-80 edges, ~10-30 KB on disk
- Active user graph after months: 30-100 nodes, ~15-50 KB on disk
The ~4000-token compaction limit per level (configurable via KG_MAX_TOKENS) keeps the active graph small. Archived nodes stay on disk but don't count toward the limit.