-
Notifications
You must be signed in to change notification settings - Fork 1
How It Works
┌─────────────────────────────────────────────────┐
│ Claude Code Sessions │
│ Session A Session B Session C │
└──────┬───────────────┬────────────────┬──────────┘
│ │ │
│ Stateless HTTP (MCP protocol) │
└───────┬───────┘ │
▼ │
┌──────────────────────┐ │
│ MCP Streamable HTTP │ │
│ Server (port 8765) │◄───────────┘
│ │
│ - MCP tools (/) │
│ - REST API (/api/*) │
│ - WebSocket (/ws) │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ MultiProjectGraphStore│
│ - In-memory graphs │
│ - Write-through │
│ - Auto-compact │
└──────────┬───────────┘
│
┌──────┴──────┐
▼ ▼
~/.knowledge-graph/
├── user.json (global)
├── sessions.json
└── projects/
├── my-app/graph.json
└── other/graph.json
Knowledge is split into two levels, all stored centrally:
| Level | What Goes Here | Storage Location |
|---|---|---|
| user | Cross-project wisdom: preferences, meta-learnings, architectural principles, personal patterns | ~/.knowledge-graph/user.json |
| project | Codebase-specific: architecture decisions, dependencies, debugging discoveries, conventions | ~/.knowledge-graph/projects/<slug>/graph.json |
The user graph is a singleton — one file, always loaded. Project graphs are loaded on demand when a session registers with a project_path.
The graph has two entry types:
Nodes — Named concepts, patterns, or insights:
{
"id": "auth-token-refresh",
"gist": "JWT refresh uses sliding window; silent failure if session expired",
"notes": ["discovered during auth debugging 2025-01"],
"touches": ["src/auth/refresh.py"]
}Edges — Relationships between nodes, files, or concepts:
{
"from": "auth-module",
"to": "session-handler",
"rel": "requires-init",
"notes": ["token validation assumes active session"]
}Edges can reference node IDs or file paths directly — you don't need a node for every file.
-
Session starts → Claude calls
kg_read(cwd)which initializes the session and returns the full active graph (both levels) plus asession_id -
During work → Claude calls
kg_put_node/kg_put_edgeto capture insights - Write-through → Every mutation saves to disk immediately (atomic write: temp file + rename)
- Background maintenance → Periodic thread (30s) runs compaction and orphan pruning
-
Multi-session → Other sessions call
kg_sync(session_id)to get changes since their start time
All operations go through the shared in-memory store. Write-through ensures disk is always up-to-date.
Compaction runs in two passes after every write:
Pass 1 — Archive (when active graph exceeds token limit):
- Score each eligible node using two percentile-ranked signals:
-
Recency —
max(last_write_ts, last_read_ts). Reading a node viakg_read(id)refreshes its recency. -
Connectedness — Weighted in/out edges to active nodes only:
in_degree × 0.66 + out_degree × 0.33. Edges to archived/orphaned nodes don't count.
-
Recency —
-
Final score =
0.33 × recency + 0.66 × connectedness(weighted sum of percentiles) - Archive lowest-scoring nodes until graph is under
COMPACTION_TARGET_RATIOof the token limit -
Grace period — Nodes created recently are never archived (see
KG_GRACE_PERIOD_DAYS). Grace is based on creation time only — updates and reads do not reset it.
Archived nodes get _archived: true. Their edges remain visible as memory traces. Use kg_read(cwd, id) to promote them back to active.
Pass 1b — Resurrection (runs after archiving): Re-scores archived + active nodes together. If any pre-existing archived node outscores a just-archived node by ≥0.05, they swap — the better-scoring archived node is restored to active. This ensures archiving history doesn't permanently strand nodes that became well-connected after the fact.
Pass 2 — Orphan (when archived section exceeds 30% of token budget):
- Count archived nodes — each costs ~5 tokens as an ID line in
kg_readoutput - When archived tokens exceed
ARCHIVED_BUDGET_RATIO(30%) ofmax_tokens, demote lowest-connectivity archived nodes to orphaned (_orphaned_ts = now) - Orphaned nodes are invisible in
kg_readandkg_sync— they no longer consume context
| State | In kg_read
|
In kg_search
|
Recovery |
|---|---|---|---|
| active | gist visible | ✓ | — |
| archived | ID + edges visible | ✓ |
kg_read(cwd, id) → promotes to active |
| orphaned | invisible | ✓ flagged | search → kg_read(cwd, id)
|
Chain rescue: reading an archived node promotes it to active AND rescues any of its orphaned neighbors back to archived — their IDs and edges reappear as crumbs to follow.
Permanent deletion: orphaned nodes with no recall after KG_ORPHAN_GRACE_DAYS (365 days) are permanently deleted from disk.
The system estimates token cost to decide when to compact:
| Component | Estimate |
|---|---|
| Base cost per node | 20 tokens |
| Node text | 1 token per 4 characters (gist + notes) |
| Per edge | 15 tokens |
| Archived nodes | Not counted toward limit |
These are rough estimates — the goal is keeping the active graph small enough to fit in Claude's context window without being wasteful.
The server uses MCP's Streamable HTTP transport in stateless mode. Each request is independent — no session IDs preserved between HTTP requests by the MCP layer.
This is because Claude Code's MCP client doesn't preserve session IDs between requests. The server was tested with stateful mode and it failed — Claude Code simply doesn't send the mcp-session-id header back.
Session tracking for sync (kg_sync) is handled at the application level, not the transport level.
- Thread-safe via
threading.RLockon the graph store - Multiple Claude Code sessions can read/write simultaneously
- Last write wins — no conflict resolution beyond that
- The
kg_synctool lets sessions pull changes made by other sessions