Skip to content

How It Works

Maxim Mironenko edited this page Oct 9, 2026 · 26 revisions

How It Works

Architecture Overview

 Claude Code · Codex CLI · Antigravity CLI · Claude Desktop
        │ MCP over stdio                  │ hooks (JSON)
        ▼                                 │
 ┌──────────────────────┐                 │              Browser
 │ kg mcp (stdio shim)  │                 │                 │
 └──────────┬───────────┘                 │                 ▼
            │ one POST per message        │       Visual editor :8766
            ▼                             ▼                 │ REST + /ws
 ┌───────────────────────────────────────────────┐          │
 │ Memory server :8765                           │◄─────────┘
 │  /        MCP tools (stateless HTTP)          │
 │  /api/*   hook endpoints, editor REST, survey │
 │  /ws      live updates for the editor         │
 └──────────────────────┬────────────────────────┘
                        ▼
 ┌───────────────────────────────────────────────┐
 │ MultiProjectGraphStore                        │
 │  in-memory graphs, write-through saves,       │
 │  compaction, one server per storage directory │
 └──────────────────────┬────────────────────────┘
                        ▼
 ~/.knowledge-graph/
 ├── user.json            (user level)
 ├── maintain.json        (the maintenance agent's own lessons)
 ├── sessions.json
 ├── recall.jsonl · useful.jsonl · chores.jsonl   (decision logs)
 └── projects/<slug>/graph.json

Every client starts kg mcp, a small stdio program that starts the server when it is down and forwards each MCP message to it as one HTTP request. Hooks post their events to /api/* directly. Claude Desktop chat has no hooks, so it uses the tools alone. The full storage layout is on Data-and-Backup.

Two-Level Model

Knowledge is split into two levels, all stored centrally. A third level, maintain, holds the maintenance agent's own craft lessons; it is never preloaded, searched, or rendered with the other two, and only maintenance runs are meant to write to it (the tool descriptions say so; the server does not enforce it).

Level What Goes Here Storage Location
user Cross-project wisdom: preferences, meta-learnings, architectural principles, personal patterns ~/.knowledge-graph/user.json
project Codebase-specific: architecture decisions, dependencies, debugging discoveries, conventions ~/.knowledge-graph/projects/<slug>/graph.json

The user graph is a singleton: one file, always loaded. A project graph is loaded on demand when a session registers with a project folder.

Nodes and Edges

The graph has two entry types.

Nodes are named lessons, concepts or patterns:

{
  "id": "auth-token-refresh",
  "gist": "JWT refresh uses sliding window; silent failure if session expired",
  "notes": ["discovered during auth debugging: an expired session returned 200 with no new token"],
  "touches": ["src/auth/refresh.py"]
}

Edges are named relationships between nodes, or between a node and a file:

{
  "from": "auth-module",
  "to": "session-handler",
  "rel": "requires-init",
  "notes": ["token validation assumes active session"]
}

An edge endpoint can be a file path, so you don't need a node for every file. A bare id endpoint should name a node, in the same graph or (from a project graph) in the user graph: when a graph loads, edges whose bare-id endpoint matches no node are removed.

Data Flow

  1. Session starts. The SessionStart hook starts the server through kg if needed, then preloads a compact core into context before the first message: the top-scored nodes of both levels, plus a session_id, within each harness's hook limit in the unit it counts (9,500 UTF-16 units in Claude Code, 9,000 bytes in Codex, 10,000 bytes in Antigravity). No tool call is needed, and a one-line message tells the user. The preload says it is a partial view. The server tracks whether the session has made the full read, kg_read(session_id), which renders the rest of the graph without repeating the preloaded gists, and the per-prompt hook keeps reminding until it happens. The "I have recalled KG Memories" announcement belongs to the full read, not the preload. Without a preload (hooks off or untrusted, Desktop chat), the agent's first kg_read(cwd) opens the session; in Codex that read adds a hint when no Codex hook has reached the server for the project. Subagents never receive the preload; the main session puts the relevant gists or kg_* instructions in their prompts.
    • Resume, fork, clear, compaction. In Claude Code and Codex, resume and fork recover the KG session from markers in the transcript; a fork gets its own copy of the session's state, and clear starts fresh. After a compaction the summary has kept only part of what the session was shown, so the session's context state is reset (what it has seen, the preload, the full-read flag) and the core is preloaded again. In the experimental Antigravity adapter, a new checkpoint in the transcript triggers the same reset; compaction after that fix, resume, fork and clear are not yet verified in a signed-in session.
  2. Every prompt. The prompt hook posts the prompt to the server. Until the full read happens the answer is the full-read reminder. After it, the server matches the prompt against both graphs and may inject up to 4 relevant nodes with the prompt (see "Ambient Recall and Capture"). When the server has nothing to say, the hook prints a short reminder from its own rotating pools.
  3. During work. The agent calls kg_put_node and kg_put_edge to capture lessons. The tool hook reports each tool call: for a file the agent reads or edits, the server answers with the memory about that file; for repeated reads and fetches, it may suggest a capture.
  4. Write-through. Content changes and promotions are saved at once, atomically. Read timestamps and maintenance changes are saved by the background thread.
  5. Background maintenance. A background thread (every 30 seconds by default, KG_SAVE_INTERVAL) runs compaction, deletes long-expired orphans and saves. If chores are switched on, a prompt can also start a small detached maintenance run (see "Chores").
  6. Several sessions. Other sessions' writes are pushed to a session on its hook and tool replies (see "Concurrency"). kg_sync(session_id) returns everything other sessions changed since the session's last sync (or its start). It does not report deletions, and a rename shows only as a changed node under its new id.

Failed saves are logged and stay marked unsaved, so they are retried and flushed at shutdown. A crash can still lose changes that were not yet saved.

A first kg_read without a project folder opens user memory only. Home-directory sessions and folderless Codex app chats are also user-only, unless a graph already exists for that folder. A later kg_read(session_id, cwd) can attach a project; a session bound to a project cannot switch to another one. Recall and capture follow the session's scope.

Ambient Recall and Capture

Since 0.9.24 the hooks only carry events: each hook posts its raw payload and prints whatever ready-made output comes back. The server makes every decision, and any hook failure falls back to silence, so a hook cannot break a session.

Recall at the prompt (POST /api/prompt_context). Only the part of the prompt a person typed takes part. Harness records (task notifications, image-paste placeholders, dragged file paths) inject nothing, and path tokens reduce to their file name, which can still match a node's touches. The remaining words (stopwords removed, at least three characters) go through the same search core as kg_search: tokens also contribute their ./_- parts, a light stem matches word forms (schedule ≈ scheduling), adjacent parts form two-word terms weighted by how rare the pair is, a match in a node's id or gist counts more than one in its notes, and IDF is sharpened so one rare term naming the right node is not outvoted by several common ones.

Three gates keep the channel quiet unless it has something:

  • a score threshold, stricter when the prompt has several terms;
  • an evidence gate: a hit must match at least two prompt terms, or a term almost unique in the graph, or a term in its own id or gist; one common word in someone's notes stays silent;
  • a novelty gate: at least one node the session has not seen yet must be among the hits or the nodes connecting them.

Up to 4 hits are injected, within 2,500 characters, nodes named in their id or gist first, then by score. Unseen nodes arrive with their gists; nodes the session already holds appear as bare id (in context) anchors, and the edges between the hits come along, so recall reads as connected knowledge. A gist never injects twice in a session. Edge lines are deduplicated only within one injection, on purpose: a repeated edge is a small reminder of a connection.

Recall at the file (POST /api/tool_event, since 0.9.41). When the agent reads or edits a file (Read, Edit, Write, MultiEdit, NotebookEdit, Codex's apply_patch, Antigravity's view_file and file-edit tools, and shell commands that plainly read one: cat, head, tail, less, sed -n, grep, jq, nl, rg with explicit file operands), the server looks the file up in a reverse index of every node's touches. Relative, ~, absolute and path:lines (anchor) forms all match the file they name. The unseen nodes that name it, archived ones included and not promoted, arrive with the tool result: at most 3 nodes within 1,200 characters, and at most 3 injections per session per 10 minutes. For Codex shell commands, a relative path counts only when the command's directory is known (an explicit absolute workdir, or an exact match to the completed command in the session's rollout); otherwise it is skipped.

Capture on re-derivation (the same endpoint, for Read, WebFetch and WebSearch, and under Codex, which has no Read tool, for shell reads). The server counts targets per project in tool_events.json. A nudge fires only when repetition shows a gap: a file read in a second distinct session, or the same URL or query fetched twice, and no node mentions the target in its touches, gist or notes. First-time reads never nudge, and transient paths (node_modules, virtualenvs, /tmp, .git and similar) never count. Throttles allow at most one nudge per 10 minutes and three per session, and one per target per day. The nudge arrives right after the tool result, while the bottom line is still in the agent's working context.

Prompt- and file-recall decisions for a known session are logged to recall.jsonl, most silences included. Capture nudges are not logged there.

Maintenance Debt

Since 0.9.25 every kg_read and preload renders a DEBT: line per graph after its HEALTH: line, for example:

DEBT: HIGH (0.72) — 14 oversized gist(s), 3 long id(s), 7 unconnected, 2 dangling touch(es), 1 lift cluster(s), never maintained, active 4/7d — smeared: billing×31→billing-service — worth a /kg-maintain pass (or a maintenance subagent) now

The score is staleness × activity × wear:

  • staleness: days since the last stamped maintenance pass, saturating at 14; a graph never maintained counts as fully stale;
  • activity: distinct active days in the last week (node read stamps, and for projects tool-event traffic);
  • wear, counted: gists over 300 characters, active nodes with no edge, long or dated ids (more than five words, or a date in the id), touches that no longer resolve to a file, clusters of dated episodes whose shared lesson has not been written down as a principle, and smeared terms (an entity re-described in prose across many nodes while one undated node plausibly owns it, shown as term×count→hub).

Smearing makes a project's central vocabulary useless to search, because the term appears everywhere. The maintenance pass consolidates one smeared entity per run: facts move into the hub, and the other nodes keep only what is new about them plus an edge. The raw factors print next to the verdict so it can be checked at a glance. Capture rules reach old graphs the same way: a rule becomes a debt factor, and maintenance pays it down.

/kg-maintain is the pass that pays debt down: bounded (capped work per category, about 25 calls), resumable (a kg_progress cursor), and self-stamping. It records kg_progress task "maintain", and only that stamp resets staleness. GET /api/maintenance_debt surveys every graph on disk, neediest first, for any dispatcher: an in-session subagent (the skill includes the dispatch prompt) or a scheduled job.

Chores

Since 0.9.37 the server can pay debt down while you work, if you switch it on (kg setup offers it; or ~/.knowledge-graph/chores.json with "enabled": true, or KG_CHORES=1). It is off by default because it spends quota. A prompt arriving is the trigger: it is the one signal that the machine is awake and this graph is in use. The server picks one debt category and one or two targets itself (up to five episodes for a lift), and launches a detached headless agent whose MCP access is limited to the kg tools, so the working session spends no context on it. Under Codex, shell and hosted web search are also off and the filesystem is read-only. A chore is that small unit. A pass is the full /kg-maintain runbook, dispatched the same way when a graph has gone 21 days without one and the week's quota use is running under its even pace.

Guards: one run at a time, spacing and daily caps, a fresh quota reading, no rename or deletion of a node a live session holds, no repeat of what an earlier pass declined, and no in-place rewrite of a node whose gist changed more than twice in the last 30 days (repeated rewriting is the one maintenance pattern measured to degrade memory; see Research). Anchor chores repair touches whose files moved, using a candidate the server found first (git's recorded rename, or the one file in the project with the same name). Lift chores write the lesson several dated episodes share, once, as a principle. Every decision, refusals included, is logged to chores.jsonl.

Since 0.10.0 a run can go through Claude Code or Codex ("runner": auto, claude, codex), each gated on its own subscription's limits: Claude's status-line gauge, or the ChatGPT plan's windows read from the newest quota event in Codex's recently written session logs (a resumed old session counts). A missing, stale or inconsistent reading refuses the run, and the log says why. Since 0.11.0 antigravity is a third runner: gated on a live agy -p /usage reading, run as a tool-less agent that can reach only the kg tools your Antigravity settings grant, and refused while paid credits could be spent. The Antigravity runner has been checked with a mock model only.

The same gauges drive budget notices (0.11.0). On hook replies it already sends, the server tells a session once per window to plan its wrap-up at 80% of the five-hour window and to wrap up at 90%, and the same at 90% and 95% of the week. Claude Code's reading comes from the status-line file (no notices without it), Codex's from the session's own rollout, Antigravity's from /usage.

Harnesses

The server makes every decision; a harness carries events in and context out.

  • Claude Code and Codex CLI run the same plugin: the same hook scripts, skills and .claude-plugin/ manifest.
  • Antigravity CLI (experimental) has its own native package in the plugin, because its hooks differ. Its MCP results are cut above about 10 KB, so any reply over 3,500 bytes is queued and delivered by the next hook in chunks, and nothing a reply implies (seen state, promotion) is recorded until the last chunk is acknowledged.
  • Claude Desktop uses the tools through kg mcp. Its chat has no hooks, so memory arrives on the first kg_read; its Code tab runs the Claude Code hooks.

What differs per harness on the server side lives in mcp_http/harness.py (plus the Antigravity adapter): which harness sent an event (from the hook's transcript path or the MCP client's User-Agent), the preload limit and the unit the client counts in, the size of one kg_read reply part, whether shell reads count as reads, whether the hook's shell directory can be trusted, and the hint a Codex session gets while the plugin's hooks are untrusted. The survey behind this split, including Cursor (not implemented), is in the repository under docs/harnesses/.

Auto-Compaction

Compaction runs after every node write and on the background tick. On each run at most one of archive, refill and rebalance acts; the orphan pass can follow any of them.

Fresh tier. The newest nodes by creation time are protected: while their render cost (each node's line plus the edge citations it brings live) fits in 30% of the level's budget, they are not scored, never archived, and first back if they were archived. A window by budget rather than days follows each project's pace. If every active node is in the fresh tier, an over-budget graph waits until newer work pushes one out.

Pass 1 — Archive (when the active graph exceeds its character budget):

  1. Score each eligible node with three percentile-ranked signals:
    • Recency: the latest of the last write, the last full node read (kg_read by id), and the latest credit (an endorsement, a repeat, a note credit or a maintenance credit). A credit counts as recent use as of its date: a gist that works never needs the full read, and a standing rule that is right never gets rewritten.
    • Connectedness: weighted edges, in_degree × 0.66 + out_degree × 0.33. An edge to an active neighbour counts 1.0, to an archived neighbour 0.2 (so a cluster that archived together isn't scored as disconnected), to an orphaned one 0. A hub floor of 0.5 × log1p(all incident edges) is used when it is larger, so a hub keeps structural credit while its neighbours are archived.
    • Usefulness: the decayed sum of the node's usefulness stamps (90-day half-life). Most come from kg_useful endorsements: the agent marks the nodes that actually helped, judged at wrap-up against real results, and at the moment it notices one was missing when it was needed (about five a session as guidance, ten at most). Three credits add the same stamps: a repeat (a newer node linked to an older lesson by instance-of, dated on the case's day), a note credit (a case added as a note to a lesson another session wrote counts as this session's endorsement, once per node), and a maintenance credit (see point 5). Reads don't count.
  2. Final score = 0.25 × recency + 0.40 × connectedness + 0.35 × usefulness (tie-aware percentile ranks).
  3. Archive the lowest-scoring nodes until the graph is under COMPACTION_TARGET_RATIO (0.8) of the character budget.
  4. Resurrection: right after archiving, archived and active nodes are scored together. If an archived node outscores a just-archived one by at least 0.05, it comes back to active instead, as long as the swap stays within the budget. A node that became well-connected after it was archived is not stranded.
  5. Maintenance is not use. Reads and rewrites by a session flagged as maintenance (kg_read(maintenance=true), or a pass or chore opening its own kg_progress task) stamp no recency and promote nothing; a rewrite keeps the activity time the node already had. A pass that judges a lesson should stay in view credits it deliberately (kg_useful with credits=1-3, at most 15 per pass). The credit fades like an endorsement, so the lesson stays only if sessions go on to find it useful. A credited orphan returns to the archive, where its score decides. The replay evaluator counts neither repeats nor maintenance credits as use.

Archived nodes get _archived: true. Their edges to active nodes stay visible as memory traces, strings you can pull. An edge between two archived nodes is hidden (you hold neither end) and reappears once either end is promoted. kg_read(session_id, id) promotes a node back to active.

Pass 1r — Refill (when the active graph sits below the fill ceiling): The reverse of archiving. While rendered characters are below 0.8 of the budget, the highest-scored archived nodes are promoted back until that ceiling is reached. After each promotion the remaining candidates are re-scored, because the promoted node's edges just became live and raise its neighbours' connectedness: pull the hub, its satellites re-rank to the top, pull them next. A candidate too large for the remaining room is skipped, so it doesn't block smaller ones behind it. Refill also runs on a run that just archived, filling the room archiving left below the ceiling in the same step; the 0.8 ceiling sits below the 1.0 archive threshold, so the two passes cannot undo each other. When the whole graph fits under the ceiling with every node active, refill brings the whole archive back at once, since promoting the last archived node also removes the archived section's header (findings F7 and F34 in formal/, fixed in 0.15.0).

Pass 1b — Rebalance (between the fill ceiling and the budget, on a run that neither archived nor refilled): The best archived node swaps with the worst active one when it scores higher by at least 0.05, at most three swaps per run, and a swap that would cross the budget is reverted. Without it, whatever was read last would stay active between the two thresholds however the ranking moved.

Pass 2 — Orphan (when archived anchors take more than 30% of the level's budget):

  1. Each archived node costs exactly its rendered id line in kg_read output.
  2. While those lines exceed ARCHIVED_BUDGET_RATIO (30%) of the budget, the archived nodes with the lowest archival score (recency, connectedness and usefulness, as above) become orphaned (_orphaned_ts = now).
  3. Orphaned nodes are invisible in kg_read; they no longer cost context.

Three-Tier Node States

State In kg_read In kg_search Recovery
active gist + its live edges ✓ —
archived id + edges to active nodes ✓ kg_read(session_id, id) promotes to active; refill and rebalance also promote
orphaned invisible ✓ flagged search, then kg_read(session_id, id); a maintenance credit returns it to archived

Chain rescue: reading an archived or orphaned node promotes it to active and brings its orphaned neighbours back to archived, so their ids and edges reappear as crumbs to follow. A maintenance session's reads do neither.

Permanent deletion: an orphaned node not read back within KG_ORPHAN_GRACE_DAYS (365 days by default) is deleted from disk.

Size Accounting (render == charge)

The compactor measures exact rendered characters: it builds the same lines kg_read shows and counts them, so the budget and the visible output cannot disagree.

Component Cost
Active node its rendered id: gist line (notes and touches are read on demand, not charged)
Archived node its rendered id "anchor" line
Orphaned node 0 (invisible)
Live edge its citation line, charged once (cited under its first-rendered endpoint)
Archived–archived edge 0 (hidden: a string you can't pull)

One predicate, core.utils.edge_is_live, decides what counts as a live edge, and both the renderer and the estimator use it. The budgets are fixed: 22,000 rendered characters per level, and a 50,000-character target for the combined full read. The target can be exceeded: active gists are never dropped, and the fresh tier (up to 30% of a level's budget) is never archived. For a graph the compactor has not caught up with, the full read hides the lowest-scored archived anchors first, then the lowest-value edges, never active gists. Node batches (kg_read(ids=[...])) have no combined budget.

How a reply arrives is a separate limit, per client. A kg_read reply longer than the client keeps whole (45,000 UTF-16 units in Claude Code, 36,000 bytes in Codex) goes out in parts cut at line ends, and kg_read(session_id, more=true) returns the next part. A node counts as seen, read or promoted only when the part showing it is delivered. Antigravity queues large replies through its hooks instead.

Transport

Every client launches kg mcp, a stdio shim that forwards each JSON-RPC message as one POST to the server's MCP endpoint and starts the server if it is down. Only a refused connection is retried, so a write never lands twice, and a server restart does not disconnect the tools.

The server runs MCP's Streamable HTTP transport in stateless mode with JSON responses: each request stands alone, with no MCP session id kept between requests. By the author's account it started as a workaround: when the server was first built (late 2025), Claude Code's HTTP MCP client did not send the mcp-session-id header back, so stateful mode failed. It also keeps the shim simple: there are no sessions or event streams to bridge.

Graph sessions are tracked at the application level instead: the session_id from the preload or the first kg_read is passed on every tool call.

Concurrency

  • One lock (threading.RLock) guards the graph store; the session registry has its own (since 0.9.44).
  • Every save is atomic (temp file, fsync, rename). A failed save keeps the graph marked unsaved, so it is retried and flushed at shutdown.
  • Only one server may use a storage directory; a second one refuses to start.
  • Many sessions, from any of the clients, can read and write at the same time.
  • Two agents editing one node are arbitrated optimistically (since 0.10.1): kg_put_node refuses a write built on a stale view (the node changed since this session last saw it) or a partial one (it would replace notes or touches this session never read), and returns the node as it stands so the agent can merge and write again. Notes and touches you send replace the stored lists, so to add a note, read the node and send the full list. The visual editor's writes are not checked, but they are stamped, so agents holding an older view are refused.
  • Other sessions' writes are pushed: prompt and tool hook replies (Claude Code, Codex) and kg_put_node/kg_search replies carry up to three gists of what other sessions wrote since this one last looked, with a pointer to kg_sync for the rest. Each change is pushed once, and only content another session wrote: a rename or a node re-sent unchanged is not announced. Maintenance writes are never pushed, and user-level writes only from a session in the same project. kg_sync returns the full diff.
  • The concurrent and stateful parts were modelled in Lean and the findings reproduced against the real code (formal/ in the repository).

Clone this wiki locally