Repository navigation
Design Decisions
This page explains the reasons behind the major choices. The project went through earlier iterations before the current design: by the changelog's one-line account, a setup built on ByteRover Cipher, then a TypeScript MCP server with Steiner-tree retrieval, then the current compression-first server. The earlier code is not in the public git history, so what this page says about it is the author's account rather than something you can check in the repository.
Earlier iterations used a graph database and a vector database (Neo4j and Qdrant, alongside PostgreSQL). They were abandoned.
What the databases cost:
- Docker containers for several services, with orchestration, health checks and startup ordering
- External API calls for embeddings, a cost on every write
- A maintenance burden far beyond the benefit for a single-user tool
Why JSON works here:
- Human-readable: inspect or edit with any text editor
- Version controllable: meaningful git diffs (the server can auto-commit the storage directory when it is a git repository)
- No dependencies: no services to manage
- Easy for a model: the text rendered from it is plain lines a model reads without any transformation
- Portable: copy the directory to back up or move the whole memory
Trade-off accepted: no database transactions. The graph store and the session registry use locks, saves use temp file + fsync + rename, only one server may hold a storage directory, and optimistic view checks refuse stale or partial node writes. Failed saves stay marked unsaved and are retried; this does not make unsaved work crash-proof.
The previous iteration put its effort into retrieval: inverted indexes, TF-IDF scoring, Steiner-tree path-finding for "surprising connections". It didn't work well enough.
The lesson drawn: clever retrieval can't fix poor storage. If what is stored is verbose and uncompressed, no search algorithm reliably surfaces what matters.
Current approach:
- The model compresses knowledge at capture time, when the insight is fresh and the full context is in its window
- Store only what matters, as a short gist, notes with the reasons, and named relationships
- Preload a compact core into context every session: the SessionStart hook injects the top-scored nodes within each client's hook limit, in the unit that client counts (9,500 UTF-16 units in Claude Code, 9,000 bytes in Codex, 10,000 bytes in Antigravity). The full
kg_readrenders the rest within its own 50,000-character target, showing preloaded gists as id-only anchors so nothing is rendered twice. Both channels use the same render-equals-charge accounting - Read the bounded graph directly; lexical search and a touches index reach deeper content and drive file-triggered recall
Retrieval did not disappear: today's search uses IDF weighting and shortest paths between hits too. The difference is its role. It is a second channel for what sits below a surface the model already reads in full, not the only way in.
How it stays bounded: compaction keeps the active graph within a fixed character budget per level, and a refill pass keeps it from sitting far below that budget. Archived nodes stay on disk, with memory traces (edges to active nodes) pointing to them, and kg_read(session_id, id) brings one back. The budget is a target, not a hard cap: active gists are never cut, and the newest nodes (the fresh tier) are never archived.
MCP supports Streamable HTTP with stateful sessions and server-sent events. By the author's account, when the server was first built (late 2025), stateful mode was tried with Claude Code and failed:
Stateful mode:
Client → POST / (initialize)
Server → Response with mcp-session-id: abc123
Client → POST / (next request)
❌ No mcp-session-id header sent
Server → ERROR
Claude Code's HTTP MCP client did not send the session id back. That was how the client behaved, not a bug in the plugin.
Current approach:
- Stateless HTTP with JSON responses: each MCP request stands alone
- Since 0.12.0 every client (Claude Code, Codex, Antigravity, Claude Desktop) connects through
kg mcp, a stdio shim that forwards each message as one POST. With a stateless server the shim has no session or stream to bridge, and a server restart does not disconnect the tools - Sessions are tracked at the application level: the preload or the first
kg_readregisters a session and returns asession_idthat the agent passes on later calls (the old standalonekg_register_sessiontool was folded intokg_readin 0.9.0) - A WebSocket exists separately for the visual editor, where the client is ours and a live view needs push
Two ways changes reach a client:
- Agents: other sessions' writes are pushed on hook replies and on
kg_put_node/kg_searchreplies; the full diff comes fromkg_sync() - Visual editor: the page subscribes to the graph on screen, so user-graph changes and changes to the project being viewed arrive live over the WebSocket (F6 in
formal/, closed by a modelled subscription protocol) - Same store underneath, different transport needs
Not all knowledge has the same scope or lifetime.
User level: "I tend to over-engineer error handling" or "pytest fixtures beat manual setup". These apply across all projects and are personal to the developer.
Project level: "Auth module requires session handler init first" or "Rate limiting uses a Redis sliding window". These belong to one codebase.
Separating them means:
- Switching projects doesn't lose personal lessons, and one project's details don't crowd another's context
- A project graph is one file (
projects/<slug>/graph.json), so it can be copied or versioned on its own. Shared team memory is not built yet - The user graph has its own storage scope; its text still reaches the model whenever the agent reads it
- Each level is archived independently, each within its own fixed budget (22,000 rendered characters)
A session with no project folder (a Desktop chat, the home directory) opens the user level alone and can attach a project later.
You can create an edge like src/auth.py --requires--> src/session.py without creating nodes for those files. This is deliberate.
Rationale: most file references don't need a node. A node is for a lesson or a concept; the edge itself says how two things relate. A node for every file would fill the graph with low-value entries.
A file endpoint (one containing / or ~) is treated as always present, since you can open the file directly, so an edge from an active node to a file is always a live string and always shown. In practice, file references usually live in a node's touches list rather than in edges; touches travel with the node, show when you read it, and drive file recall. A bare-id endpoint is different: it should name a node in the same graph or, from a project graph, a user node. When a graph loads, edges whose bare-id endpoint matches no node are removed.
Edges from active nodes to archived nodes become memory traces: visible hints that related knowledge exists, ready to pull with kg_read(session_id, id). An edge between two archived nodes is hidden (you hold neither end) and reappears the moment either end is promoted.
Nodes are scored as 0.25 × recency_pct + 0.40 × connectedness_pct + 0.35 × usefulness_pct, with tie-aware percentile ranks.
Why percentiles instead of raw values:
- Raw values have different scales (timestamps, weighted edge counts, decayed stamp sums)
- Percentile ranks put everything on 0.0–1.0
- Equal raw values share a percentile, so a column where most nodes are equal (most nodes have no endorsements) distorts nothing
Why a weighted sum, not a product:
- A product would zero out a well-connected node with stale recency, which is unfair to durable foundational concepts
- A weighted sum lets high connectedness make up for low recency
Why connectedness gets 40%:
- It is the most objective structural signal: how many other nodes depend on or reference this one
- In-degree (others point here, ×0.66) counts more than out-degree (this points elsewhere, ×0.33): incoming edges suggest the node is a dependency, outgoing ones that it is a consumer
- An edge to an active neighbour counts 1.0; to an archived neighbour 0.2, reduced but not zero, so a cluster that archived together isn't scored as disconnected and can be refilled; to an orphaned neighbour 0
- Connectedness is the larger of that weighted value and a hub floor,
0.5 × log1p(all incident edges), so a hub keeps structural credit while its neighbourhood is archived. Orphaning uses the full archival score too, endorsements included, so endorsed nodes and hubs are among the last to go
Why usefulness gets 35%: it is the one signal that says "this was needed" rather than "this was touched". It is the decayed sum (90-day half-life) of the node's usefulness stamps. Explicit kg_useful endorsements write most of them; three credits write the same stamps: a repeat (a later instance-of case under an older lesson), a note credit (a case added to a lesson another session wrote), and a maintenance credit. Reads never count, because a well-formed gist never needs the full read, so counting reads would reward the weakest gists.
Why richness was dropped: content length measures effort, not value. A crisp 80-character gist is better than a verbose one, and a length signal rewarded verbosity and was easy to game.
Why recency counts reads and credits, not just writes:
-
_last_read_tsis stamped on every full node read (kg_readwithidorids), outside maintenance sessions - Every credit counts as recent use as of its date (since 0.14.0): a node that is consulted or endorsed but rarely rewritten stays fresh
- A maintenance pass's reads and rewrites count as neither, since it reads to judge, not to use
New nodes need protection from the "capture, then immediately archive" problem, where compaction runs right after a batch of new nodes. Until 0.13.0 that protection was five days from creation. Days fit no real pace: in a project touched twice a month the grace protected nothing, and in a busy week it held half the active slots (12 of 24 on one graph). Now the newest nodes stay active while their render cost, each node's line plus the edges it brings live, fits in 30% of the level's budget, so the window follows each project's own pace. Once newer work pushes a node out of the tier, it competes on its score.
When compaction archives nodes, an older archived node may by now deserve its place more than one just archived, for example because it gained edges or endorsements after it was archived.
After each archiving pass, a resurrection pass scores archived and active nodes in one pool. An archived node that outscores a just-archived one by at least 0.05 is swapped back to active, if the swap stays within the budget. The margin prevents swapping back and forth on near-equal scores.
Archiving only ever moved knowledge down; the only way back up used to be an explicit kg_read(session_id, id). In practice graphs settled far below budget with their most valuable knowledge stranded in the archive: room existed, but nothing used it.
Refill is the symmetric pass: whenever the active graph sits below the fill ceiling (0.8 × budget), the highest-scored archived nodes are promoted back until the ceiling is reached. Three details matter:
-
Single threshold. Refill triggers below the same 0.8 ceiling it fills to. An earlier design used a separate 0.6 trigger for hysteresis, which created a dead band where graphs at 0.6–0.8 of budget never refilled at all. What prevents oscillation is the gap between the 0.8 ceiling and the 1.0 archive threshold. Archiving can overshoot below the ceiling, so refill runs in the same step and fills that gap at once; until 0.15.0 it waited for the next tick, and a node just archived and written came straight back (F7 in
formal/). - Iterative re-scoring. Promoting a node makes its edges live, which raises its archived neighbours' connectedness, so candidates are re-scored after every promotion. A dense cluster that archived together leads itself back: hub first, then its satellites.
- Skip, don't stop. A top-scored candidate too large for the remaining room is set aside, and smaller candidates behind it still promote, so one oversized node can't block the pass.
Rebalance (0.13.0) covers the band between the two thresholds. Archival acted only above the budget and refill only below the fill ceiling, so in between, whatever was read last stayed active however the ranking moved. On a run where neither acted, the best archived node now swaps with the worst active one when it scores higher by the same 0.05 margin, at most three swaps per run.
A maintenance pass reads dozens of nodes to judge them and rewrites some to tidy them. Counted as use, one pass lifted every node it inspected toward the top of the ranking and promoted low-ranked nodes into the active set. So since 0.13.0 a maintenance session's reads stamp no recency and promote nothing, and its rewrites keep the node's previous activity time. That left a pass no way to say "this lesson should stay in view", so 0.14.0 added a deliberate channel: a maintenance credit (kg_useful with credits=1-3, 15 per pass). It fades like an endorsement, so a propped lesson stays only if working sessions go on to find it useful, and the node records that it was credited, so a later pass can see a lesson that was propped and sank anyway.
Built-in auto-memory can duplicate or contradict the knowledge graph. kg setup offers to turn off Claude Code's auto-memory; with Codex, leave its own memory feature off for the same reason.
Standing instructions still have a role. Keep repository conventions and commands in versioned CLAUDE.md or AGENTS.md guidance; keep accumulated lessons and personal memory in the graph. The recommended setup loads one v3 working style through each harness's supported instruction channel. Remove duplicate copies of that style when migrating.
The first iteration used Gemini embeddings (3072-dimensional vectors) for semantic retrieval. They were dropped because of:
- Cost per write: an external API call for every memory operation
- Latency: a network round trip for every embedding
- Doubtful value: semantic similarity doesn't reliably surface important knowledge. A node about Docker networking and one about container orchestration are similar, but that doesn't make both relevant to the current task
- A bounded active graph: the model reads the compressed overview directly, and lexical search and file matching reach content below it without embedding calls
No comparison against semantic search has been run. The current design uses direct loading plus lexical retrieval because it has been adequate in daily use, not because it has been shown better in general. Embeddings remain a possible future experiment.
The search-quality work in mid-2026 pointed the same way, from live data: when search missed a fact the graph provably held, the cause was usually how it was stored, not how it was retrieved. Chronicle-style capture had re-described entities in prose across dozens of dated event nodes (one product term appeared in 48 node texts while the node that owned it held 4 edges), flattening the term's IDF until no lexical statistic could tell the owning node from the story around it. The fix went to the storage: a name things once capture rule, a write-time hint that suggests an edge to the existing hub instead of re-describing it, and a smeared debt factor that has the maintenance pass consolidate one entity at a time (facts into the hub, other nodes reduced to what is new plus an edge). So far, fixing storage has been enough to keep cheap retrieval working.
Claude Code, Codex CLI and Antigravity CLI (experimental) run against the same local server, and the server makes every decision: what to preload, what a prompt or a file should recall, what maintenance is due. A harness only carries events in and context out. That is why Claude Code and Codex share the hook scripts, the skills and the manifest unchanged, and why the per-harness differences on the server are small: one module (mcp_http/harness.py) for how to tell which harness sent an event, how much context its hooks and tool replies may carry and in which unit, and how it reads files, plus an adapter for Antigravity, whose hooks and reply limits differ enough to need its own package and a queued delivery path. A memory that lives in the server rather than in one harness's configuration is also what lets the harnesses share it: a lesson captured in Codex is recalled in Claude Code.
The survey that set this line, covering Cursor and Antigravity as well, is in the repository under docs/harnesses/. Its deciding question was whether Codex can add context on every prompt; it can, and a real Codex session confirmed the rest. Cursor is not implemented.
A graph wears: gists grow, nodes lose their edges, file pointers go stale, episodes pile up without their lesson written down. A scheduled timer on a laptop reached each graph only about once every 19 days in the 45 days it was measured, because the moment it needs (machine awake, quota readable, usage low) rarely arrives. A prompt arriving is the one signal that the machine is awake and this graph is in use, so that is when the server considers one small chore.
Chores spend quota, so they are off until switched on, and a run is gated on the quota of the harness that runs it, never the one that sent the prompt: a chore run through Codex is gated on the ChatGPT plan's windows, one run through Claude Code on Claude's. Two rules guard against known harm. A chore never renames or deletes a node a live session holds, because that session's next read of the old id would come back NOT FOUND. And no chore rewrites a node whose gist changed more than twice in the last 30 days, because repeated in-place rewriting is the one maintenance pattern measured to degrade memory (see Research).
Many sessions share one server process, and a background thread, a chore dispatcher and request handlers all touch the same state. Reading the code finds some races. Small Lean models with an exhaustive search over bounded interleavings find others, and print the shortest trace to the bad state. A finding counted only when it was reproduced against the real code, and each fix ships with a test that fails on the old code. Not every finding came from Lean: some came from questions the modelling raised, and one from a randomized search over the real compactor.
The first pass found twelve findings (F1–F12). A second pass, in 0.15.0, modelled everything added since 0.11: delivery, the dispatcher and budget notices, the server lifecycle, usefulness accounting, the compaction tick and the push of other sessions' writes. It found 26 more (F13–F38). Thirty-six are fixed or settled as policy. Two low-severity ones remain open, each waiting on a choice of behaviour: F33 (rebalance can swap the same nodes back and forth, which a live server's log does not show happening) and F38 (a hook reply lost to its one-second timeout still uses up the announcement of other sessions' writes; kg_sync keeps the record). The models also caught one candidate fix that was incomplete. The models, reproductions and evidence are in formal/ in the repository.
The graph keeps decisions, explanations and relationships rather than whole conversations. That keeps context small and gives recall useful anchors, but capture is selective and can miss or distort a fact. Session history remains a separate recovery source.
A shared local server supports concurrent agents and a visual editor for inspection and editing. It needs a running Python service, and ambient recall needs client hooks. The stored JSON is portable, while delivery limits and lifecycle behaviour remain specific to each harness. Antigravity support is still experimental.
Replay checks can establish that retrieval reproduces logged decisions and how a change would have altered them. They do not establish how much memory improves task quality: a controlled with/without benchmark has not been built yet.