Repository navigation
Home
Your coding agent forgets everything when a session ends. kg-memory gives it a memory that lasts: how your project fits together, what was decided and why, how you like to work, and which mistakes were already made once. Claude Code and Codex CLI share one memory, so what one learns the other recalls. Antigravity CLI support is experimental (see the Antigravity guide).
It is a plugin for your agent plus a small server on your machine that holds the memory. It is free and open source (MIT), and it stays that way: no paid tier, no account, no API key, no telemetry.
This wiki tracks current main, including unreleased changes; the changelog marks release boundaries.
uv tool install kg-memory # needs uv (https://docs.astral.sh/uv/), which also provides Python 3.10+
kg setup # asks before each change, backs up what it editsThen, if you use Codex, run /hooks and trust the knowledge-graph hooks; if you use Claude Desktop, fully quit and reopen it. Run kg doctor and start a new session. Linux and macOS; Windows is not supported. Details: Installation.
Without memory, every session starts from zero: the agent re-reads yesterday's files, re-derives yesterday's conclusions, and you explain your preferences again. kg-memory keeps what the agent learns and brings it back without being asked:
- At session start, the most important memories are already in context.
- When you type a prompt, memories that match it arrive with the prompt.
- When the agent reads or edits a file, what it learned about that file arrives with the tool result, usually before it changes anything.
- When another session writes something relevant, yours is told, instead of writing the same lesson twice.
What makes it more than a notes file:
- Lessons, not logs. Each memory is a small node: a one-line lesson (the gist), notes with the cases behind it, pointers into the files it concerns, and named links to related memories. The agent writes one at the moment it learns something.
- It stays small. Each session sees a fixed budget of memory. What gets used, credited and connected stays on top; the rest moves to an archive that is still searchable and one read away.
-
It tells you when it needs tidying. Every read carries a
DEBT:line: overlong entries, file pointers that no longer resolve, repeated incidents that should become one lesson./kg-maintaincleans it up, or background upkeep does, if you switch it on. - One memory for all your agents. Claude Code, Codex, Antigravity and Claude Desktop talk to the same local server.
-
Yours to inspect. Plain JSON under
~/.knowledge-graph/, versioned with git if you want history, browsable withkg editor(Visual Editor).
You don't operate the memory; the agent does. You mostly notice it in the answers.
- Start of a session. The most useful memories are loaded before your first message. The agent reads the rest once and says "I have recalled KG Memories".
- During work. You ask about the deploy script, and the note about its one non-obvious flag arrives with your question. The agent opens a config file, and the memory that says "generated, edit the template instead" arrives with it.
- When something is learned. The agent traces a failing test to a timezone assumption, or you say you prefer small commits. It writes that down there and then. You can also just say "remember that…".
- When memory is wrong or missing. Correct the agent as you would anyway. It updates the lesson. If the memory existed but didn't surface, the agent reads it back and reports the miss, which keeps that memory from sinking again.
- Wrapping up. Say you're wrapping up. That is when the agent writes down what the session learned and credits the memories that helped.
- Over weeks. Memories nobody uses sink into the archive; the ones that keep helping stay visible.
To seed a new memory faster, /kg-extract maps a codebase and /kg-scout mines your past sessions for lessons (Skills Reference).
Solid. kg-memory is built by one maintainer and has been in their daily use on real projects since spring 2026, across about 60 releases. Claude Code is the primary client; Codex CLI runs the same plugin against the same memory and is verified against real Codex sessions. Changes reach main through pull requests that run the test suite in CI. The concurrent parts (parallel writes, session forks, renames, saves) were checked with formal models. Saves are atomic, with a rolling backup and optional git history, and releases so far have carried existing memory forward.
Measured, within limits. Prompt and file recall decisions and every credit are logged locally, and a replay evaluator scores ranking changes against that record. There is no published measurement yet of how much the memory improves an agent's results: the public evidence so far is daily use and these logs, not a headline number.
Still moving. It is pre-1.0: minor releases still retune scoring, budgets and the instructions the agent follows. Antigravity support is experimental, and Windows is not supported. Projects are told apart by their folder name, so two projects in folders with the same name share one memory, and project folders must be inside your home directory.
What it costs. Some context: on a mature memory, the session-start preload plus one full read come to about 60,000 characters (roughly 15,000 tokens of English text). Quota only if you switch on background upkeep, which runs short headless sessions on your subscription while your limits have room to spare; that is why it is off by default. Memories the agent reads reach your model provider like any other context; the memory server sends no memory data anywhere on its own.
A local server (Python, port 8765) keeps the graph in memory and saves it as JSON under ~/.knowledge-graph/. Every harness connects through kg mcp, a small stdio bridge that starts the server when it is down. Memory lives at two levels: user (what holds across projects) and project (one codebase). A third level, maintain, is the maintenance agent's own memory and never appears in sessions.
Each level has a fixed budget of 22,000 rendered characters. When a level outgrows it, the lowest-scored nodes are archived; when room frees up, the best archived ones come back. The score blends recency, connectedness and usefulness, where usefulness comes from the agent's explicit kg_useful endorsements of the memories that actually helped. The newest nodes, up to 30% of the budget, are never archived. Hooks deliver the preload, prompt recall and file recall; the agent writes with tools.
See How It Works for the full picture and Design Decisions for why it is built this way and what the design trades.
| Page | What's in it |
|---|---|
| Installation | Requirements, what kg setup changes, steps only you can take, updates, uninstall |
| How It Works | Architecture, data flow, recall, compaction, maintenance, harnesses |
| Server Management |
kg commands, logs, the systemd service, endpoints, troubleshooting |
| Knowledge Graph API | All MCP tools with inputs, outputs, examples |
| Skills Reference | Core, maintain, scout, extract, ops: when and how to use them |
| Visual Editor | Setup, features, limitations |
| Data and Backup | File locations, crash protection, recovery, git history |
| Configuration | Environment variables, background upkeep settings, fixed budgets |
| Design Decisions | Why JSON, why one server, why compress on entry, and the trade-offs |
| Research | What the 2026 agent-memory literature says, read against this system |
Agents can follow the same material as recipes: "run /kg-ops and fix the memory server" is a complete instruction.