A working method to cut token cost, extend agent sessions, and improve answer quality — built on the levers that actually move provider bills in 2026, and honest about the ones that don't.
Version: 0.2.0
CACP is a small, portable toolkit (ctx.py) plus a durable memory vault and
vendor-neutral agent adapters. It stops coding agents from re-reading the whole
repo and long histories every turn, keeps the cached prompt prefix stable so
repeated context is nearly free, and measures the result against real provider
usage instead of estimates.
Works with Claude Code, Codex, Cursor, Copilot, Gemini, Cline, Roo, OpenCode, and any agent that can read files or run shell commands.
Full method write-up: METHOD.md. Reproducible measurement: docs/measurement-protocol.md.
| Pillar | Command / rule | What it does |
|---|---|---|
| 1. Stable prefix | ctx pack |
One deterministic startup packet (rules → memory index → repo map → optional digests). A byte-stable prefix keeps the prompt cache hot: ~0.1x input on the API, a longer window on a subscription. |
| 2. Tiered admission | map → digest → read → run |
Climb cheap→expensive; stop at the first rung that answers. No full-repo reads. |
| 3. Durable memory | memory query |
Local, LLM-free top-k retrieval — load only the relevant note blocks, not the whole vault. |
| 4. Gated output | adapter rule | Terse output only when it nets positive (long replies); never on short coding answers; code/commands/errors byte-exact. |
| 5. Measured | measure (+ report) |
Real billed tokens and cache-read share from provider usage logs. Confirm every saving; never assume. |
Retry turns are the most expensive tokens: prefer small patches, verify with
ctx run -- <tests>, and record decisions in memory so the next session doesn't
re-derive them.
- Subscription (Claude Code Max/Pro): fewer admitted tokens per turn + no re-reads + a hot auto-cache → the rolling window lasts longer.
- API (pay per token): a repeated cached prefix bills at ~0.1x input; trimmed admission cuts billed input; gated output cuts the higher-priced output tokens.
CACP will not quote a fixed savings percentage — the real number depends on
your repo and task mix. ctx measure tells you what it actually was.
Requirements: Python 3.10+ (stdlib only). Optional: Obsidian (a human-facing viewer for the memory vault).
ctx.py is a single self-contained file — the fastest integration is to drop
that one file into your project and run init.
Run this in the root of the project you want to set up:
# macOS / Linux / WSL / Git Bash
curl -fsSL https://raw.githubusercontent.com/renf0x/cacp/main/install.sh | bash# Windows PowerShell
irm https://raw.githubusercontent.com/renf0x/cacp/main/install.ps1 | iexIt downloads ctx.py, then runs python ctx.py init to scaffold the memory
vault, agent adapters, and a first cache-stable packet. Non-destructive.
curl -fsSL https://raw.githubusercontent.com/renf0x/cacp/main/ctx.py -o ctx.py
python ctx.py init # scaffolds everything; --agents generic,claude to narrowgit clone https://github.com/renf0x/cacp.git
cd cacp
python install.py C:\path\to\your-project --agents all --open-obsidianPaste this to Claude Code / Codex / Cursor in the target project:
Integrate CACP into this project: download
ctx.pyfromhttps://raw.githubusercontent.com/renf0x/cacp/main/ctx.pyinto the repo root, runpython ctx.py init, then read.ctx/startup-packet.mdand followAGENT_CONTEXT.mdfrom now on.
Any of the above drops the toolkit into the root of the target project:
your-project/
ctx.py # the single-file tool
AGENT_CONTEXT.md # vendor-neutral protocol
AGENTS.md / CLAUDE.md # marked, idempotent adapter block (only agents you pick)
handoff.md
memory/ # durable Obsidian vault
.ctx/startup-packet.md # first cache-stable packet (gitignored)
It never overwrites existing AGENTS.md / CLAUDE.md — it appends one marked,
idempotent adapter block. Pick the agents you use:
| Option | Wires the protocol into |
|---|---|
generic |
AGENT_CONTEXT.md — the vendor-neutral protocol every agent follows |
codex |
AGENTS.md — read by Codex, Cursor, Copilot, Gemini, Cline, Roo, … |
claude |
CLAUDE.md — read by Claude Code |
all (default) |
all of the above |
install.py drops a local ctx.py invoked as python ctx.py <cmd>. For a global
ctx command that resolves from any folder:
pip install .A local ctx.py still takes priority when present, so per-project pinning works.
Install once; the agent then applies the protocol automatically (its instruction
file points at AGENT_CONTEXT.md). At session start:
python ctx.py pack --out .ctx/startup-packet.md # cache-stable startup packet; read it onceThe commands that cover almost everything:
ctx pack --out .ctx/startup-packet.md # pillar 1: stable prefix (read once, append after)
ctx map # see what is expensive before opening files
ctx digest src/large-file.ts # structure instead of the whole file
ctx read src/config.json # full read when unavoidable (logged honestly)
ctx run -- npm test # filter a noisy command; full log saved locally
ctx memory query "how does X work?" # local top-k retrieval (no LLM); --scope project for repo-wide
ctx report # input-side planning estimate from the ledger
ctx measure # REAL billed tokens + cache-read sharectx measure reads actual provider usage — it never estimates.
# Subscription: auto-detects ~/.claude/projects/<slug>/, or pass a transcript file
ctx measure --transcript path\to\session.jsonl
# API: feed the response usage objects and your prices
ctx measure --usage-json usage.json --in-price 5 --out-price 25
# A/B: honest one-command diff of two real runs (baseline vs CACP)
ctx measure --compare runA.jsonl runB.jsonlThe report is split honestly in two: TOOL-CONTROLLABLE (new input admitted
per turn, output, turns — the only block attributable to a workflow/tool) and
PLATFORM CACHE (the provider's automatic 0.1x discount — informational,
never claimed as the tool's saving). A low cache-read share means something is
invalidating the prefix — rebuild it with ctx pack and stop editing context
mid-session.
ctx report aggregates the local ledger (admitted-vs-avoided, input-side,
heuristic chars/3.5). Treat it as a planning signal for climbing the ladder;
confirm dollars/limit impact with ctx measure. See
docs/measurement-protocol.md for a reproducible
A/B you run on your own usage.
ctx init --agents claude wires both hooks automatically (never overwriting an
existing .claude/settings.json):
ctx guard(PreToolUse) — the deterministic ladder. Denies a fullReadof any file above ~4k est tokens and points the agent todigest/retrieval/ ranged reads (offset/limit always pass). Measured on a real run: the model adapted instantly and net effective input dropped 40% at equal answer quality — where instruction-only protocols had failed (see METHOD.md, "Measured reality check").ctx hook(PostToolUse) — logs reads that bypass ctx so the ledger's coverage is honest.
{
"hooks": {
"PreToolUse": [
{ "matcher": "Read",
"hooks": [ { "type": "command", "command": "python ctx.py guard" } ] }
],
"PostToolUse": [
{ "matcher": "Read|Bash",
"hooks": [ { "type": "command", "command": "python ctx.py hook", "async": true } ] }
]
}
}The hook is silent, always exits 0, never blocks the agent, skips ctx's own
commands (so they are not double-counted), and de-duplicates a tool call by its
tool_use_id.
Permissions. Allow only the installed command —
Bash(ctx *)— neverBash(python ctx.py *). A localctx.pytakes priority over the global install, so auto-allowingpython ctx.pywould let any cloned repo'sctx.pyrun without a prompt.
Admission discipline (guard/digest) attacks the wrong axis in a long session:
the dominant cost is the conversation itself — every turn re-sends the whole
history, and only the harness can truncate the live context (/compact,
/clear). No external tool can shrink it. What ctx session does instead is
make truncation early, informed and lossless, so it actually gets used:
session gauge(UserPromptSubmit hook) — reads REAL usage from the transcript and injects a one-line warning only when the live context crosses a threshold (default 80k tok; "compact NOW" at 120k). Below it: silent, free.session save --stdin | --note "..."— the agent writes a small distillate (goal, decisions, open items, key paths) to.ctx/session-state.md. That file survives/compact,/clear, restarts.session snapshot(PreCompact hook) — deterministic floor under any compaction: extracts the last real user asks + files edited from the transcript, even if the agent never saved anything.session restore(SessionStart hook, matchercompact|clear|resume) — prints the state file so the harness re-injects it into the fresh context: ~1k tokens instead of re-reading files to re-derive where the work stood.
The state file is two managed sections (## Agent notes written by save,
## Auto snapshot written by the hook); updating one preserves the other.
.ctx/ is gitignored — session state is volatile by design; durable knowledge
still belongs in memory/.
Hook configs are captured at session start — after wiring these, restart the agent session for them to take effect.
memory/ is a plugin-free Obsidian vault living in the repository, so it is
shared: clone the project and the memory comes with it, readable by every agent.
- Record durable findings in the right journal:
architecture.md,decisions.md,bugs.md,investigations.md,operations.md. Link with[[wiki-links]]. - Keep
MEMORY.mda thin index. Never store source files, large logs, or secrets. ctx memory query "<q>"retrieves the relevant blocks;ctx memory checkvalidates links, size limits, and protected rules;ctx memory rotatearchives closed journal entries.
ctx memory open --install-obsidianmemory/project-rules.md belongs to the user — agents must not weaken it without
approval. After an approved change: ctx memory rules-approve --user-approved.
handoff.md carries only tasks (Now / Next / Blocked / Done this session). It is volatile, so it is kept out of the cached packet — append it
after the packet. This lets another agent continue from verified state without the
previous agent's full conversation.
Earlier versions bundled an RLM sub-agent (9 provider backends + OAuth) and a CodeGraph integration. By 2026 those duplicate native agent subagents/Task tools and IDE/LSP symbol graphs, while adding maintenance, key-management surface, and unproven answer quality — so they were removed, along with all shipped benchmark numbers. What remains is the deterministic, self-hostable core plus the two levers those tools missed: cache-stable layout and measurement against real usage.
.ctx/
__pycache__/
node_modules/
dist/
build/
Commit the whole memory/ vault (notes + .obsidian/app.json /
templates.json) — it is the shared, cross-agent durable knowledge. Only
memory/.obsidian/workspace.json and cache stay machine-local.
CACP is a method, not a guarantee of lower billing or better answers in every
session. It is designed so you can prove or disprove each lever on your own
repositories and providers with ctx measure. If a lever doesn't help your
workload, turn it off.
MIT