Repository navigation
Design
This page explains why dotclaude works the way it does. Source labels (official, binary, capture, measured, reported, inference) are in Home.
- 0.27.0 rebuilt the core plugin from an empty tree on official extension points. The next section lists the facts behind that choice.
- A bound that Claude Code enforces holds. A bound that only prose states did not hold in the measured week.
- One set of usage bounds, sized for Pro, applies on every plan.
-
plugins/dotclaude/lib/budget.mjsowns each number. - Hooks read state.
They never judge tone or architecture.
One
prompthook onStopjudges the claims of a reply (D27, Stop prompt hooks). - The turn-limit handoff, the agent loop, and the routing rule are 0.19 history. 0.20.0 removed them.
- The rejected alternatives still hold.
Source tags: official is a doc of Claude Code, binary is the 2.1.292 bundle, reported is a public best-practices source or user feedback, and inventory is a count in this repository.
- A plugin
settings.jsoncan keep onlyagentandsubagentStatusLine(official,plugins/components.md). So dotclaude gives a plain settings snippet for your ownsettings.json, and writes no setting. -
force-for-pluginin an output style overrides theoutputStyleof the user.keep-coding-instructionsdefaults to false (official,output-styles.md). - No settings key replaces the system prompt, and only CLI flags do (official,
cli-reference.md). - The prompt has a
sharedblock and asessionblock.CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=1swapssharedfor a lean body (binary). -
DISABLE_COMPACT=1withCLAUDE_CODE_MAX_CONTEXT_TOKENSsets the window, and at the limit the turn ends withblocking_limit(binary). - Opus 5.5 is on the "billed past 200K" list of 2.1.292, so
/contextand the context warnings use a 200K window for it. The block point stays at the full window.CLAUDE_CODE_AUTO_COMPACT_WINDOW=300000moves the display and the warnings to 300K, and a sandbox/contextthen showed19.4k / 300k(binary). The variable takes a plain integer from 100,000 to 1,000,000, and the window of the model caps it. A value such as500kreads as 500 and is raised to the 100K minimum (official,env-vars.md). - The stop at the limit was measured on 2.1.292 in
just sandboxwith the 300K window (measured):- A first prompt larger than the window ended with
terminal_reason: "blocking_limit"and "Prompt is too long", with 0 tokens used. - In a turn that read large files, the context went from 223,461 to 245,427 to 267,792 tokens.
The next read ended the turn with
blocking_limit, after 4 turns, for $0.72. - The debug log had no compaction line, and the transcript had no compact summary.
The 1M window has "no premium for tokens beyond 200K" (official,
model-config.md).DISABLE_COMPACTalso disables/compact(official,env-vars.md).
- A first prompt larger than the window ended with
- Opus 5.5, Sonnet 5.5, and Fable 5.1 have 1M native windows, and Haiku 4.5 has 200K and no effort (binary).
- The 300K window is a judgment. Reports put context rot around 300K to 400K on a 1M window, and one says to keep use under 40% (reported, Open items).
- A top-level
effortLeveldoes not count for Opus 5.5 (official,model-config.md). An effort change keeps the cache, and a model switch, a tool-set change, and compaction break it (official,prompt-caching). - Each tool definition goes in each request, so the snippet removes the tools that dotclaude does not use: Artifact, Workflow, ScheduleWakeup, ReportFindings, and the advisor.
/contextat the start of a sandbox session fell from 32k to 8.1k tokens (measured, Prompt surface). - Deny rules are not a security boundary, so the snippet pairs them with the sandbox (official,
permissions.md). A sandbox cut permission prompts by 84% in one report (reported). - Users want asks only for destroy, reach-out, and secret reads (reported). A guardrail pack made no difference in an A/B, and prompt-only rules decay (reported).
- Hook processes per
Bashcall fell from 7 to 2 with the module in 0.19 (inventory).
0.28.0 makes numbered decisions from D1 to D35. These five set the frame:
- D1: use the least powerful tool for each need. Look first for no key, then a settings key, then a part to turn off, then a mod handler, and last a script.
- D2: dotclaude writes no managed file and runs at tier
userfromsettings.json. - D3: cuts of the prompt surface go by exact id, and an unknown id passes on unchanged.
- D4: the start-of-session bound is a measure plus 5% headroom.
- D5: only an entry that writes user configuration is user-invoked only.
Decisions lists D1 to D35 with the page of each. Rule D35 sets one more frame: dotclaude never makes up a format of its own, and takes only what Claude Code provides. The CHANGELOG lists what it removed.
| Principle | What it means |
|---|---|
| Mechanisms over prose | Prose stays only where no mechanism exists (Usage evidence). |
| Sized for Pro | Larger plans reach their limits later. 0.27.0 has no plan-specific value (Plans and models). |
| One owner for each number |
plugins/dotclaude/lib/budget.mjs holds the bounds. Tests pin its copies in code and config, not in prose. |
| No banned-phrase lists | Claude routes around them with synonyms. The rules name what each behavior does and why. |
| No hooks that judge tone or architecture | A regex cannot tell a needed abstraction from a speculative one. The system prompt and the reviewer agents cover those. |
No agent-type hooks, and one prompt hook |
They spend usage on every event. One prompt hook on Stop, never on SubagentStop (D27). |
| Only documented extension points | dotclaude uses hooks, output styles, skills, agents, userConfig, and settings keys. It never patches Claude Code. |
| Fail open | A bug in a guard must not stop the work of the user. The guards are a best-effort parser, not a sandbox. |
| Layered imports |
plugins/dotclaude/lib/ imports only itself. A folder in features/ imports only lib/ and itself. The entry points (hooks/register.mjs, features/*/cli.mjs, and the skill scripts) import lib/ and features. |
A reply can be in any language. A regex on its phrases blocks right replies and misses wrong ones.
- 0.18.1 removed the announced-work
Stophook, the reply-text tests of theverifygate, and the reply graders of the evals. - The gates read state, such as the edit ledger.
- The same applies to the prompt of the user. 0.19.0 removed the regex that found "remove the tests" in the last prompt and then dropped the assertion ask. The ask now stays, and the user approves it once.
A rule, test, or grader built from a list of failures primes the failures it names. It rewards the wording and not the result. 0.18.1 removed the rules, graders, and prose tests of that shape. It cut the working rules from 8.4 KB to 5.1 KB.
0.27.0 loads plugins/dotclaude/hooks/register.mjs as a hooks module.
The module registers several events, and hooks.json also has classic hooks on PreToolUse, PostToolUse, and Stop (Claude mods, Hooks).
The pure rules are in features/guard/rules.mjs, and the lab in plugins/dotclaude/tests/mod.test.ts runs the module with stubs.
0.22 to 0.26 had a larger module and two classic hooks.
0.19 started hooks/dispatch.mjs once for each event.
measured (2026-09-29, 0.19, 50 Bash calls): 7 hook processes per call became 2.
CPU time fell from about 164 ms to 86 ms per call.
A bare bun start takes 5 ms.
| Scenario | Bound | Mechanism | Check |
|---|---|---|---|
| A subagent works a long task |
maxTurns in the agent file: 16 for test-runner, 31 for investigator, 41 for reviewer, and 61 for implementer and debugger
|
Claude Code ends the agent at the limit | tests/dotclaude/ |
| A subagent model and effort | The efforts in SUBAGENT_EFFORTS for each model, in features/statusline/limits.mjs
|
The model and effort of each agent file |
tests/dotclaude/ |
| The main conversation grows | A 300K window and no compaction |
CLAUDE_CODE_MAX_CONTEXT_TOKENS, DISABLE_COMPACT, and autoCompactEnabled:false in the snippet |
tests/dotclaude/ |
| A repo of another owner with an AI policy | An ask |
tool.check in hooks/register.mjs
|
plugins/dotclaude/tests/mod.test.ts |
A recursive rm
|
An ask |
permissions.ask rules of the settings snippet, with the sandbox |
tests/dotclaude/settings-pin.test.mjs |
| Fan-out | 5 subagents, and 5 agents per workflow, at once | snippet env | none |
| A subagent runs in the background | It runs in the foreground and causes no wake turns |
CLAUDE_CODE_FORK_SUBAGENT=0 in the snippet |
none |
| Text of dotclaude on every request | The first request text is 8 KB to 16 KB, each CLAUDE.md is at most 6 KB, and skill metadata and bodies have bounds |
tools/limits.mjs |
tests/support/sizes.test.mjs fails above a high bound and prints an advisory below a low bound, and just measure checks the first request |
| Weekly review | The usage shares in Usage evidence are reproducible |
dotclaude-usage (dotclaude-sessions) |
none |
The snippet rows need your copy of the settings snippet.
The maxTurns, effort, and ask rows need only the plugin.
0.19 to 0.26 showed the subagent context against AUTO_COMPACT_TOKENS (117k) on the status line, because a subagent compacted at the same point as the main conversation.
0.27.0 has no such status line row.
- measured (914 subagent runs, 2026-09-28 to 2026-10-05): the peak context of a run stopped at about 117k, and 1 run reached 162k.
- 78 of 457
implementerruns passed 100k, so the earlier 100k bound showed a warning for normal runs. - 0.19 enforced a 100k bound (150k for the
reviewer) with a hook, and 0.20.0 removed it.
Fork mode (CLAUDE_CODE_FORK_SUBAGENT) forces every subagent into the background.
A foreground agent returns its report in the turn that spawned it, so it causes no wake turn.
Agents spawned in one message still run together.
- The expected saving is the redundant notification turns, up to about $135 a week, or 7% (inference).
- Claude Code 2.1.285 removed a second, redundant reply that came after each background report in auto mode.
- The first wake turn remains, so forks stay off.
- The main session waits while agents run. Esc interrupts it.
0.19 added a hook that refused every tool except the report tool near the turn limit of an agent.
0.20.0 removed it, and maxTurns alone bounds an agent.
The measurements stay as the reason to keep tasks short.
Measurements and reasons
Claude Code delivers nothing from an agent that it stops at its turn limit.
- measured: Of 13 capped runs after agents learned their limit up front, none reported before the cap. All were still calling tools.
- The 0.19 hook refused every tool except the report tool once about 5% of the limit remained, with a minimum of 3 turns.
- Claude Code counts the limit per invocation. A resume or a wake-up starts the count again. measured: resumed runs reached 100 to 380 calls under a cap of 80.
- Context grows with every turn. The cache-read cost of a task therefore grows with the square of its length. A resume keeps that growth. A fresh agent that starts from a report starts small.
- inference: Splitting a 115-turn task in two cuts its cache reads by about 25 to 30%.
-
reviewerhas a 60-turn limit. A capped review loses its findings. A cappedimplementeronly splits its work. - The
implementerlimit stays 80. measured: in the week to 2026-09-29, 31 of 252 runs reached it, and 27 of those were past 90k context. Most capped briefs named one behavior. Long runs passed 100k context near turn 20, so the 0.19 context bound ended them first. A higher limit gives no more finished work. 0.20.0 keeps 80 forimplementer.
0.20.0 removed the slices skill and the guards that held its oracle, because /goal covers a large change.
The 0.19 loop and the worktree hypothesis
The 0.19 slices skill took the workflow of the Bun, GitHub Copilot, and pnpm v12 Rust ports (reported).
Each port used four parts:
- A guide, written first.
.dotclaude/loop/GUIDE.mdholds the goal, the invariants, and the idiom map. - Slices that start at the leaves. Each slice has one behavior and at most 5 files.
- A frozen test oracle.
loop.jsonlists it asprotectedglobs. The edit and Bash guards deny a change by a subagent to a match. - A reviewer that sees only the diff.
reviewerwith thedifflens gets the git range and the guide, not the report of the implementer. TheStophook blocks once for a slice with the statusimplemented.
The 0.19 mechanisms were hooks because an oracle that the implementer can edit shows nothing. The main conversation is not limited, so the user can still fix a wrong oracle.
One hypothesis was wrong. It said that worktree agents ran commands outside their worktree. measured: the 50 errors were refusals by Claude Code of commands that it cannot check stay in the worktree. The brief of the 0.19 skill told each agent to run plain commands from the worktree root. dotclaude added no guard for this, because the refusal already stops the command.
0.20.0 removed the routing rule and the delegation note, because users report cost blow-ups from subagents.
dotclaude-usage still reports the delegation share.
What the rule said and the evidence
0.19.0 replaced the subagent rule "Work in the main conversation, and use a subagent only for …" with a routing rule. The routing rule tells Claude to delegate work whose tool results it does not need later. It also tells Claude to match the work to the agent descriptions.
- The old rule came from one week in which subagents were over half the cost.
The causes were fan-out and
general-purposeagents, which the snippet andmaxTurnsnow limit. - Each main Opus turn reads the whole main context again (99.9% cache reads, see Usage evidence).
- Users report that the main agent does almost all the work itself.
- measured before the change (2026-10-03, 7 days): 3.8 subagent runs per 100 main turns. The median was 19169 tool-result tokens per main session (baseline).
- The rule is gone, so no one measured whether it changed this share.
The settings snippet sets promptCacheTtl to "5m".
- measured: In 7 days of transcripts, 98.7% of the calls of the main agent came 5 minutes or less after the call before.
- At API prices, the 5-minute cache would have cost 8.5% less than the 1-hour cache.
- A 5-minute write costs 1.25x, and a 1-hour write costs 2x (see the subagent numbers below). So the 1-hour cache pays only when a pause of 5 to 60 minutes is common.
| Alternative | Why rejected |
|---|---|
| A 1-hour subagent cache TTL | About $170 a week worse. See the details below. |
| More or better prose | The 400k handoff rule and "pick the most specific agent" were both in the 0.19 output style during the measured week. |
| Haiku for more subagents | The cost is context times turns. No source that claims a saving gives measurements. Weaker models tend to take more turns. |
| Fixing cache invalidation first | Full cache rewrites were 6% of cost. Most open issues about them (#96101, #96163, #97262, #97342, #97335) are in Claude Code, where a plugin cannot fix them. |
| Trimming the skill listing of the user | The user skills total 1.4 KB of descriptions. dotclaude cuts only its own skill and agent descriptions, to about 150 characters each. |
| A per-session subagent count cap | Concurrency and the context budget already bound the cost. A hard count stops legitimate long sessions. |
| Headroom | Removed in 0.10.0. It compresses tool output with loss. In this repository it dropped words from files that Claude then read as exact. The saving is on tool results only, which are a small part of a cached context. Every agent needed its retrieval tool. |
| Codex delegation | Removed in 0.7.0. It was a second model catalog, quota reader, and profile writer. It also added Claude turns for the relay agents and for review of every GPT diff. inference: one morning of the Codex sessions of the maintainer, scaled from Pro 20x, would use 60 to 100% of a ChatGPT Plus week. |
The 1-hour subagent cache TTL in numbers
- Subagent cache writes in the measured week (2026-09-21 to 09-28) were about 71M tokens. They cost $355 at the 5-minute price and about $568 at the 1-hour price.
- The 1-hour TTL could save at most $46 of rewrites after idle periods. The net result is about $170 a week worse.
- measured again on 2026-10-05 from the request times of 981 subagent runs in 7 days. The figures below are from this second measurement, and not from the measured week. The time between two requests of one subagent includes the tool time, such as a test run. 21,765 gaps were under 5 minutes, 41 were 5 to 60 minutes, and 1 was longer.
- The documented price ratios are 1.25x for a 5-minute write, 2x for a 1-hour write, and 0.1x for a read. At these ratios, the 5-minute TTL cost $796 and the 1-hour TTL cost $967, so 1 hour costs $171 more.
- The 41 long gaps, such as 10 to 30 minute test runs, saved 2.54M tokens of rewrites. But all 65.5M tokens of writes in these 7 days paid the higher price.
- A new agent that reads the cached prefix of an earlier agent of the same kind saves at most $5 more.
- The 1-hour TTL pays only when about each subagent run has a pause of 5 to 60 minutes.
- The snippet does not set
subagentPromptCacheTtl.
- Overview
- Quickstart
- Install
- Plugins
- Settings
- Hooks
- Troubleshooting
- Undocumented reads
- Development
- Design
- Decisions
- Changelog
- Other