Repository navigation
Releases: Valiant-Codex/claude-fleet-codex
Release list
v0.12.0 — skillify interviews the owner; agent-audit reads a transcript
See CHANGELOG.md, section [0.12.0].
v0.11.2 — advisor-review: choosing the model, and briefing a second round
See CHANGELOG.md, section [0.11.2].
v0.11.1 — What loads at launch is stated once; the other copies become pointers
See CHANGELOG.md, section [0.11.1]: agent-template.md carries the single statement of the always-on set; six restatements become pointers; a tracked .pyc removed; CONTRIBUTING gains a retired-fact sweep and a read of the release diff, both marked as judgement.
v0.11.0 — Fleet-common conduct stated once, imported by every CLAUDE.md
See CHANGELOG.md, section [0.11.0]: a model-independent fleet-conduct.md imported by every CLAUDE.md via @shared/fleet-conduct.md; the template CLAUDE.md brought back in line with the runtime it describes; the 20 KB budget claims removed from the docs.
claude-fleet-codex v0.10.1 — Renamed to claude-fleet-codex: named for what it is, a Claude Code fleet, not a portable one
Nothing executable changes in this release; it is a patch by the owner's call because every script,
unit and template behaves exactly as in 0.10.0. What changes is what the repo says about itself.
Changed
agentic-codexis nowclaude-fleet-codex. Three reasons. The repo runs on Claude Code and on
nothing else, and the old name did not say so. "Agentic" had stopped meaning anything. And a name of
the shapeclaude-…-codexneeds a real word between the two —fleet, which is already the repo's
own vocabulary (fleet-agents,fleet-dormant, the roster) — becauseclaudeandcodexadjacent
read as a hybrid of two vendors' products, which this is not. GitHub redirects the old name for
clones, fetches and links; update your remotes when convenient
(git remote set-url origin https://github.com/<ORG>/claude-fleet-codex.git). Tags and release
subjects before this version keep the old name in their text; they are history.- The claim that brain content ports to another harness is retired. It was stated with an honest
accounting since 0.7.0 and it stopped being true as the framework grew:CLAUDE.mdand its@-imports,
the auto-memory store and its index cliff that the memory model is built around, theSKILL.mdformat,
the sub-agent-based skills (decision-loop,advisor-review,agent-audit), the settings schema and the
SessionStart hook, and a topic's very identity (a Remote Control bridge) are all Claude Code's. Point
another harness at a brain and you get Markdown it cannot act on; nothing here was ever exercised
elsewhere.docs/portability.mdis rewritten around the two things that are true: the deployment
moves between machines (Git as source of truth plus one provisioning step — exercised), and the
ideas travel — a brain as a Git repo of Markdown on GitHub, a shared governance repo every agent reads,
one explicit provisioning boundary, persistent agents as supervised services under separate Unix users,
a nightly memory mirror, a structural drift check, a dead-man's switch, decision records with a
human-gated autonomy line. Rebuild those on any harness; do not copy these files to one. The README,
architecture.md,config-model.md,reference-architecture.md, the brain template's README and the
shared bootstrap no longer promise otherwise.
agentic-codex v0.10.0 — Topics resume at boot; rotation is a verb you run on purpose
Since 0.6.5 a rotate-on-boot unit abandoned every topic's conversation at each reboot, on the
premise that a reboot kills the server-side Remote Control bridge and a resumed session would announce
a dead id forever. On a current runtime the premise no longer holds, and the cost — every topic blank
after every kernel or glibc reboot — was being paid for nothing. The reference deployment measured it:
on Claude Code 2.1.270, fifteen plain-resumed topics reconnected to their pre-reboot bridge ids with
full history, once by a manual restore an hour after a reboot and once at boot itself. The vendor now
documents the mechanism (claude --resume reconnects to the session recorded in the conversation, and
since v2.1.232 mints a replacement itself when the server reports it gone). The boot rotation was
throwing history away to obtain a bridge the runtime keeps.
Changed
- The
rotate-on-bootunit is gone. At boot eachclaude-topic@<key>unit does what it did between
boots: resume the stored conversation.templates/infra/systemd/claude-topic-rotate-on-boot.serviceis
removed;provision-agentno longer installs or enables it and instead removes a stale copy (unit
file plusWantedBylink) from an agent provisioned before this release — a leftover would blank every
conversation at the next reboot.agentic-divergence-checkstops excluding the unit from its
"unexpected user units" assertion for the same reason: a leftover is drift now, not fleet code.
Runtime dependency stated indocs/runtime.md: ≥ 2.1.232. The vendor's own "reconnect history"
note shows the resume-time behaviour changed four times between 2.1.200 and 2.1.232. - What this does not cover is stated in the same place: a topic that comes back announcing its old
bridge id while the server has forgotten it — the client sees no failure, so nothing is minted. That
shape was observed once, after a crash reboot on 2.1.224, and was the reason the boot rotation existed.
Nothing local can detect it (the standing "active is not reachable" rule is unchanged). After a reboot,
open the apps; a topic that answers "session can't be found" getsfork.
Added
claude-topic fork <key>(templates/infra/bin/claude-topic): same conversation, new sessionId,
new bridge id, history carried over —claude --resume <sid> --fork-session --remote-control.
The recovery verb for a bridge that is really dead, whererotatewould cost the context. Same shape
asrotate: the source sessionId is appended totopics.rotatedbefore anything changes and the verb
refuses if it cannot; the new sessionId is captured after start. The fork happens insiderunat exec
time through a one-shot marker (~/.config/agent/fork-next/<key>) thatrunconsumes before exec,
so a fail-loop can never fork twice. The history upload is observed behaviour, not documented by the
vendor — which is exactly why it is a verb you run on purpose and not the boot path.claude-topic runwaits for the network before exec — bounded (120 s,CLAUDE_TOPIC_NET_WAIT),
fails closed (logs and proceeds on timeout). With rotation gone, nothing else restarts a topic that came
up before the network and stayed alive with Remote Control off:Restart=alwaysonly sees a process
that exits, and the old boot unit was, by accident, the late-boot retry that hid this. It is a
readiness gate ("does the API host answer at all"), never a reachability check.claude-topic rotate-all --agent <user>: rotate another fleet agent's topics from any topic that
still answers, by delegating to the target's ownsystemctl --userunder the target's uid with a
clean environment (XDG_RUNTIME_DIR,DBUS_SESSION_BUS_ADDRESSunset,HOMEfrom passwd). Refuses a
user outside the roster and any dormant agent (fleet-dormant), and needs root or passwordless sudo —
it refuses rather than half-rotating and reporting success. Written in the reference deployment on
2026-08-30 for exactly the day this release makes rarer; ported because a fleet-wide dead-bridge day is
still possible and a hand-writtenrunuserincantation is how it used to be recovered.
Fixed
agentic-divergence-checkmeasuresCLAUDE.mdas the runtime sees it. Block-level HTML comments
are stripped before injection, so counting them reported an overage that did not exist and pushed
people into trimming real content to clear it; the reference scan strips them for the same reason — a
path named only in a maintainer note is not a claim the agent can act on. Fixed in the reference
deployment on 2026-08-25 and never propagated; found by this release's diff, which is what the release
procedure exists for.
Docs
docs/runtime.md: the "after a reboot" section rewritten around resume,fork, the network wait, the
runtime dependency and the uncovered failure shape; bring-up recap updated.docs/portability.md: the
rotation digest is now written only when someone runsrotate-all.templates/infra/README.md: unit
row removed.manage-agents(skill and its ops reference): verbs and the post-reboot procedure.
Template hygiene: one reference-deployment name that had survived intemplates/infra/fleet-dormant.
Left private on purpose
install-host-servicesin the reference deployment installs ann8n-backuptimer and resolves its
workspace path from a fixed org; both are deployment-specific and stay out of the template.
v0.9.0 — A console is yours to build
The reference deployment built a tailnet-only web console for its fleet. Its code is deliberately
not in this repo — the reasoning is in the new doc, and it is the same reasoning as app-layer.md:
a console encodes one operator's stack, and a root-privileged web app with a dependency tree would
contradict the "no extra always-on gateway to keep patched" line this repo leads with. What crosses
is the framework-level verb it needed, and everything that was learned building it.
Added
claude-topic list --json(templates/infra/bin/claude-topic). One JSON document for machine
consumers: everystatus --porcelainfield per topic, joined to the live process status the
runtime reports (busy/idle/waitingand what it is waiting for), plus any session of that
user with no topic unit underother_sessions. The join is the point: the runtime knows sessions
by pid and by a display name it invents, which matches neither the topic key nor the registry
name — the one fact both sides share is that/proc/<pid>/cgroupnames the unit the process runs
in. A missing or brokenclaude agentsyieldsstatus: unknownplus awarning, never a silent
idle; proved against three fixtures (binary absent, bogus pid, non-JSON output).other_sessions
publishes both ids, becauseclaude stop|rmtake the runtime's short id and answer
No job matching …if handed the session UUID — a real bug in the reference deployment's console,
found in its own action log.docs/console.md— the pattern, documented not templated: what a console is
for, the identity model on a tailnet, the verb discipline (call the CLI you already have; a missing
verb is a missing verb), background jobs for long commands, confirmations that state what is lost,
and how to verify a page a headless browser is not allowed to reach. Listed indocs/README.md,
and the README says plainly that the code is absent on purpose.
Security
- Documented, with its reproduction: a tailnet identity header is not an authentication check on a
multi-user box. Tailscale Serve addsTailscale-User-Loginand strips client copies, but it
proxies to loopback — so any local process, including an unprivileged agent that ingests untrusted
content, can present the same header and be indistinguishable from the operator. Ask the kernel
instead: the process that owns the socket resolves the client's uid (/proc/net/tcp, or the
Tailscale local API). No header is a refusal, never "local, therefore allowed". - And the trap that nearly shipped, because the lesson generalises. The Node adapter the console
was built on emits two servers: the handler the custom wrapper imports, and the adapter's own
standalone server. Both run the same identity hook; only one owns the socket. The standalone one
binds0.0.0.0by default and forwards client headers untouched — so the uid the app read was
whatever the caller typed. It sat unused in the installed tree on a host with a public IP; started
by hand, an unprivileged local user with two forged headers got a200. Nothing ran it, so nothing
was breached, and the application code was correct throughout: the hole was a second path to it.
The fixes are in the doc — a per-process token the wrapper mints and stamps, checked before
anything else, plus not shipping the other entry point at all. The general rule: when a check
depends on owning the socket, make it structurally impossible to reach the app without having done
it.
Changed
- De-identification and release-hygiene fixes the pre-tag checklist is supposed to catch, and had
not. Four comments acrosstemplates/infra/bin/claude-topic,
templates/infra/bin/claude-topic-session-hookand
templates/infra/scripts/agentic-divergence-checkstill named the reference deployment's agents
and its owner; placeholders now, per CONTRIBUTING ground rule 1. And0.8.9and0.8.10shipped
without their CHANGELOG link definitions, so those two versions' links did not resolve — added,
along with this one. The checklist itself is fine; it was run by eye. Both were found by running
it mechanically this time, which is the only way it fires.
v0.8.10 — A budget you raise whenever it stings is not measuring anything
Changed
- The always-on context budget is retired; the divergence check now asserts the runtime's own read
limit on the auto-memory index. A total budget overCLAUDE.md+ its@-imports + the index
derived from nothing on disk: in the reference deployment it was set at 20,000 bytes, then 30,000,
then 40,000 within three days, moved each time it stung, and it watched a number with no loss
attached. Measured on the same box, it would have fired about eight times before the one limit that
actually costs you something — the runtime loads only the first 200 lines or 25 KB ofMEMORY.md
at session start and silently drops the rest, so a memory past that point is already invisible
to recall. Worse, its remedy said to shorten the index, which is the recall surface. Now:
MEMIDX_MAX_BYTES/MEMIDX_MAX_LINESdrift on breach, an[info]line past 80%, and the full
total still prints every run because the trend is the signal. These are vendor constants, not
derived ones — the one deliberate exception to
decisions' derive-the-expectation rule — so the script pins the docs
URL and the date they were read, and takes 25 KB as 25,000 so a stale reading fails towards firing
early rather than towards silence. Proven against lowered thresholds before shipping: silent at the
defaults, drift at 10,000 bytes and at 50 lines,[info]at 16,000. agent-audit's memory pass gains the two growth steps, and the cross-tier checks from the
reference deployment's 2026-09-03 pass. What keeps the index under the runtime's limit is this
pass, with a human reading the diff — never a threshold, and never the session that happened to
notice the number. The steps: retire dead episodic index lines (the line goes, the memory file
stays, so a future session can stillgrepfor it) and reduce curated bullets to verdict + pointer
wherever the bullet already names the decision record holding its reasoning. On the reference agent
13 of 79 index lines were dead episodic notes, and the curated file had grown 2.4× in fifteen days
by carrying reasoning that a named record already held. The same pass also picks up the checks that
had not been ported: whether the two tiers agree rather than merely not duplicating, and ordering
present-state re-verification by the frontmattermodifiedfield instead of file mtime — a bulk
restore resets mtime for a whole store at once, and on the reference store 26 of 72 memories shared
a single mtime minute.
Fixed
- The date of the second over-trim of the reference deployment's curated memory file, in
agent-audit: 2026-08-18, not 2026-08-22 — the 22nd was the restore. Caught by an advisor pass.
v0.8.9 — A restart that reports success can still have lost three days
The runtime mints sessionIds and re-keys them silently on /clear and on compaction. The wrapper
wrote topics.state only when one of its own verbs ran, so between a fork and the next invocation
the stored pointer was stale and nothing said so — and the next restart resumed the stale branch,
orphaned everything said since, and reported success, because by its own definition it had
succeeded. In the reference deployment: six orphaning events across two agents, the largest
losing 8.5 MB and three days. One of them was caused by a self-restart armed specifically to
demonstrate that restarting was safe; the journal recorded nothing but Started.
Added
claude-topic-session-hook— a root-ownedSessionStarthook, installed by
install-host-servicesand registered in each brain'sdeploy/claude-settings.json. It writes
the state at the instant the runtime creates an id, forstartup|resume|clear|compact— exactly
the set that mints them. Root-owned deliberately: an agent must not be able to rewrite the record
of where its own history lives.- Bridge-derived identity in
claude-topic.sid_newest_on_bridgeandbridge_of_sidrecover
the live branch from transcripts on disk. The Remote Control bridge is the durable identity a
sessionId is not, so same bridge = same topic. Derived from disk on purpose: it answers after a
crash and at boot, when there is no live process to ask. claude-topic runadopts. The unattended path systemd re-enters on every restart — where the
damage actually happened — now adopts the newest branch on the bridge when the stored pointer is
stale. Verified counterfactually against both real incidents: it returns the branch that was
live, where the stored pointer returned one frozen for weeks.claude-topic status --porcelain— field/value lines with a computeddriftverdict
(yes/no/unknown/unanchored) andnewest_on_bridge.- The divergence check watches the hook, asserting both that it is installed and that the
agent's settings register it, and reports pointer adoptions from the last 24 hours.
Fixed
rotate-allaborted at the first failing topic, on the boot path. The subshell the comment
credited with preventing exactly this does not: undererrexit, a failing( cmd ); rc=$?kills
the loop beforercis read. One bad topic left every other one announcing a dead bridge while
activeandenabled— the outage shape the verb exists to prevent.rotate-all --dry-runnamed the wrong conversation. It printed the stored id whilerotate
abandoned the live one, so the preview was wrong on precisely the drifted topics that are the
only reason to run it. Both now resolve through one helper.rotate-allaccepted only the exact string--dry-runand fell through to the destructive
path for anything else —--dryrunwould have blanked every conversation on the box.- Transcripts were ordered by the timestamp they declare about themselves, in a tree the agent
can write. A two-line file claiming 2099 won outright and would have beenexec'd. Ordering is
now by mtime. removediscarded the pointer with no record, whererotaterefuses to discard one it
cannot write down.- The hook wrote
topics.statewithout the lock the wrapper uses for that same file. - The divergence check recovered the live id with a regex over prose, then
continued when it
found nothing. That regex returns empty — silently — for a stopped topic, for empty output, for
a reworded line, and for a change of case. - An already-completed split was invisible. Comparing stored against live only ever sees a
split still in progress: two topics reported CLEAN with orphaned branches sitting on disk.
Notes
Do not bind a hook to a systemd Environment=. The first version of this fix passed the topic
key as Environment=TOPIC_KEY=%i. It was wrong in both directions. Absent when needed: the units
were running from before the unit file changed, so twelve of twelve topics refused, and the
mechanism was inert for an afternoon while being reported as verified — the verification had run
against a throwaway topic created after the edit, with no Remote Control bridge at all, so it
exercised neither the unit path nor a bridge. Present when not wanted, which is worse: a systemd
Environment= is inherited by the whole process tree, so any nested claude fires SessionStart
carrying the topic's key and a foreign session id, and the hook writes that foreign id into the
topic's row. A binding that can be absent when you need it and present when you don't is not a
binding. The key is now derived from the bridge in the session's own transcript, which cannot
name a session that is not on that bridge.
The hook runs detached. SessionStart fires before the runtime has written the transcript —
measured: a synchronous hook exhausted its budget at 18:12:21 and the transcript was created at
18:12:21. It lost by being early. A wait long enough to win would be paid by every ordinary
claude run that is not a topic.
mtime ordering is detection, not prevention. It defeats pre-planting; a file planted just
before a restart gets mtime "now" for free. That cannot be closed at that layer — the tree belongs
to the principal it would defend against. Every adoption is therefore recorded and reported.
This supersedes the claude-topic port in 0.8.8. That release shipped the first version of
this fix — bridge derivation plus a TOPIC_KEY-driven hook — from the reference deployment before
the deployment had finished reviewing it. Everything above was found afterwards, by an adversarial
review and by running the mechanism against deliberately broken fixtures rather than a clean one.
If you deployed 0.8.8, the Environment=TOPIC_KEY=%i line it added to claude-topic@.service is
the one thing to remove first.
A structurally better fix exists and was deliberately deferred: one working directory per topic
(WorkingDirectory=%h/.topics/%i). The runtime derives its project directory purely from cwd, so
the key would be readable from dirname(transcript_path) — no bridge, no parsing, no ambiguity —
and it would delete both derivation functions outright. It needs a transcript migration and it
moves where agents' relative paths resolve, so it is a decision to take calmly rather than at the
tail end of an incident fix.
v0.8.8 — Four fixes the reference deployment had already made and never shipped
Divergence between this repo and the deployment it was distilled from is the design. Divergence in the
things that must not diverge — the safety logic, and the skills that shape how an agent decides — is
rot, and it had set in unnoticed in four places. Found by hand, because the propagation check that
should have caught them compares only the infra/scripts pairs whose raises it can parse, and says so
now.
Fixed
claude-topic restartresumed the wrong conversation, silently. It read the STORED session id.
That id is written at create/restart/remember and at no other time, while the running conversation's
id can change underneath it — a compaction, a slash command, a new chat started in the app. From
that moment the state is stale, and the next restart resumes the OLD branch while reporting success,
because by its own definition it succeeded. In the reference deployment that meant a three-day jump
backwards that nobody could see from inside.restartnow reads the live id first, always, and
adopts it when the two disagree; the stored branch stays on disk.agentic-divergence-checkgains
the matching daily assertion, so the drift is visible before anyone restarts.fleet-brain-changeprescribed a commit recipe that commits without pushing. The example fed the
message by heredoc and then chained&& git push.bash -cexecutes each complete command as it
parses, so the commit runs and only the following line fails to parse — leaving the repo
committed-but-unpushed, which is exactly what makes an ff-only sync skip it silently and forever.
The recipe is fixed rather than annotated, and the file gained the Gotchas section that two places
in it already pointed at and which did not exist.
Changed
decision-looprewritten. The shipped version justified its central step with an unsourced
"reacting beats answering open questions 3-4x". A research pass found no controlled study behind it,
and found the adjacent evidence pointing the other way — a written proposal measurably pulls the
reader toward it. The skill was breaking its own rule against numbers a reader cannot check, in its
founding line. It now carries a verifiable four-part trigger with a negative list, a gated
research fan-out (1–5 sub-agents, and none at all when no external source can change the answer, with
the mandatory brief constraints in a newreferences/research-brief.md), a premortem in place of the
deleted claim, an independent advisor pass, the choices put as structured questions, and the record
written in the same turn.- The local-addendum instruction is now conditional. It asserted that every agent keeps one. An
agent whose decisions all land in one shared place needs no addendum and should not grow one for
symmetry: a thin skill overlapping an existing one degrades selection for both. agent-audit's Memory step was one sentence; it is now four checks. Including the one that
exists because a curated memory tier was over-trimmed twice by sessions following the old wording
literally — prove a duplicate with two greps against both the mirror and the runtime store, cite
file:linefor each, or keep it.
Notes
docs/skills.md described decision-loop accurately enough to catch a regression in its own rewrite:
the falsify-assumptions-early step had been dropped. A downstream description found an upstream
defect, which is the argument for keeping such descriptions specific rather than vague.