Skip to content

Releases: Valiant-Codex/claude-fleet-codex

v0.12.0 — skillify interviews the owner; agent-audit reads a transcript

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 01 Oct 18:48

See CHANGELOG.md, section [0.12.0].

v0.11.2 — advisor-review: choosing the model, and briefing a second round

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 14 Sep 08:28

See CHANGELOG.md, section [0.11.2].

v0.11.1 — What loads at launch is stated once; the other copies become pointers

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 13 Sep 20:39

See CHANGELOG.md, section [0.11.1]: agent-template.md carries the single statement of the always-on set; six restatements become pointers; a tracked .pyc removed; CONTRIBUTING gains a retired-fact sweep and a read of the release diff, both marked as judgement.

v0.11.0 — Fleet-common conduct stated once, imported by every CLAUDE.md

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 13 Sep 19:09

See CHANGELOG.md, section [0.11.0]: a model-independent fleet-conduct.md imported by every CLAUDE.md via @shared/fleet-conduct.md; the template CLAUDE.md brought back in line with the runtime it describes; the 20 KB budget claims removed from the docs.

claude-fleet-codex v0.10.1 — Renamed to claude-fleet-codex: named for what it is, a Claude Code fleet, not a portable one

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 12 Sep 21:09

Nothing executable changes in this release; it is a patch by the owner's call because every script,
unit and template behaves exactly as in 0.10.0. What changes is what the repo says about itself.

Changed

  • agentic-codex is now claude-fleet-codex. Three reasons. The repo runs on Claude Code and on
    nothing else, and the old name did not say so. "Agentic" had stopped meaning anything. And a name of
    the shape claude-…-codex needs a real word between the two — fleet, which is already the repo's
    own vocabulary (fleet-agents, fleet-dormant, the roster) — because claude and codex adjacent
    read as a hybrid of two vendors' products, which this is not. GitHub redirects the old name for
    clones, fetches and links; update your remotes when convenient
    (git remote set-url origin https://github.com/<ORG>/claude-fleet-codex.git). Tags and release
    subjects before this version keep the old name in their text; they are history.
  • The claim that brain content ports to another harness is retired. It was stated with an honest
    accounting since 0.7.0 and it stopped being true as the framework grew: CLAUDE.md and its @-imports,
    the auto-memory store and its index cliff that the memory model is built around, the SKILL.md format,
    the sub-agent-based skills (decision-loop, advisor-review, agent-audit), the settings schema and the
    SessionStart hook, and a topic's very identity (a Remote Control bridge) are all Claude Code's. Point
    another harness at a brain and you get Markdown it cannot act on; nothing here was ever exercised
    elsewhere. docs/portability.md is rewritten around the two things that are true: the deployment
    moves between machines
    (Git as source of truth plus one provisioning step — exercised), and the
    ideas travel
    — a brain as a Git repo of Markdown on GitHub, a shared governance repo every agent reads,
    one explicit provisioning boundary, persistent agents as supervised services under separate Unix users,
    a nightly memory mirror, a structural drift check, a dead-man's switch, decision records with a
    human-gated autonomy line. Rebuild those on any harness; do not copy these files to one. The README,
    architecture.md, config-model.md, reference-architecture.md, the brain template's README and the
    shared bootstrap no longer promise otherwise.

agentic-codex v0.10.0 — Topics resume at boot; rotation is a verb you run on purpose

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 12 Sep 20:34

Since 0.6.5 a rotate-on-boot unit abandoned every topic's conversation at each reboot, on the
premise that a reboot kills the server-side Remote Control bridge and a resumed session would announce
a dead id forever. On a current runtime the premise no longer holds, and the cost — every topic blank
after every kernel or glibc reboot — was being paid for nothing. The reference deployment measured it:
on Claude Code 2.1.270, fifteen plain-resumed topics reconnected to their pre-reboot bridge ids with
full history, once by a manual restore an hour after a reboot and once at boot itself. The vendor now
documents the mechanism (claude --resume reconnects to the session recorded in the conversation, and
since v2.1.232 mints a replacement itself when the server reports it gone). The boot rotation was
throwing history away to obtain a bridge the runtime keeps.

Changed

  • The rotate-on-boot unit is gone. At boot each claude-topic@<key> unit does what it did between
    boots: resume the stored conversation. templates/infra/systemd/claude-topic-rotate-on-boot.service is
    removed; provision-agent no longer installs or enables it and instead removes a stale copy (unit
    file plus WantedBy link) from an agent provisioned before this release — a leftover would blank every
    conversation at the next reboot. agentic-divergence-check stops excluding the unit from its
    "unexpected user units" assertion for the same reason: a leftover is drift now, not fleet code.
    Runtime dependency stated in docs/runtime.md: ≥ 2.1.232. The vendor's own "reconnect history"
    note shows the resume-time behaviour changed four times between 2.1.200 and 2.1.232.
  • What this does not cover is stated in the same place: a topic that comes back announcing its old
    bridge id while the server has forgotten it — the client sees no failure, so nothing is minted. That
    shape was observed once, after a crash reboot on 2.1.224, and was the reason the boot rotation existed.
    Nothing local can detect it (the standing "active is not reachable" rule is unchanged). After a reboot,
    open the apps; a topic that answers "session can't be found" gets fork.

Added

  • claude-topic fork <key> (templates/infra/bin/claude-topic): same conversation, new sessionId,
    new bridge id, history carried over — claude --resume <sid> --fork-session --remote-control.
    The recovery verb for a bridge that is really dead, where rotate would cost the context. Same shape
    as rotate: the source sessionId is appended to topics.rotated before anything changes and the verb
    refuses if it cannot; the new sessionId is captured after start. The fork happens inside run at exec
    time through a one-shot marker (~/.config/agent/fork-next/<key>) that run consumes before exec,
    so a fail-loop can never fork twice. The history upload is observed behaviour, not documented by the
    vendor — which is exactly why it is a verb you run on purpose and not the boot path.
  • claude-topic run waits for the network before exec — bounded (120 s, CLAUDE_TOPIC_NET_WAIT),
    fails closed (logs and proceeds on timeout). With rotation gone, nothing else restarts a topic that came
    up before the network and stayed alive with Remote Control off: Restart=always only sees a process
    that exits, and the old boot unit was, by accident, the late-boot retry that hid this. It is a
    readiness gate ("does the API host answer at all"), never a reachability check.
  • claude-topic rotate-all --agent <user>: rotate another fleet agent's topics from any topic that
    still answers, by delegating to the target's own systemctl --user under the target's uid with a
    clean environment (XDG_RUNTIME_DIR, DBUS_SESSION_BUS_ADDRESS unset, HOME from passwd). Refuses a
    user outside the roster and any dormant agent (fleet-dormant), and needs root or passwordless sudo —
    it refuses rather than half-rotating and reporting success. Written in the reference deployment on
    2026-08-30 for exactly the day this release makes rarer; ported because a fleet-wide dead-bridge day is
    still possible and a hand-written runuser incantation is how it used to be recovered.

Fixed

  • agentic-divergence-check measures CLAUDE.md as the runtime sees it. Block-level HTML comments
    are stripped before injection, so counting them reported an overage that did not exist and pushed
    people into trimming real content to clear it; the reference scan strips them for the same reason — a
    path named only in a maintainer note is not a claim the agent can act on. Fixed in the reference
    deployment on 2026-08-25 and never propagated; found by this release's diff, which is what the release
    procedure exists for.

Docs

  • docs/runtime.md: the "after a reboot" section rewritten around resume, fork, the network wait, the
    runtime dependency and the uncovered failure shape; bring-up recap updated. docs/portability.md: the
    rotation digest is now written only when someone runs rotate-all. templates/infra/README.md: unit
    row removed. manage-agents (skill and its ops reference): verbs and the post-reboot procedure.
    Template hygiene: one reference-deployment name that had survived in templates/infra/fleet-dormant.

Left private on purpose

  • install-host-services in the reference deployment installs an n8n-backup timer and resolves its
    workspace path from a fixed org; both are deployment-specific and stay out of the template.

v0.9.0 — A console is yours to build

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 10 Sep 10:59

The reference deployment built a tailnet-only web console for its fleet. Its code is deliberately
not in this repo
— the reasoning is in the new doc, and it is the same reasoning as app-layer.md:
a console encodes one operator's stack, and a root-privileged web app with a dependency tree would
contradict the "no extra always-on gateway to keep patched" line this repo leads with. What crosses
is the framework-level verb it needed, and everything that was learned building it.

Added

  • claude-topic list --json (templates/infra/bin/claude-topic). One JSON document for machine
    consumers: every status --porcelain field per topic, joined to the live process status the
    runtime reports (busy / idle / waiting and what it is waiting for), plus any session of that
    user with no topic unit under other_sessions. The join is the point: the runtime knows sessions
    by pid and by a display name it invents, which matches neither the topic key nor the registry
    name — the one fact both sides share is that /proc/<pid>/cgroup names the unit the process runs
    in. A missing or broken claude agents yields status: unknown plus a warning, never a silent
    idle; proved against three fixtures (binary absent, bogus pid, non-JSON output). other_sessions
    publishes both ids, because claude stop|rm take the runtime's short id and answer
    No job matching … if handed the session UUID — a real bug in the reference deployment's console,
    found in its own action log.
  • docs/console.md — the pattern, documented not templated: what a console is
    for, the identity model on a tailnet, the verb discipline (call the CLI you already have; a missing
    verb is a missing verb), background jobs for long commands, confirmations that state what is lost,
    and how to verify a page a headless browser is not allowed to reach. Listed in docs/README.md,
    and the README says plainly that the code is absent on purpose.

Security

  • Documented, with its reproduction: a tailnet identity header is not an authentication check on a
    multi-user box.
    Tailscale Serve adds Tailscale-User-Login and strips client copies, but it
    proxies to loopback — so any local process, including an unprivileged agent that ingests untrusted
    content, can present the same header and be indistinguishable from the operator. Ask the kernel
    instead: the process that owns the socket resolves the client's uid (/proc/net/tcp, or the
    Tailscale local API). No header is a refusal, never "local, therefore allowed".
  • And the trap that nearly shipped, because the lesson generalises. The Node adapter the console
    was built on emits two servers: the handler the custom wrapper imports, and the adapter's own
    standalone server. Both run the same identity hook; only one owns the socket. The standalone one
    binds 0.0.0.0 by default and forwards client headers untouched — so the uid the app read was
    whatever the caller typed. It sat unused in the installed tree on a host with a public IP; started
    by hand, an unprivileged local user with two forged headers got a 200. Nothing ran it, so nothing
    was breached, and the application code was correct throughout: the hole was a second path to it.
    The fixes are in the doc — a per-process token the wrapper mints and stamps, checked before
    anything else, plus not shipping the other entry point at all. The general rule: when a check
    depends on owning the socket, make it structurally impossible to reach the app without having done
    it.

Changed

  • De-identification and release-hygiene fixes the pre-tag checklist is supposed to catch, and had
    not.
    Four comments across templates/infra/bin/claude-topic,
    templates/infra/bin/claude-topic-session-hook and
    templates/infra/scripts/agentic-divergence-check still named the reference deployment's agents
    and its owner; placeholders now, per CONTRIBUTING ground rule 1. And 0.8.9 and 0.8.10 shipped
    without their CHANGELOG link definitions, so those two versions' links did not resolve — added,
    along with this one. The checklist itself is fine; it was run by eye. Both were found by running
    it mechanically this time, which is the only way it fires.

v0.8.10 — A budget you raise whenever it stings is not measuring anything

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 06 Sep 23:06

Changed

  • The always-on context budget is retired; the divergence check now asserts the runtime's own read
    limit on the auto-memory index.
    A total budget over CLAUDE.md + its @-imports + the index
    derived from nothing on disk: in the reference deployment it was set at 20,000 bytes, then 30,000,
    then 40,000 within three days, moved each time it stung, and it watched a number with no loss
    attached. Measured on the same box, it would have fired about eight times before the one limit that
    actually costs you something — the runtime loads only the first 200 lines or 25 KB of MEMORY.md
    at session start and silently drops the rest, so a memory past that point is already invisible
    to recall. Worse, its remedy said to shorten the index, which is the recall surface. Now:
    MEMIDX_MAX_BYTES / MEMIDX_MAX_LINES drift on breach, an [info] line past 80%, and the full
    total still prints every run because the trend is the signal. These are vendor constants, not
    derived ones — the one deliberate exception to
    decisions' derive-the-expectation rule — so the script pins the docs
    URL and the date they were read, and takes 25 KB as 25,000 so a stale reading fails towards firing
    early rather than towards silence. Proven against lowered thresholds before shipping: silent at the
    defaults, drift at 10,000 bytes and at 50 lines, [info] at 16,000.
  • agent-audit's memory pass gains the two growth steps, and the cross-tier checks from the
    reference deployment's 2026-09-03 pass.
    What keeps the index under the runtime's limit is this
    pass, with a human reading the diff — never a threshold, and never the session that happened to
    notice the number. The steps: retire dead episodic index lines (the line goes, the memory file
    stays, so a future session can still grep for it) and reduce curated bullets to verdict + pointer
    wherever the bullet already names the decision record holding its reasoning. On the reference agent
    13 of 79 index lines were dead episodic notes, and the curated file had grown 2.4× in fifteen days
    by carrying reasoning that a named record already held. The same pass also picks up the checks that
    had not been ported: whether the two tiers agree rather than merely not duplicating, and ordering
    present-state re-verification by the frontmatter modified field instead of file mtime — a bulk
    restore resets mtime for a whole store at once, and on the reference store 26 of 72 memories shared
    a single mtime minute.

Fixed

  • The date of the second over-trim of the reference deployment's curated memory file, in
    agent-audit: 2026-08-18, not 2026-08-22 — the 22nd was the restore. Caught by an advisor pass.

v0.8.9 — A restart that reports success can still have lost three days

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 06 Sep 23:06

The runtime mints sessionIds and re-keys them silently on /clear and on compaction. The wrapper
wrote topics.state only when one of its own verbs ran, so between a fork and the next invocation
the stored pointer was stale and nothing said so — and the next restart resumed the stale branch,
orphaned everything said since, and reported success, because by its own definition it had
succeeded
. In the reference deployment: six orphaning events across two agents, the largest
losing 8.5 MB and three days. One of them was caused by a self-restart armed specifically to
demonstrate that restarting was safe; the journal recorded nothing but Started.

Added

  • claude-topic-session-hook — a root-owned SessionStart hook, installed by
    install-host-services and registered in each brain's deploy/claude-settings.json. It writes
    the state at the instant the runtime creates an id, for startup|resume|clear|compact — exactly
    the set that mints them. Root-owned deliberately: an agent must not be able to rewrite the record
    of where its own history lives.
  • Bridge-derived identity in claude-topic. sid_newest_on_bridge and bridge_of_sid recover
    the live branch from transcripts on disk. The Remote Control bridge is the durable identity a
    sessionId is not, so same bridge = same topic. Derived from disk on purpose: it answers after a
    crash and at boot, when there is no live process to ask.
  • claude-topic run adopts. The unattended path systemd re-enters on every restart — where the
    damage actually happened — now adopts the newest branch on the bridge when the stored pointer is
    stale. Verified counterfactually against both real incidents: it returns the branch that was
    live, where the stored pointer returned one frozen for weeks.
  • claude-topic status --porcelain — field/value lines with a computed drift verdict
    (yes / no / unknown / unanchored) and newest_on_bridge.
  • The divergence check watches the hook, asserting both that it is installed and that the
    agent's settings register it, and reports pointer adoptions from the last 24 hours.

Fixed

  • rotate-all aborted at the first failing topic, on the boot path. The subshell the comment
    credited with preventing exactly this does not: under errexit, a failing ( cmd ); rc=$? kills
    the loop before rc is read. One bad topic left every other one announcing a dead bridge while
    active and enabled — the outage shape the verb exists to prevent.
  • rotate-all --dry-run named the wrong conversation. It printed the stored id while rotate
    abandoned the live one, so the preview was wrong on precisely the drifted topics that are the
    only reason to run it. Both now resolve through one helper.
  • rotate-all accepted only the exact string --dry-run and fell through to the destructive
    path for anything else — --dryrun would have blanked every conversation on the box.
  • Transcripts were ordered by the timestamp they declare about themselves, in a tree the agent
    can write. A two-line file claiming 2099 won outright and would have been exec'd. Ordering is
    now by mtime.
  • remove discarded the pointer with no record, where rotate refuses to discard one it
    cannot write down.
  • The hook wrote topics.state without the lock the wrapper uses for that same file.
  • The divergence check recovered the live id with a regex over prose, then continued when it
    found nothing. That regex returns empty — silently — for a stopped topic, for empty output, for
    a reworded line, and for a change of case.
  • An already-completed split was invisible. Comparing stored against live only ever sees a
    split still in progress: two topics reported CLEAN with orphaned branches sitting on disk.

Notes

Do not bind a hook to a systemd Environment=. The first version of this fix passed the topic
key as Environment=TOPIC_KEY=%i. It was wrong in both directions. Absent when needed: the units
were running from before the unit file changed, so twelve of twelve topics refused, and the
mechanism was inert for an afternoon while being reported as verified — the verification had run
against a throwaway topic created after the edit, with no Remote Control bridge at all, so it
exercised neither the unit path nor a bridge. Present when not wanted, which is worse: a systemd
Environment= is inherited by the whole process tree, so any nested claude fires SessionStart
carrying the topic's key and a foreign session id, and the hook writes that foreign id into the
topic's row. A binding that can be absent when you need it and present when you don't is not a
binding.
The key is now derived from the bridge in the session's own transcript, which cannot
name a session that is not on that bridge.

The hook runs detached. SessionStart fires before the runtime has written the transcript —
measured: a synchronous hook exhausted its budget at 18:12:21 and the transcript was created at
18:12:21. It lost by being early. A wait long enough to win would be paid by every ordinary
claude run that is not a topic.

mtime ordering is detection, not prevention. It defeats pre-planting; a file planted just
before a restart gets mtime "now" for free. That cannot be closed at that layer — the tree belongs
to the principal it would defend against. Every adoption is therefore recorded and reported.

This supersedes the claude-topic port in 0.8.8. That release shipped the first version of
this fix — bridge derivation plus a TOPIC_KEY-driven hook — from the reference deployment before
the deployment had finished reviewing it. Everything above was found afterwards, by an adversarial
review and by running the mechanism against deliberately broken fixtures rather than a clean one.
If you deployed 0.8.8, the Environment=TOPIC_KEY=%i line it added to claude-topic@.service is
the one thing to remove first.

A structurally better fix exists and was deliberately deferred: one working directory per topic
(WorkingDirectory=%h/.topics/%i). The runtime derives its project directory purely from cwd, so
the key would be readable from dirname(transcript_path) — no bridge, no parsing, no ambiguity —
and it would delete both derivation functions outright. It needs a transcript migration and it
moves where agents' relative paths resolve, so it is a decision to take calmly rather than at the
tail end of an incident fix.

v0.8.8 — Four fixes the reference deployment had already made and never shipped

Choose a tag to compare

@valiant-durin-bot valiant-durin-bot released this 30 Aug 17:47

Divergence between this repo and the deployment it was distilled from is the design. Divergence in the
things that must not diverge — the safety logic, and the skills that shape how an agent decides — is
rot, and it had set in unnoticed in four places. Found by hand, because the propagation check that
should have caught them compares only the infra/scripts pairs whose raises it can parse, and says so
now.

Fixed

  • claude-topic restart resumed the wrong conversation, silently. It read the STORED session id.
    That id is written at create/restart/remember and at no other time, while the running conversation's
    id can change underneath it — a compaction, a slash command, a new chat started in the app. From
    that moment the state is stale, and the next restart resumes the OLD branch while reporting success,
    because by its own definition it succeeded. In the reference deployment that meant a three-day jump
    backwards that nobody could see from inside. restart now reads the live id first, always, and
    adopts it when the two disagree; the stored branch stays on disk. agentic-divergence-check gains
    the matching daily assertion, so the drift is visible before anyone restarts.
  • fleet-brain-change prescribed a commit recipe that commits without pushing. The example fed the
    message by heredoc and then chained && git push. bash -c executes each complete command as it
    parses, so the commit runs and only the following line fails to parse — leaving the repo
    committed-but-unpushed, which is exactly what makes an ff-only sync skip it silently and forever.
    The recipe is fixed rather than annotated, and the file gained the Gotchas section that two places
    in it already pointed at and which did not exist
    .

Changed

  • decision-loop rewritten. The shipped version justified its central step with an unsourced
    "reacting beats answering open questions 3-4x". A research pass found no controlled study behind it,
    and found the adjacent evidence pointing the other way — a written proposal measurably pulls the
    reader toward it. The skill was breaking its own rule against numbers a reader cannot check, in its
    founding line. It now carries a verifiable four-part trigger with a negative list, a gated
    research fan-out (1–5 sub-agents, and none at all when no external source can change the answer, with
    the mandatory brief constraints in a new references/research-brief.md), a premortem in place of the
    deleted claim, an independent advisor pass, the choices put as structured questions, and the record
    written in the same turn.
  • The local-addendum instruction is now conditional. It asserted that every agent keeps one. An
    agent whose decisions all land in one shared place needs no addendum and should not grow one for
    symmetry: a thin skill overlapping an existing one degrades selection for both.
  • agent-audit's Memory step was one sentence; it is now four checks. Including the one that
    exists because a curated memory tier was over-trimmed twice by sessions following the old wording
    literally — prove a duplicate with two greps against both the mirror and the runtime store, cite
    file:line for each, or keep it.

Notes

docs/skills.md described decision-loop accurately enough to catch a regression in its own rewrite:
the falsify-assumptions-early step had been dropped. A downstream description found an upstream
defect, which is the argument for keeping such descriptions specific rather than vague.