Work since 1.3.0 (epic #49): supervision is enforced rather than described, the
credential scan covers what the worker actually receives, configuration refuses bad
values instead of silently defaulting, and prose counts are held against the code by a
new offline check.
Follow-ups (epic #58)
- Interruption and scanning (Refs #52). A signal after the worker returned but
before the round was saved un-counted the round and left a stale in-flight marker; an
ordinary error was recorded as an interruption; a slowgit ls-filescrashed task and
fleet; read-onlymuse-asknever scanned for credentials. Each is fixed and
reproduced offline, and a scan that cannot list files now refuses. - Harvest excludes follow git's glob rules (Refs #53):
*and?stay inside one
path segment, so--exclude 'a/*.py'no longer dropsa/b/c.py. - The supervisor's documented commands pass its own guard (Refs #54), the block
reason prints a runnable finish command, and a recorded check is allowed only inside
the owning task's worktree. - Ten more mutants killed (Refs #55) and tests that could not fail now can
(Refs #56).mutate.shderives its per-mutant timeout from a timed control run and
drops to one job under load, so a busy host no longer reports false timeouts. model:frontmatter measured per permission mode (Refs #57): it takes effect in
default, acceptEdits and bypassPermissions, and auto and plan keep the session model.- The eval judge is replayed byte-for-byte (Refs #50) and the README records
measured vote splits instead of the best run. On a macOS host the eval sandbox blocks
the Command Line Tools, so cases that need git or python3 are measured on Linux
(Refs #59).
Added
- A PreToolUse guard denies a supervisor's writes (Refs #43). Omitting
Writeand
Editfromallowed-toolsonly left them prompted, and the skill pre-approved broad
Bashgrants — so the supervisor keptBashand a hook (hooks/supervisor_guard.py,
matcherBash|Write|Edit|NotebookEdit) now denies its Write/Edit/NotebookEdit and any
Bash beyond muse shims, read-only git, readers and the recorded check. The skill
pre-approves only read-only status/doctor shims and readers. - A SubagentStop block plus a PostToolUse census for the central claim (Refs #40).
SessionEnd output is discarded, so leftover worktrees were reported to nobody, and the
stop check never matched the supervisor's agent type. SubagentStop now matches
^muse:muse-supervisor$and blocks (at most twice per task) a supervisor that stops
with no verdict, PostToolUse onAgenttells the orchestrator where the on-disk
artifacts disagree with the summary, SessionEnd is removed, and the SessionStart
preflight names recorded muse/ and fleet/ worktrees still open. - Fleet supervisors run as muse-supervisor, and unverified tasks are counted from disk
(Refs #44). The workflow spawned supervisors without the agent type, trusted each
supervisor's self-reported verified flag, and interpolated briefs into double-quoted
shell; the census now reads eachtask.json, and briefs and checks travel via files. - Shims on PATH with narrow grants (Refs #42).
bin/shipsmuse-ask,
muse-cleanup,muse-doctor,muse-fleet,muse-model,muse-statusand
muse-task; docs and grants name them bare. - An offline stub round trip against a moving base branch (Refs #32), so diffing
against a ref name can no longer silently stand in for the pinned sha. scripts/mutate.shplants each known bug and fails when a must-kill one survives
(Refs #32).- A frontmatter checker with a probe self-test, and a pinned CI toolchain
(Refs #34).claude plugin validate --strictcatches YAML parse errors only
(measured on 2.1.280), soscripts/check_frontmatter.pyholds keys and value types;
CI installs the CLI atMIN_CLAUDE_VERSION, pins actions by SHA, and runs shellcheck
at severity=error. - A doc-claims check holds prose counts against the code (Refs #47).
tests/test_doc_claims.shfails when a command, shim, auto-triggering surface, flag
or hook event is added without updating the prose — each with a probe proving the
failure names the planted defect. - A monitor and a status line for supervised tasks (Refs #48).
muse-taskand
muse-fleetappend round, verify and verdict events; a plugin monitor streams each as
one notification line, andsettings.jsonsets asubagentStatusLineshowing the
task, round N/M and the last check.claude plugin detailsdoes not inventory monitors
or settings on 2.1.280, sotests/claude_validate_monitors.shis the proof they load. - The evals run for real (Refs #33). Every case is a
case.yamlwhose scaffold builds
its fixture repo and puts a stub muse on PATH, so nothing leaves the machine; cases get
the tools their behaviour needs (Workflow is gated on 2.1.280) andevals/README.md
records measured scores. All nine triggering graders pass; one behaviour judge
(trust-the-self-report) is still inconsistent.
Changed
- userConfig refuses bad values and reaches every driver (Refs #41). Unparseable
values were silently replaced by defaults, and some keys never reached fleet, ask or
the workflow. Every driver now refuses before spawning, all five keys reach task,
fleet, ask and the supervised workflow as flags, worktree roots resolve against the
repo and refuse one inside it, and doctor reports each value with its source. Each
script resolves flag first, thenCLAUDE_PLUGIN_OPTION_<KEY>, then the default, and
refuses a set-but-invalid value rather than falling back. - The credential scan covers exactly what the worker receives (Refs #39). It missed
dotfiles, used narrow patterns, and scanned a different directory than the worker
started in; one scan now coversgit ls-filesplus seeded and linked paths and the
gitignored files of the worker's start directory, and all three drivers honour it. - The token diet (Refs #47). The supervisor agent's description carried three
<example>blocks (~600 always-on tokens) and invited auto-triggering, which the docs
never intended. It is now a two-line description; the skill description is trimmed to
about half with every trigger cue and exclusion kept. Always-on falls from ~1,202 to
~625 tokens, against the 750 ceiling (the skill description regained the trigger cues
two positive eval cases had stopped firing on). - The Windows Git Bash leg is blocking (Refs #32). It runs the same free offline
suite; the tests stub their host dependencies instead of asserting the runner's
health, strip CR from Python output, resolve executables throughshutil.which, and
kill the whole tree on timeout.
Fixed
- Worktree safety: cleanup, harvest, resolution, collisions (Refs #36, Refs #35,
Refs #45, Refs #46). Cleanup could delete the repo, human worktrees and unharvested
work; harvest dropped tracked build files and binaries and mishandled nested
excludes;--outand the repo resolved against the cwd instead of the owning
repository; fleet never checked branch or patch collisions. Each now refuses or
resolves correctly, with an offline reproduction ending in the data intact. - Robustness: bytes, signals and round-in-flight (Refs #37, Refs #38). Non-UTF-8
worker or check output raised after the round ran but before it was recorded, so the
round never counted; a signal left the worker running under a reaped worktree; a
second round could start on the same task. Output is bytes-safe with the round
recorded before fingerprinting, signals kill the process tree, and a marker answers
round_in_flightwhile the worker lives. - Test and CI integrity: count guard, mutate.sh and the frontmatter checker refuse
what they let through (Refs #32). A misspelled mutate id ran only the control and
exited 0, totals were hard-coded, a hung mutant had no timeout and a crash counted as
killed; the count guard counted skips; the stamp guard grepped instead of parsing.
All eighteen known mutants are now killed andMUST_KILLholds every one. - An unset userConfig key no longer blocks delegation (Refs #41). Claude Code
substitutes${user_config.KEY}only for keys the user set and leaves the rest
literal; the prose double-quoted them, so the shell failed with "bad substitution"
before any script ran. Placeholders are single-quoted, and a literal placeholder for a
flag's own key counts as not given, so the default applies. Found by a real
no-config end-to-end delegation, which now completes.
Corrected
Claims in 1.3.0 that were false, corrected here rather than silently edited there:
- The SessionEnd hook reported open worktrees to nobody — SessionEnd output is
discarded — and is removed in favour of the SessionStart leftover report (Refs #40). - SubagentStop plain output never reached the orchestrator; the hook now blocks, and
the PostToolUse census carries the report (Refs #40). - The 1.3.0 note that the validator rejects
optionson userConfig does not hold on
2.1.280, which accepts it;plugin.jsonnow usesoptionsondefault_effort
(Refs #41). model: haikuon the reporting commands (Refs #29, Refs #57) is now measured on
2.1.280: headless runs with--model sonnetserve/muse:statusfrom
claude-haiku-4-5-20251001 in default permission mode, and amodel: haiku
skill control is likewise served by claude-haiku-4-5-20251001 in acceptEdits
and bypassPermissions modes, while auto and plan keep the session model with
the debug warning that names skills and commands together as one mechanism.
The frontmatter stays, and the runs are recorded in references/field-notes.md.SKILL.mdsaid six commands and lists seven; the Running-it table and the details
section pointed atreferences/workflow.mdas the workflow script, which that file
says it is not;--outwas said to resolve against the agent's cwd, when it resolves
against the owning repository; and the docs said one auto-triggering surface where
there are three — one skill, the supervisor agent, the registered workflow
(Refs #47).