Releases: mike-bronner/workbench-dev-team
Release list
0.51.1 — Scheduled runs in auto mode, agents clean their own scratch
🤖 Scheduled runs start in auto mode
dispatch-agent.shstarts each run with--permission-mode auto --permission-prompts noneinstead of--dangerously-skip-permissions. Auto mode's own rules call that flag an unsafe way to start an agent loop.- The Index and memory MCP tools are allowed by rule, so the classifier never judges a board write or a memory call.
- Every refused action is written to the run log as one
Permission denied:line, sub-agents included. A headless run exits 0 even when actions were refused, so the log is the only place to see them. - Commands with no rule, such as
gh pr create, are now judged by the classifier. Check the logs after the first runs.
🧹 Agents clean up their own scratch
- Watson, Holmes, the lens helpers, and Lestrade make scratch under the session scratchpad, or
~/Developer/scratchpadwhen none is named. They delete it themselves before they report. - A refused delete is retried with the literal path. It is never handed to you as a
!command. - Each scheduled run's folder now lives in
~/Developer/scratchpadand is deleted when the run ends, including on failure. Runs no longer leave folders in$TMPDIR. agents/lint-scratch-cleanup.shchecks that the rule stays in place.
⬆️ Upgrade
- Update the plugin.
- Run
/workbench-dev-team:setup. It reinstallsdispatch-agent.sh, so scheduled runs use auto mode only after this step. - After the first scheduled runs, check for refusals with
grep '^Permission denied: ' ~/.claude-workbench/dev-team-logs/*.log.
Folders left in $TMPDIR by earlier runs are not removed.
Full Changelog: 0.51.0...0.51.1
0.51.0 — Claude Code's own prompt approves commits
💥 Breaking changes
- The
approvecommand is gone. Claude Code's own permission prompt now asks before every foreground commit, push, and pull request merge. Re-run/workbench-dev-team:setupafter upgrading. Until you do, those commands run with no prompt. - Every brief needs an
Acceptance:slot. Watson, Holmes, and Lestrade refuse a five-slot brief and name the missing slot.Item IDandRepo sweepruns are unaffected, so the scheduled pipeline keeps working.
🔐 Commits and pushes: ask rules instead of a parser
The old gate parsed shell to decide what to prompt for. Every review found a spelling it missed, and its approval record could be written by any process.
- Setup installs ten
permissions.askrules forgit commit,git push, andgh pr merge, bare forms included. - The approval is still your "commit it" in chat after review. The prompt is the mechanical backstop, not a security boundary.
commit-guard.shrefuses what the rules cannot see: a sub-agent's commit or push, any merge by a sub-agent or the pipeline, a force or delete push, and a commit or push behindbash -c,env,eval, aNAME=valueprefix, or a program path.
🤖 The pipeline stays in its own folders
pipeline-scope.shanswers the scheduled pipeline's prompts. It allows one plaingit -C,rm, orrmdirline whose paths and repository stay inside$TMPDIRand the scratch roots.- A push must send one local branch that is not the clone's default. Pull request merges are never allowed.
- Each run starts in a fresh
mktemp -dfolder, so this plugin's own repo is never a run's project folder. - The 24 tools pipeline runs used to lose to this repo's local settings are denied with
--disallowedTools.
🕵️ Reviews
- Holmes's helper reviewers run on a new read-only agent type,
holmes-lens. - The review guard is now a static rule: Holmes and
holmes-lensmay not write outside the scratch roots. The per-review hold, its records, and its two-hour expiry are gone. - Holmes's Local mode reviews against the brief's Acceptance list.
/developgrades its three options against the same list and points at/workbench-core:intake.
🧰 Setup
- Every setup block now passes workbench-core's destructive-scope guard, so setup runs from a session.
- Setup prints
! rmlines for the old approval script, the old approval records, and the old review-guard records. Nothing is deleted for you.
⬆️ Upgrade
- Update the plugin.
- Run
/workbench-dev-team:setup. - Run the
! rmlines setup prints. - After each Claude Code upgrade, re-run the manual pipeline check in the README ("Re-check after a Claude Code upgrade").
Release workbench-core's Acceptance-slot change only after this one.
Full Changelog: 0.50.1...0.51.0
0.50.1
Fixed
- 🔒️ Close the Opus 5.5 self-audit gaps and the gate's false positives by @mikebronner in #47
Full Changelog: 0.50.0...0.50.1
0.50.0 — A foreground push asks first
🚦 A foreground push now asks first
A foreground git push ran with no prompt at all. The prose also called push the human's own, so agents stopped short and asked Mike to push.
The gate now prompts for a push on the same route as a commit. The human approves one exact command, in one directory, and it runs once.
🧱 The gate stopped parsing shell
Two review rounds tried to teach the gate more bash. Each round found new shell that it read wrong. The worst case was two apostrophes in # comments pairing up and hiding a commit and a push.
So the gate now sorts every command into three classes.
| Class | Example | Result |
|---|---|---|
| Plain form | git commit -m "...", git push origin main |
🔐 prompted |
| Could hide one | a comment, $(...), a loop, git pull && git push |
🛑 refused, with a request for the plain form |
| Everything else | git status, git log, gh pr view |
✅ passes |
The refusal test is a case-insensitive substring match with no parsing. No quote, comment, or line break can hide a verb from it.
🔒 What an approval is bound to
- A commit is bound to HEAD and the staged diff. It also covers the working tree when the commit takes files from there.
- A push is bound to its branches, tags, HEAD, and remote, branch, push and url config.
- Any change before the approved command runs voids the approval.
Force pushes and pushes that delete remote refs are refused outright, in every spelling git accepts.
📝 Also in this release
- The warmup and the develop, orchestrate and git-commit skills teach the plain form. Stage with its own
git add, and put a multi-line message in a scratchpad file for-F. - Sub-agents stay refused on every hidden shape, and the pipeline stays exempt.
- Shell aliases such as
gpstay outside the gate, and the README says so.
0.49.0 — Every agent pinned to Opus 5.5 at medium
⚠️ Your existing config does not change on its own
Setup never rewrites ~/.claude-workbench/dev-team-config.json without your say-so. After updating, re-run /workbench-dev-team:setup.
It compares each agent's model and effort against the new pin. For every agent that differs, it shows the current value and the replacement, then asks. Until you answer Replace, dispatch keeps using your old values, because config flags beat agent frontmatter.
📌 What changed
All three agents now ship the same exact pin:
| Agent | Before | Now |
|---|---|---|
| Lestrade | sonnet / high |
claude-opus-5-5[1m] / medium |
| Holmes | opus / high |
claude-opus-5-5[1m] / medium |
| Watson | opus / session effort |
claude-opus-5-5[1m] / medium |
lensModel, fanout, fallback, and maxBudgetUsd are unchanged.
🧭 Why
- Exact ID, not the
opusalias. The alias moves to the next Opus release without anyone approving it. [1m]. It keeps the 1M context window the agents budget their working context against.mediumfor all three, Holmes included. Anthropic publishes Opus 5.5 guidance on this. Atmedium, Opus 5.5 beats Opus 5 athighon coding and code review. At the same level, Opus 5.5 also thinks more per turn than Opus 5.
Watch Holmes's bounce and escalation rate after the switch. Raise it to high only if either gets worse.
🔧 How the pin reaches both dispatch paths
- Interactive. The Agent tool's
modelparameter accepts aliases only, and it overrides frontmatter. So setup's Step 6a now stampsmodelinto agent frontmatter alongsideeffort, and the orchestrate skill never passes that parameter. - Scheduled.
bin/dispatch-agent.shpasses--modeland--effortonly when the config names them. It keeps no baked-in model default. - Deliberate opt-outs.
CLAUDE_CODE_SUBAGENT_MODELwithCLAUDE_CODE_SUBAGENT_MODEL_FORCE=1beats the frontmatter model.CLAUDE_CODE_EFFORT_LEVELbeats--efforton the scheduled path.
🛡️ Hardening found in review
Two Holmes local reviews found six blockers, and all six are fixed:
- The model validator accepted a value spanning several lines. It now refuses any newline.
- A wrong-shaped config entry (for example
"watson": "opus") is reported for a manual fix, never offered for replacement. - A failed config write, including a failed
mv, now exits 1 and leaves the file byte-for-byte unchanged.
✅ Tests
New commands/test-config-pin.sh has 28 cases. agents/test-effort-stamp.sh and bin/test-dispatch-agent.sh now enforce the pin for all three agents. Every new guard was broken on purpose, and its test went red. The full suite passes.
0.48.0 — The review guard stops missing perl -pi
🔒 The spelling that got through
hooks/scripts/local-review-guard.sh stops a Holmes review lens writing to the tree it is reviewing. It matched -i as a whole token only.
So perl -i -pe was refused and perl -pi -e was allowed — the commonest way anyone writes an in-place Perl edit.
Five more clustered spellings went with it: perl -ni, perl -pi.bak, perl -lpi, sed -ie, and ruby -pi. Ruby was not in the interpreter table at all.
🔬 How it surfaced
Not by reading the code. During a Holmes Local-mode review of an uncommitted workbench-core tree, a correctness lens ran perl -pi against assets/permissions/rails.json to test a mutation, despite an explicit no-write instruction in its own prompt. The guard allowed it.
The same lens's later perl -i attempt on another file was correctly refused. That inconsistency is what made the hole visible. The lens self-restored and the tree was verified intact, so nothing was lost.
The 79-case suite covered no clustered form, which is why this shipped.
✨ The fix
The flag rule now walks each single-dash cluster left to right, stopping at the first switch that takes the rest of the token as its value. So perl -pes/i/j/ stays a read-only one-liner rather than becoming a false positive.
Which letter means in-place is read per interpreter, from the same table as those terminators, because the interpreters disagree:
-Iis an in-place edit to BSD and macOSsed, and is refused there. Toperlandrubyit names an include directory, soperl -Ilib -ne printstays allowed.sed's-lstays out of the terminator set. GNU's-l Ntakes a line length and BSD's takes nothing, and listing it let a joinedsed -lithrough.
Both halves come from one subscript, so an interpreter cannot be present in one lookup and missing from the other. That drift would raise KeyError, crash the hook, and a crashed hook fails open.
✅ Verified from outside
The session that reported the hole re-ran its original matrix against the fix rather than taking it on trust.
All eight in-place forms are now refused, including the six that used to pass. Zero false positives across perl -Ilib -ne print, perl -pes/i/j/, perl -ne print, sed -n 1,5p, git diff HEAD, and grep -i foo.
hooks/scripts/test-local-review-guard.sh goes from 79 to 125 cases. Every test and lint script in the repo passes.
⚠️ This refuses commands that previously ran
A minor bump rather than a patch, because the behaviour change is user-visible. Six command shapes a review agent could run before are now blocked. That is the point of the release, and it is worth knowing before you update.
0.47.0 — Approve a commit whose message lives in a file
⚡ After you update
Restart any session that should be governed. A running session holds the plugin it loaded at start, so an open conversation keeps the old hooks until it restarts.
If you want Watson to stop carrying an effort default, edit your own config. Setup never overwrites ~/.claude-workbench/dev-team-config.json, which is deliberate — your edits survive every update. The consequence here is that an existing install keeps whatever watson.effort it already had, and the next setup run stamps that value straight back into the frontmatter. Remove the key yourself, then re-run /workbench-dev-team:setup.
✨ Approve a commit whose message lives in a file
The approval gate refused git commit -F <path>, which is the shape it should have preferred. The gate makes the command run twice, so a message written inline lands in your transcript twice in full.
The subject check looked for the subject in the command text, and a -F command carries only a path. It now reads the first line of the file git will read, in every spelling git accepts — -F <path>, --file=<path>, --file <path>, and short clusters like -aF. Variables in the path resolve from the command's own assignments and from the environment, never by running a shell on it: the recorded command is caller-controlled text, so expanding it through a shell would be the injection this gate exists to prevent.
The file is display. The command is still the decision. Every check that decides whether an approval may be granted finishes before the first byte of a message file is read. A file can refuse an approval or name a commit. It can never widen one, change which command an approval covers, or stand in for an approval that was never granted.
It fails closed. A path that will not resolve, a file that will not open, an empty message, and -F - each refuse and name their own case. A command that promised no subject at all, such as --amend --no-edit, falls back to the request id as before.
🔥 No shipped effort default for Watson
Watson carried effort: xhigh in the shipped config from 2026-06-12, and in its frontmatter from 2026-09-11, when the setup stamper first made the interactive dispatch path read frontmatter. Nobody read either line for months.
Correcting the number would have fixed one value and left the next wrong one just as invisible, so the key goes instead. A value that is never shipped cannot drift.
The config key and the frontmatter line had to go together. The stamper rewrites frontmatter from the config on every setup run, and setup runs at install, so removing the line alone would have been cosmetic.
The consequence, stated rather than buried: Dispatch stops passing --effort for Watson too. Per-agent effort control for Watson is gone on both paths, not only the interactive one. Watson inherits the session's effort instead.
Holmes and Lestrade keep high. The same drift argument applies to them identically, but between them they ran three dispatches in fourteen days, so changing them buys nothing against a review- or triage-quality risk. The asymmetry is deliberate and each document now says so.
📐 An advisory working-context budget on every agent
All three agent definitions now carry a target of roughly 250k tokens of working context for a single task.
Nothing enforces that figure, and nothing in the harness can. maxBudgetUsd is passed on the scheduled dispatch path and reaches no other, and the Agent tool that spawns an agent from a live conversation exposes no budget parameter at all. The wording says so outright, because a limit presented as enforced when it is not gets trusted and then silently exceeded.
The five-slot brief template stays free of figures. The two numbers measure opposite sides of a handoff: a brief-length ceiling would bound the prose a sender writes, while this bounds the context a receiver accumulates. A shorter brief does not buy a cheaper run, which is why the brief-length figure was removed in the first place and stays removed.
0.46.0 — Holmes reviews work that never reached the board
⚡ After you update
Restart any session that should be governed. A running session holds the plugin it loaded at start. An open conversation keeps the old hooks until it restarts.
No setup re-run is needed. This release changes agent prompts, skills, and hooks. All three are read live. 0.44.0's requirement to run /workbench-dev-team:setup still stands if you have not run it.
✨ Holmes reviews work that never reached the board
Holmes accepted Item ID: <n> and nothing else. Every review therefore needed a board item behind it.
Watson has had no such coupling. His Direct mode takes a prose brief. It runs on any local repo.
That asymmetry had a cost. Watson's Direct mode hands work back as an uncommitted working tree. That tree had no review path. You reviewed it yourself, or it shipped unreviewed.
Local mode is now Holmes's default. Hand him a five-slot brief. He reviews the uncommitted working tree in its Workdir:. Index mode is entered only on an explicit item-ID token.
| Local mode | Index mode | |
|---|---|---|
| Entered by | a prose brief (the default) | Item ID: <n> |
| Reviews | the uncommitted working tree | the PR on the board |
| Rubric | the brief's Goal: and Done when: |
the issue's acceptance criteria |
| Tests | runs the repo's own suite | reads CI status |
| Verdict | prose, to the dispatching session | an App-signed review |
| Writes | one vault note | the board and the PR |
Ambiguous prose resolves to Local mode. It never resolves to Index mode. The two mistakes cost different amounts. A misread brief wastes a report. A guessed id posts a signed verdict onto somebody else's PR.
The rubric binds exactly as acceptance criteria bind. Goal: and Done when: go verbatim into the lens prompts. Holmes never amends them. A rubric that is itself wrong comes back as a dispute with three options.
Local mode makes no The Index call and no GitHub write. There is no board item to move. A local review carries no consent to post under your identity.
The lens fan-out, the adversarial verification, the memory pass, and the finding-routing matrix are unchanged.
🔒 A guard, because the prose did not hold
Local mode reads your live working directory. There the uncommitted change is the only copy of the work.
The rule protecting that tree started as prose at five sites. It sat verbatim in every sub-agent prompt Local mode dispatches.
On the mode's first real exercise, a lens sub-agent ran chmod against that directory. It changed a script from 755 to 644. It disclosed the breach itself. The prohibition had reached it and it acted anyway.
Prose in an agent prompt is advisory. A second PreToolUse hook now enforces the rule. It sits beside the commit gate rather than inside it.
| Event | Behaviour |
|---|---|
| A Holmes dispatch carrying a prose brief | arms a record for that session |
| A Bash call from a sub-agent of an armed session | denied when it mutates a tree |
| The dispatch returns | releases the hold |
Nothing asks the reviewed agent to arm anything. The hook reads the Agent tool's own subagent type and prompt. It applies Holmes's own mode detection. The prose that drifts is not in the loop.
The signal is the harness-supplied session_id plus a non-empty agent_id. Your own window keeps working. A concurrent session is untouched. A host-wide marker would be the watson.lock leak with the sign reversed.
What it refuses is a class, not a roster. Git's reading verbs are enumerated. Every other git verb is refused. That is how it sees git restore, which the old prose list never named. File metadata joins it, since that is what the breach used.
Reads and the repository's own suite stay legal. That is the binding constraint. A guard that stops either makes the mode useless.
What it costs you
While a review is armed, every sub-agent of that session is read-only. That includes ones unrelated to the review. It is bounded by the release and a two-hour expiry.
A review that was never armed fails open silently. The hook is a backstop. The prose stays.
🔍 The approval prompt now names the commit
The commit-approval gate denied a commit. It said nothing about the description on the approval command. Each session therefore worded that line itself. One wrote "Request human approval for this commit". That names the action and never the commit.
A prompt carrying no decision content gets cleared unread. That habit reaches the other rules on the same ask list. git reset --hard and gh pr merge are among them.
The denial now dictates the line rather than describing an intent. It names the parameter, and gives the literal shape Commit: <first line of the commit message>. It stops at the subject. Nothing about what the gate approves has changed.
🧪 Tests
| Suite | Cases |
|---|---|
hooks/scripts/test-local-review-guard.sh |
79, new |
hooks/scripts/test-commit-approval-gate.sh |
84, up from 83 |
agents/lint-holmes-local-mode.sh |
8 checks across four files, new |
The lint keeps no copy of the forbidden-verb rule. It feeds the reference's fenced blocks to the guard's own classifier. The documentation and the enforcement therefore cannot drift apart.
0.45.0 — A sub-agent gets no commit, no merge, and no push
⚡ After you update
Restart any session that should be governed. A running session holds the plugin it loaded at start. An open conversation therefore keeps the old gate until it restarts.
No setup re-run is needed for this release. 0.44.0's requirement to run /workbench-dev-team:setup still stands if you have not run it.
🔴 The bug
0.44.0 made the gate deny, and routed approval through a permissions.ask rule. The reasoning: a permission rule is evaluated before the classifier in every mode. A rule therefore forces a real prompt.
That holds in a foreground session. It fails in a background sub-agent.
The sub-agent's request is background and non-interactive. The ask therefore resolves with nobody attached. The approval command is a Bash command the agent itself holds. The agent became both the requester and the approver.
Measured across every session transcript on one machine, during the gate's first day:
| Measurement | Value |
|---|---|
| Sub-agent self-approvals | 48 |
| Median delay, denial to approval | 3.4 seconds |
Unapproved commits on main from a single run |
17 |
Three seconds is not a human reading a diff.
✅ The fix
Three lanes, decided in this order.
| Lane | Signal | Verdict |
|---|---|---|
| Scheduled Index pipeline | WORKBENCH_DEV_TEAM_PIPELINE=1 |
Silent. Commits unattended, as before |
| Any sub-agent | Non-empty agent_id |
Refused, with no approval path offered |
| Foreground session | Empty agent_id |
Unchanged from 0.44.0 |
A sub-agent's denial prints no request id and names no command. It writes no pending record. The lane is decided before any record is read. An approval planted by hand therefore buys nothing either.
There is deliberately nothing for the agent to run. Any command an agent can run is not an approval.
The lane signal is the payload's agent_id. The harness supplies it, rather than the command. agent_type cannot serve. It is present for a scheduled run and an interactively dispatched one alike.
What is refused
A rule, not a list:
| Category | Verbs |
|---|---|
| Writes a commit | commit, revert, cherry-pick, am |
| Integrates another history | merge, rebase, pull |
| Publishes one | push, plus gh pr merge |
Matching the commit verb alone is what left merge and push open.
Reads stay open. A sub-agent can still fetch, diff, log, and view a pull request.
What a refused sub-agent does
It hands the work back. An uncommitted tree, the diff, and the proposed message go to the dispatching session. That session commits it, where a prompt reaches a human.
Watson's Direct mode, the develop standard, the orchestration guidance, and the README all say so.
One routing change
An Index item dispatched from a conversation now goes through bin/dispatch-agent.sh. That is the only path which can set the pipeline flag. The Agent tool keeps Direct mode.
🧪 Tests
83 gate cases and 31 approval cases. Every new guard is mutation-verified. Each one was broken in a sandbox copy and confirmed to redden. The control stayed green.
⚠️ Honest limits
- Still a protocol gate, not anti-evasion machinery. The gate reads the command the agent asked to run. A verb hidden inside
bash -cis not its subject. Neither is one written into a script. The observed failure was an agent following printed instructions. Removing the instruction removes that failure. - A commit message that discusses git commands can trip the gate. The command text includes heredoc bodies. A message line beginning with a gated verb is read as a command. Reflow the line.
git pullis refused, because it merges. A Direct-mode sub-agent cannot update its checkout.git fetchstays open, so it can still read remote state.
0.44.0
The commit approval gate now actually stops a commit
The bug
The gate returned permissionDecision: "ask". A hook's ask is classifier-approvable, so under permissions.defaultMode: "auto" the auto-mode classifier answered it and no human was ever prompted. The gate has never stopped a commit.
Verified against the Claude Code 2.1.268 binary: only deny binds from a hook, and the hook output schema exposes no way to mark a decision unapprovable. The harness's own safety checks use an internal field that hooks cannot set.
The fix
The gate denies, and the denial is an instruction. It tells the agent to show the human the staged diff and the proposed message, then run a specific approval command. That command is covered by two permissions.ask rules, and rules are evaluated before the classifier in every mode, so running it forces a real prompt. The human's answer is the approval.
One approval covers one commit. The request id hashes the session id, the agent id, and the exact command text, so two agents in one session can never share an approval. The commit that uses a record deletes it. An unused record expires in 15 minutes.
Every failure of the machinery refuses the commit: missing rules, an unwritable or undeletable record, no session id, or a copy of the script the rules do not name.
The scheduled pipeline is unchanged
The WORKBENCH_DEV_TEAM_PIPELINE=1 carve-out still runs first, before the payload is read, and is still per-process. A test proves the scheduled lane commits with no session id and no writable state directory.
Upgrade note, required
Run /workbench-dev-team:setup immediately after updating. Until it runs, the approval script is not at its stable path and its permission rules are absent, so the gate denies every interactive commit with no way to clear it. That is fail-closed by design, and it is loud rather than silent: both the denial and the script name setup as the remedy.
Honest limits
- This is a protocol gate, not a barrier against a hostile agent. Anything holding Bash can write an approval record. The guarantee is narrower and worth stating plainly: an unapproved commit cannot happen silently.
- An invocation spelled differently from the two rules, such as
sh <path>or the path withoutbash, matches no rule and prompts nobody. The denial prints the exact command for that reason. A process-inspection check was rejected as keying the gate to an undocumented harness detail.
Tests
48 gate cases, up from 26. 27 new approval-script cases. 19 new setup cases. 19 mutants, all of which redden.