Skip to content

feat(review-toolkit): add OpenAI Codex as a cross-vendor orchestrator review surface#177

Merged
kyle-sexton merged 5 commits into
mainfrom
feat/codex-orchestrator-review-surface
Jul 14, 2026
Merged

feat(review-toolkit): add OpenAI Codex as a cross-vendor orchestrator review surface#177
kyle-sexton merged 5 commits into
mainfrom
feat/codex-orchestrator-review-surface

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

Problem

code-review-fanout's ## Orchestrator plugins set was a closed pair of same-vendor Claude plugins — pr-review-toolkit and code-review. Combined with the plugin's own Claude leaf agents, every finding-producing surface in the fan-out ran on one vendor's model, so they share one vendor's blind spots. A breadth-review skill whose surfaces are all correlated is narrower than it looks.

Change

Add the OpenAI Codex plugin (codex@openai-codex) as a third orchestrator surface:

  • /codex:review — read-only, so it satisfies the review modes' report-only contract with no PR-mutation gate (unlike code-review's PR mode).
  • /codex:adversarial-review — red-teams the diff, fitting the section's "intentional adversarial breadth" intent.

It is the first surface backed by a different model — its blind spots are uncorrelated with the same-vendor leaf agents and Claude orchestrators, so a finding only Codex raises is signal the rest structurally cannot see.

Also adds the matching main-thread invocation line to run-everything-mode.md step 2 — the exhaustive sweep is where the different-model angle earns its cost. Kept a graceful enhancement (absent → note the skip), not a hard dependency, consistent with the other two.

Verification

  • markdownlint-cli2 on both changed files — 0 errors
  • Invocation names match the enabled plugin (codex@openai-codex; /codex:review, /codex:adversarial-review)

🤖 Generated with Claude Code


Note

Low Risk
Documentation-only changes to review skill orchestration; no runtime code, auth, or data paths.

Overview
Extends code-review-fanout from two Claude-only orchestrators to a third optional surface: the OpenAI Codex plugin (codex@openai-codex), documented as the first different-model reviewer so its blind spots are uncorrelated with same-vendor leaves and orchestrators.

SKILL.md now lists Codex alongside pr-review-toolkit and code-review, with /codex:review (read-only, no PR-mutation gate) and /codex:adversarial-review (red-team). Invocation contract is spelled out: --wait so normalization sees foreground output, and --base <review-base> aligned with the skill’s shared review diff base.

run-everything-mode reorders the flow so the review diff base is resolved before orchestrators (step 2), then documents explicit Codex calls in the orchestrator step, including the committed-only vs dirty-tree coverage gap and when to call that out in ## Surfaces. Step numbers and the diffability pre-check move accordingly (now before leaf fan-out at step 5).

Reviewed by Cursor Bugbot for commit b81b15b. Bugbot is set up for automated code reviews on this repo. Configure here.

… review surface

The code-review-fanout "Orchestrator plugins" set was a closed pair of
same-vendor Claude plugins (pr-review-toolkit, code-review) — every
finding-producing surface in the fan-out shared one vendor's model and
therefore one vendor's blind spots.

Add the OpenAI Codex plugin (codex@openai-codex) as a third orchestrator
surface: /codex:review (read-only, so it satisfies the review modes'
report-only contract with no PR-mutation gate) and /codex:adversarial-review
(red-teams the diff, fitting the intentional adversarial-breadth clause).
It is the first surface backed by a DIFFERENT model, so its blind spots are
uncorrelated with the same-vendor leaves and Claude orchestrators.

Add the matching main-thread invocation to run-everything mode, where the
exhaustive sweep makes the different-model angle worth its cost. Kept a
graceful enhancement, not a hard dependency.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Jul 14, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_29d703f4-6180-4af4-8167-a4905a837c57)

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale comment

Risk: low. Not approving because Cursor Bugbot did not complete successfully (skipped due to usage limit), so the required automated review signal is missing. Human review is needed before merge; no reviewers were assigned.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Router and Approver

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 95afb68185

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/review-toolkit/skills/code-review-fanout/context/run-everything-mode.md Outdated
…findings

Codex review on the PR: /codex:review and /codex:adversarial-review with
no flag prompt the user and may run in a Claude background task that
returns only a status handle (verified against openai/codex-plugin-cc's
command docs). The fan-out's step 5 gathers raw orchestrator outputs and
normalizes them in the same turn, so a backgrounded Codex run would
contribute an empty surface and be silently dropped from the report.

Invoke both commands with --wait so they run in the foreground and return
findings synchronously, and document why the flag is load-bearing.
@cursor

cursor Bot commented Jul 14, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_d4f4383d-71de-4b85-b0c3-895d22c953b4)

@kyle-sexton

Copy link
Copy Markdown
Contributor Author

@codex review

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale comment

Risk: low. Not approving because Cursor Bugbot skipped (usage limit reached), so the required automated review signal is missing. Human review is needed; no non-author reviewers were available to assign.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Router and Approver

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e7fff61205

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/review-toolkit/skills/code-review-fanout/context/run-everything-mode.md Outdated
Codex review: run-everything invoked /codex:review in step 2, before the shared
review diff base is resolved in step 3, so on a PR whose base is not the default
branch Codex would auto-pick the working tree or default branch and review a
different diff than every other surface. Pass --base <review-base> (verified as a
real flag in openai/codex-plugin-cc's command docs) with the resolved base, and
require resolving that base before the call.
@cursor

cursor Bot commented Jul 14, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_73030b92-45ec-4fe6-9a93-a8e8e537c0ff)

@claude

claude Bot commented Jul 14, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 1m 48s —— View job


PR Review

  • Gather context from PR description and comments
  • Read changed files
  • Review against CLAUDE.md guidelines and correctness
  • Post findings

Summary

Documentation-only PR. The three-commit history (feat → --wait fix → --base fix) shows good responsiveness to the two Codex review rounds, and both load-bearing flags are now in place. The cross-vendor framing, PR-mutation gate analysis, and graceful-absent pattern are all correct. Two small issues below.


Finding 1 — Garbled sentence in SKILL.md (minor, correctness)

Line 70:

"Invoke it with --wait so the review runs in the foreground and returns findings in the same turn — its default prompts and may run in a background task the synchronous normalization step would miss."

"its default prompts and may run" is missing a word (likely dropped during the multi-commit fixup). The parallel text in run-everything-mode.md says "without a flag the command prompts or backgrounds, returning only a status handle" — the intended phrasing in SKILL.md is probably "without this flag it prompts and may run in a background task" or "its default behavior is to prompt and may run in a background task".

Fix this →


Finding 2 — Forward-reference ordering tension in run-everything-mode.md (minor, clarity)

Line 10 (step 2):

"resolve that base (step 3 below) BEFORE this call"

Step 3 is where the review base is normally resolved ("Resolve the roster … and resolve the review diff base"), but step 2 now needs that base earlier. Telling the reader "step 3 below, but do it before step 2" is a minor contradiction with the numbered flow. A cleaner fix would be to split step 3 into "3a. Resolve review base" and "3b. Resolve roster", then reference "3a" in the step 2 note — but that's a restructuring judgment call, not a blocking issue. The current wording is functional (an executing model will read the full flow before acting), just slightly confusing for a human reader skimming linearly.


Positive observations

  • --wait and --base <review-base> are both documented with clear "why this flag is load-bearing" rationale — this is exactly the right level of detail for a skill-instruction doc that an agent must follow correctly.
  • The PR-mutation gate analysis for Codex is accurate: /codex:review is read-only, so no opt-in gate is needed unlike /code-review:code-review.
  • The graceful-absent pattern (Absent → note the skip) is consistent with the other two orchestrators.
  • The cross-vendor framing ("uncorrelated blind spots") is well-reasoned and worth the explanation in both files.
  • codex@openai-codex follows the plugin@namespace identifier convention the PR description verified against the enabled plugin.

Overall: LGTM pending the one-word grammar fix in SKILL.md. Finding 2 is an optional clarity improvement.

@kyle-sexton

Copy link
Copy Markdown
Contributor Author

@codex review

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale comment

Risk: low. Not approving because Cursor Bugbot skipped (usage limit reached), so the required automated review signal is missing. Human review is needed; no non-author reviewers were available to assign.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Router and Approver

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: dd7d2171d2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/review-toolkit/skills/code-review-fanout/SKILL.md Outdated
Comment thread plugins/review-toolkit/skills/code-review-fanout/context/run-everything-mode.md Outdated
…equire --base on both paths

Codex review round 3: the previous fix was a band-aid — it told the reader to
resolve the review base "before this call" while the numbered flow still computed
it in a later step, and it added --base only to run-everything, not to the
SKILL.md orchestrator bullet that governs default medium/large reviews.

Root-cause fix: promote review-base resolution to its own step (step 2) ahead of
the main-thread orchestrators, so Codex's --base is already resolved when it runs;
renumber the flow and its internal references accordingly. Add the --base
requirement to the SKILL.md codex bullet too, so default-mode dispatches Codex
against the shared base like every other surface rather than letting it auto-pick
the working tree or default branch.
@cursor

cursor Bot commented Jul 14, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_9ba6a57e-066d-4478-b22f-f27d84a2d7df)

@claude

claude Bot commented Jul 14, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 1m 45s —— View job


PR Review

  • Gather context from PR description and comments
  • Read changed files (current HEAD: 71b3326)
  • Review against CLAUDE.md guidelines and correctness
  • Post findings

Summary

Four-commit chain from the initial feat through three iterative fixes driven by Codex's own review rounds. The final HEAD (71b3326) correctly addresses both structural issues Codex raised: --wait for synchronous findings and --base <review-base> for diff-base alignment. One grammar defect from the prior Claude review (Finding 1 below) was not closed in any of the fix commits and is still present.


Finding — Garbled parenthetical in SKILL.md (minor, correctness)

Line 70:

"…returns findings in the same turn (its default prompts and may run in a background task the synchronous normalization step would miss)…"

The phrase "its default prompts and" is dropped/garbled — the parallel text in run-everything-mode.md step 3 reads "without a flag the command prompts or backgrounds, returning only a status handle". The intended SKILL.md phrasing is likely: "without this flag it prompts and may run in a background task the synchronous normalization step would miss".

Fix this →


Positive observations

71b3326 fixes are correct:

  • run-everything-mode.md: promoting review-base resolution to its own step 2 (ahead of step 3 orchestrators) resolves the ordering contradiction Codex flagged — Codex's --base now has a resolved value before the call.
  • SKILL.md line 70 now requires --base <review-base> with a reference back to "Shared inputs" for how to resolve it — so the default-mode path (medium/large tiers) gets the same constraint.

Flow numbering is internally consistent: "diffability pre-check (before step 5)" correctly targets the leaf fan-out step after renumbering. No orphaned step references.

Cross-vendor framing is accurate: the "uncorrelated blind spots" rationale holds — Codex is the only surface in the fan-out not backed by the same vendor model. A finding only Codex raises is structurally unverifiable by the Claude leaves, which is the right framing for why it earns its cost.

Graceful-absent pattern is consistent across all three orchestrators.

PR-mutation gate analysis is correct: /codex:review is read-only, no opt-in gate needed, unlike /code-review:code-review.


Overall: LGTM pending the one-word grammar fix in SKILL.md line 70. The --wait and --base issues raised across the Codex review rounds have been correctly resolved.

@kyle-sexton

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 71b3326567

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/review-toolkit/skills/code-review-fanout/context/run-everything-mode.md Outdated
Codex review round 4: --base forces the Codex companion into branch mode
(git diff <base>..HEAD, committed only), which coincides with the leaves'
git diff <base> on a clean branch but excludes uncommitted tracked edits a
dirty branch also carries. Rather than chase perfect diff-scope parity for
the mixed branch+dirty case (an edge of the exhaustive mode), document the
coverage boundary and require naming it in the ## Surfaces line — consistent
with this skill's "name the skip" discipline.
@cursor

cursor Bot commented Jul 14, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_d0fcc8a6-698b-49bb-9cee-4bcaf5662ad5)

@claude

claude Bot commented Jul 14, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 1m 51s —— View job


PR Review

  • Gather context from PR description and comments
  • Read changed files (HEAD: b81b15b)
  • Review CLAUDE.md guidelines alignment
  • Assess final Codex finding (branch-only vs working-tree coverage gap)
  • Post findings

Summary

Five-commit chain showing good iterative responsiveness: --wait, then --base, then promoting diff-base resolution to step 2 (before orchestrators), then noting the branch-mode coverage boundary. The structural concerns from all three Codex rounds are addressed. Two issues remain — one grammar defect that survived every review round unfixed, and a SKILL.md accuracy gap that the run-everything fix missed.


Finding 1 — Garbled phrase in SKILL.md, still unfixed after five commits (minor, correctness)

SKILL.md line 70:

"…returns findings in the same turn (its default prompts and may run in a background task the synchronous normalization step would miss)…"

"its default prompts and" is a dropped/garbled fragment. The parallel text in run-everything-mode.md step 3 reads "without a flag the command prompts or backgrounds, returning only a status handle" — the intended SKILL.md phrase is:

"…(without this flag it prompts and may run in a background task the synchronous normalization step would miss)…"

Both prior Claude reviews flagged this; neither of the two --wait/--base fix commits addressed it. This is purely documentation but the current text is misleading to anyone reading SKILL.md.

Fix this →


Finding 2 — SKILL.md overstates diff-set parity for dirty-branch dispatches (minor, accuracy)

SKILL.md line 70 (same line):

"…pass --base <review-base> carrying this skill's resolved review diff base ("Shared inputs") so Codex diffs the same change set as every other dispatched surface…"

Commit b81b15b ("note Codex --base branch-mode coverage boundary") correctly added to run-everything-mode.md step 3:

"Coverage boundary: --base runs Codex in branch mode (git diff <base>..HEAD, committed only), which coincides with the leaves' git diff <base> on a clean branch … but NOT when the branch also carries uncommitted tracked edits."

But this caveat landed only in run-everything-mode.md. SKILL.md governs the default-mode Codex dispatch path (medium/large tier) too — a model following SKILL.md to dispatch Codex in default mode on a dirty branch will read "same change set" with no qualification. The unconditional claim in SKILL.md is now known-false for mixed branch+dirty runs. The fix is a short qualifier matching the run-everything note:

"…so Codex diffs the same change set as every other dispatched surface on a clean branch — on a dirty branch, --base forces committed-only mode and misses uncommitted tracked edits; name that gap in ## Surfaces…"

Fix this →


Positive observations

  • Ordering fix is correct: promoting diff-base resolution to step 2 (before orchestrators) eliminates the forward-reference contradiction — --base has a resolved value before the Codex call.
  • run-everything-mode.md step 3 now has the correct coverage boundary note; a model running the exhaustive sweep is told to name the dirty-tree gap rather than assuming identical change sets.
  • --wait and --base rationale is documented at the right level of detail — the "why this flag is load-bearing" explanations in both files are exactly what an agent needs to follow these instructions correctly.
  • PR-mutation gate analysis is accurate: /codex:review is read-only, no opt-in gate required unlike /code-review:code-review.
  • Graceful-absent pattern is consistent across all three orchestrators.
  • Cross-vendor framing ("uncorrelated blind spots") is well-reasoned — a Codex-only finding is structurally unverifiable by the same-vendor leaves and orchestrators, which justifies the cost.
  • codex@openai-codex plugin identifier follows the plugin@namespace convention verified in the PR description.

Overall: LGTM pending the two SKILL.md fixes. Finding 1 (grammar) is a one-word change. Finding 2 (accuracy) requires carrying the dirty-tree caveat from run-everything-mode.md into SKILL.md so default-mode dispatches get the same warning.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Risk: low. Not approving because Cursor Bugbot skipped (usage limit reached), so the required automated review signal is missing. Human review is needed; no non-author reviewers were available to assign.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Router and Approver

@kyle-sexton
kyle-sexton merged commit 1724a25 into main Jul 14, 2026
17 of 19 checks passed
@kyle-sexton
kyle-sexton deleted the feat/codex-orchestrator-review-surface branch July 14, 2026 07:55

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b81b15bae6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


- **`pr-review-toolkit`** — `/pr-review-toolkit:review-pr`: aspect-scoped agent fan-out. Absent → this plugin's leaf agents cover most of the same dimensions; note that orchestrator breadth was skipped.
- **`code-review`** — `/code-review:code-review`: parallel reviewers + confidence scorer for an existing PR. **PR-mutation gate:** its PR mode posts findings as a PR comment, which violates the review modes' report-only contract; when the branch has an open PR, dispatch it only on explicit user opt-in ("post the review comment"), otherwise skip it and name the skip in `## Surfaces`. Absent → note the skip; a repository's own CI review bot (when present) still provides PR coverage.
- **`codex`** (OpenAI Codex) — `/codex:review`: read-only cross-vendor review, so it satisfies the review modes' report-only contract with no PR-mutation gate; `/codex:adversarial-review`: red-teams the diff, fitting the intentional adversarial-breadth intent. The first surface backed by a **different model** — its blind spots are uncorrelated with the same-vendor leaf agents and Claude orchestrators, so a finding only Codex raises is signal the rest structurally cannot see. Invoke it with `--wait` so the review runs in the foreground and returns findings in the same turn (its default prompts and may run in a background task the synchronous normalization step would miss), and pass `--base <review-base>` carrying this skill's resolved review diff base ("Shared inputs") so Codex diffs the same change set as every other dispatched surface — without it Codex auto-picks the working tree or default branch. Absent → note the skip; the same-vendor surfaces still cover most dimensions.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Don't claim Codex covers dirty default reviews

Fresh evidence after the prior fix: the dirty-tree caveat was added only to run-everything-mode.md, but this shared SKILL.md bullet still governs default medium/large orchestrator dispatch and says --base makes Codex diff the same change set. In a default review of an ahead branch with staged/unstaged edits, the leaves use git diff <review-base> and include those edits, while I checked codex-plugin-cc's companion code and any baseRef resolves to branch mode, whose branch context diffs only <merge-base>..HEAD; the persisted report will therefore overstate Codex coverage unless the same gap is documented here or default mode runs a separate working-tree Codex pass.

Useful? React with 👍 / 👎.

kyle-sexton added a commit that referenced this pull request Jul 24, 2026
…#1240)

## Summary

Close-out of the ci-workflows work-loop topic (work-loop PLAN Phases 1-7
+ babysit track B1-B3, all [DONE]). Advances the final two phase tags
with recorded Phase 7 / B3 evidence, then prunes the branch-lived
contract slice per the topic-docs lifecycle. Durable outcomes already
live outside the slice: the loop-lane convention
(docs/conventions/loop-lane/ — the durable decision record; no separate
ADR meets the admission test), the rate-limit-guard plugin, the
work-items work-loop/attend-queue skills, the source-control
babysit-loop skill + layered seam keys, and the ci-workflows tracker
binding.

Phase 6 disposition: standards branch docs/activate-lease-claiming
remains pushed and deliberately UNFILED — operator files it (plan
constraint); the verified one-liner is in the operator's hands.

## Test plan

- Docs-only: repo lint gates (markdownlint; changelog-parity untouched —
no plugin content changed)
- Tag audit: 0 [TODO]/[DOING] remaining in the slice at commit 1; slice
absent after commit 2

## Related

No linked issue — plan close-out per the planning skill's close-out
procedure. The final babysit-PLAN.md (B1-B3, same close-out treatment)
is preserved at commit df5213c
`docs/topics/ci-workflows-work-loop/babysit-PLAN.md` — referenced rather
than pasted to stay under the PR-body size cap. Lane wind-down
telemetry: ci-workflows#231/#232.

<details><summary>Final PLAN.md (work-loop) at close-out</summary>

```markdown
# ci-workflows-work-loop

## Brief

**TLDR**: Drain repo issue backlogs autonomously via a 3-session-per-repo topology — worker loop
authors PRs (never merges), babysit lane owns merges at a merge-capable tier (worker minimum;
autopilot where only bot PRs are in scope), HITL decides — built as reusable org-scope plugin
lanes; ci-workflows (21 open issues) is the first instantiation. Sibling deliverable with its own
Brief: `/source-control:babysit-loop` (`babysit-PLAN.md`, same directory — succession from the
terminated babysit session).

### Goal

1. Two new skills in the `work-items` plugin (`melodic-software/claude-code-plugins`):
   `work-loop` (dynamic self-paced drain: mechanical-triage intake sweep, adaptive parallel
   dispatch, escalation, heartbeat, exit-condition evaluation) and `hitl-queue` (attended lane:
   polls escalations + untriaged intake, composes `/work-items:triage` + `/planning:interview`,
   writes answers back, flips items to agent-ready). Work-lane-specific mechanics ONLY — shared
   cross-lane concerns live in the convention doc (goal 1b), which both plugins cite.
   Interim-executor rationale cites `plugins/autonomy/reference/runner.md`.
1b. New repo-level convention `docs/conventions/loop-lane/` + registry row
   (per `docs/PLUGIN-PHILOSOPHY.md`: sibling-plugin imports are defects; a cross-plugin
   convention lands in an owner doc before a second plugin adopts it — the babysit lane in
   `source-control` is that second adopter). Owns: 3-session topology, escalation contract,
   model-routing capability tiers (IDs live-resolved), and the loop-layer invariants — stop
   shapes (standing vs drain), `#691` cycle-budget semantics (resolution (a): budget restarts the
   SESSION, never ends the LOOP), idle backoff, `#502` per-lane edited status comment,
   headless-config floor (a headless lane launch never blocks on an interview: explicit/persisted
   config or tier defaults + log), and the subagent discipline preamble. `autonomy` remains the
   governing-policy pointer (guardrail matrix), not a mechanics home.
2. New small `rate-limit-guard` plugin: statusline tee wrapper (writes `rate_limits` + timestamp to
   fixed path), reader contract (pause ≥90% of either window until later `resets_at`; Monitor
   watches tee file so account switches wake the pause early — MANDATORY, not optional: the
   statusline schema carries no account-identifier field, verified live), `StopFailure` matcher
   `"rate_limit"` reactive fallback (side-effect-only detection — hook output/exit ignored,
   payload carries no reset/quota data; resume timing comes from the tee file or parsed error
   text, verified live), capability-detect fail-open (absence/absurd values = unknown,
   reactive-only — covers enterprise/API-key auth, where limits may exist but are unobservable).
   Implementation notes: shell-run fields reject `${user_config.*}` (use exec-form args/env/config
   file); plugin `experimental.monitors` auto-start is a candidate watcher surface; babysit-loop
   and hitl sessions are consumers two and three of the shared windows.
3. ci-workflows instantiation: `.work-item-tracker.json` (github provider) via PR; initial triage
   pass over the 14 unlabeled issues (attended, this session); loop run to exit condition.
4. `standards` doc PR: record activation of the deferred assignee+lease claiming mechanism
   (`conventions/process/issue-tracker.md` status axis).

### Constraints

- Worker loop is PR-only; merging belongs to babysit worker tier; judgment belongs to HITL.
- Exit condition: every open issue closed OR has an active non-draft PR (native closing-keyword
  linkage; the seam's in-flight-PR exclusion is the drain signal).
- Label set is github-iac-owned: no ad-hoc label creation; role labels resolved via
  `.work-item-tracker.json` `config.role_labels`, never compared as literals. Worker-escalated vs
  parked discriminated by machine-marked bot comment, not a new label.
- Skill bodies are thin delta only — invoke owned mechanics (`/work-items:work`,
  `/implementation:implement-dispatch`, `/source-control:worktree`, `/work-items:triage`,
  `/planning:interview`); never restate selection/claim/lease/dispatch/triage machinery.
- Worker dispatch prompts enumerate required skills per phase; subagent authority is an open-ended
  duty ("escalate decisions that are the operator's — e.g. contract, security-posture,
  enforcement-scope, issue-goal changes"), not a closed list.
- Work-class gate, in autonomy contract terms (`reference/guardrails/work-classes.md` +
  `admission-policy.md`, read this session): `C2` mechanical → autonomous-eligible; `C3` scoped →
  autonomous for bug-fix-shaped items, human-gated for feature-shaped (operator tightening —
  permitted without justification); `C4` structural/contract and `C5` untrusted-provenance →
  human-gated / per-contract floors; unclassified → fail-closed human-gated (contract rule).
- Model routing: capability TIERS in the convention doc, defined by capability ORDER, never by
  family name (verified: family ≠ capability across generations — the advisor matrix ranks
  Sonnet 5 equal to Opus 4.6). Live resolution: model aliases (`fable`/`opus`/`sonnet`/`haiku` —
  the only live-updating handles; dateless IDs are pinned snapshots) + the Models API list
  endpoint; account for provider-dependent alias resolution and org `availableModels` overrides.
  Advisor rule as documented: "at least as capable" as main (equal pairings valid); fast-main +
  strong-advisor pairing is officially recommended (advisor docs + blog, cited in
  scratchpad model-routing-verification/RESEARCH.md). Defaults: fast-tier orchestrator +
  advisor ≥ main; strong-tier workers;
  frontier tier for complex-stamped items; fast tier for mechanical; reviewer/verifier never
  weaker than implementer; tier signal from triage briefing (issue body, not a label);
  security-surface work-class → frontier tier, ALWAYS (operator directive). Windows: 5-hour
  ROLLING + weekly FIXED-ANCHOR (not rolling — verified); the per-model weekly cap's family is
  mid-transition in official docs (Sonnet vs Opus conflict) — LINK the support article, never
  hard-code the family; Fable's 50%-of-weekly allowance on Max is documented.
- Adaptive item-level cap: start 2, +1 after 3 consecutive clean items (ceiling 3), −1 on any dirty
  item (floor 1); no ramp-up while a rate-limit warning is live; composed budget with
  implement-dispatch's per-item wave cap stated explicitly.
- Rate limits: drain-then-pause default (finish in-flight, stop claiming, ScheduleWakeup/Monitor
  until reset, report); hard stop only by user request; model downshift only as explicit opt-in.
- Org-agnostic throughout; no machine-/user-specific paths in skill bodies.
- ci-workflows `CLAUDE.md` security ground rules are inviolable by workers (escalate, never edit).
- Authoring governed by `playbooks:skill-authoring` + `skill-quality:check`; WebFetch current
  official plugin/skills docs during build (claude-code-plugins CLAUDE.md mandate); hook wiring
  (StopFailure) via `update-config`; model-routing tier tables built from a live official-docs
  fetch at build time, never recall.
- Prep before build: pull ci-workflows (behind 11) + standards (behind 19); claude-code-plugins →
  main + pull (current branch's upstream deleted).

### Acceptance criteria

- Loop runs unattended on ci-workflows until the exit condition; final report lists items closed,
  PR'd, and escalated.
- Escalated items carry `needs-human` (role-resolved) + machine-marked question comment; attend-queue
  surfaces them and untriaged intake in one attention view.
- Two parallel workers never hold the same item (seam lease race observed working).
- Guard pauses claiming at ≥90% of either window and resumes after reset without operator action
  (subscription auth); on non-subscription auth it declares reactive-only and never throttles.
- Both lanes + guard plugin pass `skill-quality:check`.
- Second-repo adoption = one binding PR + three sessions; zero plugin edits.

### Captured assumptions

- Subscription (Pro/Max) auth in loop sessions for proactive throttle; anything else fail-opens to
  reactive.
- work-items `#572` (autonomous branch/worktree provisioning seam) worked around in-template via
  explicit provisioning before dispatch; upstream fix optional follow-up.
- Statusline runs (and tees) in interactive sessions; headless `claude -p` unverified and unused.
- `/login` multi-account rotation is undocumented community practice — manual operator action
  only; the design never automates or depends on it.

### Out of scope

- The autonomy runner build (stays charter-gated; this loop is the interim executor earning its
  triggers — cited, not built).
- Routine-catalog graduation of heartbeat/triage sweeps (trigger recorded: loop proves stable).
- Babysit plugin lane (owned by the user's separate session).
- Worker-loop merging; new labels; the unofficial `api/oauth/usage` endpoint.

### Deferred questions

- Multi-repo scale-out shape (per-repo loops vs round-robin worker) — trigger: second repo adopts.
  Arbiter: USER-RESERVED.
- Upstream work-items PR for `#572` once the in-template workaround proves out. Arbiter:
  /planning:plan.
- Build-time verifications: StopFailure `rate_limit`
  matcher behavior on API-key auth; implement-dispatch per-worker model override wiring (substrate
  confirmed: agent frontmatter supports `model`, `effort`, `skills` preload, `isolation:
  worktree`); worker
  commit signing vs merge strategy against the `signing` ruleset (`required_signatures` on main —
  verified live); machine-readable surfacing of Q18's quality signals (design intent; check
  pipeline-shape.md). Arbiter: /planning:plan.

## Plan

> Scope-change note (2026-07-23, approval round): three Brief statements superseded by
> operator decision — (1) "babysit merges at worker tier minimum / autopilot for bot-only
> repos" → autonomy ladder with human-merge default (Phase 2 item 1; babysit-PLAN B3 tier =
> worker); (2) skill name `hitl-queue` → `attend-queue` (naming-grammar rule); (3) exit
> condition gains the drain-terminal state. Brief text above is historical; this Plan is
> authoritative for the deltas.

### Standards grounding

No standards index (`docs/standards/README.md` with `standards-contract` frontmatter) exists in
claude-code-plugins, ci-workflows, or standards — resolution ladder rung 4/6 (inference; assumption
surfaced). Grounding sources read this session:

| Surface | Source | Provenance |
|---|---|---|
| Plugin/skill design | claude-code-plugins `CLAUDE.md` (fresh-docs mandate L6-11, design rules L37-50); `docs/PLUGIN-PHILOSOPHY.md` (seam rules L22-33, convention registry L329-353, hook contract L326-328, Windows exec-form L128, Monitors stance "Wait" L134) | team (repo) |
| Work-items seam | `plugins/work-items/reference/tracker-seam.md`, `label-taxonomy.md`, `tools/work-item-tracker/CONTRACT.md` (lease protocol L207-231) | team (repo) |
| Autonomy contracts | `plugins/autonomy/reference/guardrails.md` (matrix, escalation L68-94), `guardrails/work-classes.md`, `admission-policy.md`, `runner.md` (interim-executor L8-51) | team (repo) |
| Source-control | `plugins/source-control/skills/babysit-prs/SKILL.md` + `reference/loop.md`, `reference/config-resolution.md` | team (repo) |
| Lane ops | `plugins/claude-ops/skills/lanes/SKILL.md` (launcher, `telemetry-upsert.sh` #502 interim, #480 prompt seam) | team (repo) |
| Security | ci-workflows `CLAUDE.md` L7-50 (inviolable ground rules) | team (repo) |
| Tracker process | standards `conventions/process/issue-tracker.md` (L37 deferred lease, L52-58 label IaC governance) | team (repo) |

### Build-technique note

No viability unknown remains (interview + research resolved design); risk is integration. Kept
tracer-bullet slice = Phase 5+7 head: bind the tracker seam and run ONE work-loop cycle
end-to-end on ci-workflows before declaring the lane operational. Phases 2–4 are the
convention/plugin substrate that slice consumes.

### Phase 1: Topic-contract move + branches [DONE]

Move both Briefs to the tracked contract slice in claude-code-plugins (operator-confirmed
placement; topic-docs convention `contract_tier: branch`):

| File | Action | What |
|---|---|---|
| `claude-code-plugins/docs/topics/ci-workflows-work-loop/PLAN.md` | Create | this file (Brief + Plan) |
| `claude-code-plugins/docs/topics/ci-workflows-work-loop/babysit-PLAN.md` | Create | babysit Brief + Plan |
| `claude-code-plugins/docs/topics/ci-workflows-work-loop/design/design-resolution.md` | Create | design gate file — a contract kind per topic-docs convention (README tier table), moves with the slice |
| `.work/ci-workflows-work-loop/*` | KEEP | working state; research + ledgers stay memory-slice |

Branch per PR-phase (`feat/loop-lane-convention`, `feat/rate-limit-guard`,
`feat/work-items-loop-lanes`, …) — topic docs ride the first build branch and each later branch
carries its own phase-status updates.

- **Sanity Check:** `git -C <claude-code-plugins> ls-files docs/topics/ci-workflows-work-loop/` lists both PLAN files + `design/design-resolution.md` (exit 0, 3 lines).

### Phase 2: `docs/conventions/loop-lane/` convention + registry row [DONE]

Owner doc for every shared cross-lane concern (PLUGIN-PHILOSOPHY L331-333: owner doc lands BEFORE
the second adopter). One PR to claude-code-plugins.

| File | Action | What |
|---|---|---|
| `docs/conventions/loop-lane/README.md` | Create | full contract (content below) |
| `docs/conventions/loop-lane/CHANGELOG.md` | Create | `## 1.0.0 — <date>` |
| `docs/PLUGIN-PHILOSOPHY.md` | Modify | one registry row: `\| Loop-lane topology, escalation, capability tiers, loop invariants \| [docs/conventions/loop-lane/](conventions/loop-lane/README.md) \|` |

README.md contract sections (pointer-not-copy — cite owners, never restate):

1. **3-session topology** — worker loop authors PRs (never merges); babysit lane owns merges
   WITHIN the autonomy guardrail matrix's merge-policy column; HITL decides. Merge authority
   (operator-resolved 2026-07-23): a configurable AUTONOMY LADDER with the safest default —
   **human merge is the shipped default for everything except gate-proven C2-mechanical
   PRs** (dependabot/small/easy/safe candidates; a work-class test irrespective of author — bot
   authorship alone never qualifies, per the loop-lane convention); higher rungs up to full autonomy (frontier
   subagents resolving conflicts, responding to comments, making changes, driving to merge)
   are opt-in per repo. The ratification mechanism: raising a merge rung is a config change on
   the TRACKED layered seam (reviewable, versioned) — exactly the matrix's required
   "human-ratified knob flip recorded on the governance surface" (`work-classes.md` Promotion
   L57-59), so C3 autonomous merge is reachable only through a recorded flip, never by
   default. Operator may adjust rungs on the fly; each adjustment is a tracked config edit.
2. **Escalation contract** — `needs-human` role label resolved via `.work-item-tracker.json`
   `config.role_labels` (cite `plugins/work-items/reference/label-taxonomy.md` canonical roles)
   \+ machine-marked bot comment discriminator (worker-escalated vs parked). Cites autonomy
   `reference/guardrails.md#escalation` event classes as governing policy; no second channel.
3. **Capability tiers** — order-defined tiers (never family names); RUNTIME resolution via
   aliases ONLY (the only handles guaranteed under subscription OAuth — stress-test finding:
   the Models API list endpoint needs API-key auth loop sessions may lack); Models API is the
   build/audit-time verification path; re-derivation trigger recorded in the convention
   CHANGELOG: any new model release re-audits the tier table. Advisor "at least as capable";
   security-surface class → frontier ALWAYS; weekly-cap family LINKED to the support article,
   never hard-coded. Tier tables built from live official-docs fetch at build time (fresh-docs
   mandate).
4. **Loop-layer invariants** — stop shapes (standing vs drain, PLUS the drain-terminal state
   `[FALLBACK — confirm or override]`: when every remaining open item is human-gated/escalated
   and no PR is in flight, drain reports + stops cleanly instead of idling forever — reviewer
   finding: without it an overnight drain deadlocks on the first unanswered escalation);
   #691 resolution (a): cycle budget restarts the SESSION, never ends the LOOP — restart is
   launcher/operator-initiated (a running loop cannot `/clear` itself; cite `claude-ops` lanes
   SKILL.md "A relaunch is the only context reset"), so budget-hit = emit restart-request
   telemetry + stop loop cleanly; idle backoff; #502 per-lane edited status comment — the
   convention FIXES the comment contract (machine sentinel marker, edit-in-place, one comment
   per lane, per #502) and each lane INLINES the small `gh api` upsert (operator coupling
   directive 2026-07-23 + philosophy: an installed plugin cannot invoke a sibling plugin's
   scripts; `claude-ops` `telemetry-upsert.sh` is cited as the interim #502 home and a
   compatible consumer — `morning-brief` reads the same surface — coupling stays
   one-directional);
   **durable loop state across compaction** — adaptive-cap streak counter, rate-limit-warning
   latch, and cycle count persist in a machine-readable block of the lane's #502 telemetry
   comment, re-read at each cycle start (conversation context alone is compaction-lossy);
   headless-config floor (headless launch never blocks on interview: explicit/persisted config
   or tier defaults + logged notice); subagent discipline preamble (sweep-all-disciplines +
   use-your-skills + do-your-research); seam exit 8 (provider unavailable / secondary GitHub
   limits under one credential) → backoff-and-retry, counted as a dirty signal for the adaptive
   cap; drain exit evaluated against a cycle-start snapshot — new automated intake is REPORTED,
   never chased.
5. **Consumers + launch surfaces** — work-items `work-loop`/`attend-queue`, source-control
   `babysit-loop`; launchable interactively via `/loop` (primary, dependency-free) or headless
   via `claude-ops:lanes` (`prompt_dir` seam, #480 forward dependency) — lanes is a SUPPORTING
   launcher, strictly one-directional (operator coupling directive 2026-07-23): lanes launches
   the skills; no loop skill ever references, requires, or degrades without `claude-ops`; every
   mention in skill bodies is guarded if-installed with the `/loop` fallback documented.
   `autonomy` remains governing-policy pointer.
6. **Rate-limit guard binding** — all three lanes are consumers of the shared subscription
   windows. Operability rule (reviewer finding): an installed plugin cannot read sibling-plugin
   files or this repo's `docs/` at runtime — each consuming skill body INLINES the operable
   floor (tee-file path, 90% threshold, staleness rule, drain-then-pause) and cites the guard
   reader contract for provenance only. The convention records this inline-floor rule so the
   values stay byte-identical across consumers (fleet audit checks conformance per row).
   Single-account-per-machine invariant (stress-test finding: tee file is last-writer-wins with
   no account id — a mid-drain `/login` to another account feeds healthy windows to lanes on
   the exhausted one; the guard cannot detect this, the convention says so); lane telemetry
   records guard mode (proactive/reactive/unknown) each cycle so silent degradation is visible
   in the #502 surface.

- **Sanity Check:** `grep -c "loop-lane" docs/PLUGIN-PHILOSOPHY.md` ≥ 1 (registry row present).
- **Sanity Check:** `test -f docs/conventions/loop-lane/README.md && test -f docs/conventions/loop-lane/CHANGELOG.md` exit 0.
- **Sanity Check:** repo CI green on the PR (`ci-status` required check).
- **Sanity Check:** `grep -En "plugins/(work-items|source-control|claude-ops|autonomy)/" docs/conventions/loop-lane/README.md` returns citations only in pointer form (reviewed: no restated mechanics blocks >3 lines quoting a plugin file).

### Phase 3: `rate-limit-guard` plugin [DONE]

Review: security
New small plugin in claude-code-plugins. One PR. Pre-flight consumer check N/A (new surface, no
existing consumers). Fresh-docs mandate: WebFetch statusline + hooks reference pages before
authoring; cite URLs in the PR.

| File | Action | What |
|---|---|---|
| `plugins/rate-limit-guard/.claude-plugin/plugin.json` | Create | manifest, semver 0.1.0, `userConfig`: hook kill switch ONLY (reviewer finding: threshold/tee-path overrides are dead config — `${user_config.*}` is plugin-scoped and cross-plugin consumers read the documented fixed values; contract path + 90% are FIXED) |
| `plugins/rate-limit-guard/hooks/hooks.json` | Create | `StopFailure` matcher `"rate_limit"` → `record-rate-limit-stop.sh` |
| `plugins/rate-limit-guard/hooks/record-rate-limit-stop.sh` + `.test.sh` | Create | side-effect-only detection record (hook output/exit ignored by harness — verified); test-first |
| `plugins/rate-limit-guard/hooks/hook-utils.sh` | Create | synced copy via `scripts/sync-hook-utils.sh` (CI byte-parity gate) |
| `plugins/rate-limit-guard/scripts/statusline-tee.sh` + `.test.sh` | Create | wraps the user's statusline command; tees stdin JSON `rate_limits` + ISO timestamp to the contract path via ATOMIC write (temp file + rename — 3+ concurrent sessions write one path; readers must never see torn JSON; Windows: rename-over-open-target can EACCES without FILE_SHARE_DELETE — retry-then-skip, and a tee failure NEVER propagates into the statusline pipeline; `.test.sh` covers the locked-target case); tees any session-distinguishing field present so a future account-id field is adopted automatically; defined no-statusline path (installs as standalone minimal statusline when none configured); Windows/Git Bash shell requirement stated in the script header and verified live (first instantiation is the operator's Windows machine); test-first |
| `plugins/rate-limit-guard/skills/setup/SKILL.md` | Create | CHECK-ONLY setup (PLUGIN-PHILOSOPHY L270: setup must NOT mutate Claude Code user settings): `check` verifies jq, probes tee freshness (distinguishing "no statusline configured" from "wrapper missing"), and PRINTS the exact `settings.json` statusline edit for the operator to apply; surfaces the chezmoi/dotfiles tracking proposal for the operator-changed user-scope file |
| `plugins/rate-limit-guard/reference/reader-contract.md` | Create | pause when EITHER window ≥90%, until the TRIPPED window's `resets_at` (the later `resets_at` applies only when BOTH windows trip — reviewer finding: the earlier phrasing over-pauses a 5-hour trip until the weekly reset); staleness rule (no account-id field — consumers MUST arm a session Monitor on the tee file; plugin ships no monitor config, fleet Monitors stance is "Wait" per PLUGIN-PHILOSOPHY L134); capability-detect fail-open (absent/absurd = unknown = reactive-only); drain-then-pause; fixed tee path + threshold as the operable constants consumers inline; consumer list (3 lanes) |
| `plugins/rate-limit-guard/README.md`, `CHANGELOG.md` | Create | per repo norms |
| `.claude-plugin/marketplace.json` | Modify | plugin entry (no version field) |
| `scripts/skill-leaf-name-registry.txt` | Modify | register `setup` leaf if registry requires per-plugin entries (verify gate script semantics at build) |

Build-time verifications (Brief "Deferred questions", arbiter: this plan): StopFailure
`rate_limit` matcher behavior on API-key auth (empirical: forced-failure probe); statusline
stdin schema fields present on this CLI version (live probe); shell-run fields reject
`${user_config.*}` → exec-form args/env only (already verified; re-verify against current docs).

- **Sanity Check:** `bash scripts/validate-plugins.sh` exit 0.
- **Sanity Check:** `bash scripts/run-plugin-tests.sh` exit 0 (includes new `.test.sh` files).
- **Sanity Check:** `bash scripts/sync-hook-utils.sh --check` exit 0.
- **Sanity Check:** live probe — with wrapper wired, `jq -e '.rate_limits and .captured_at' <tee-file>` exit 0 after one statusline refresh. `[OPERATOR-GATED — the statusline wiring is the operator's settings.json edit (setup is check-only by philosophy; the skill prints the exact edit); the build verified the same probe against the wrapper fed the documented schema payload in a scratch HOME. Runs live as a Phase 7 pre-flight once the operator applies the edit.]`
- **Sanity Check:** `/skill-quality:check setup` PASS.

### Phase 4: work-items `work-loop` + `attend-queue` skills [DONE]

Review: architecture
Two thin-delta skills in `plugins/work-items/skills/`. One PR. Bodies invoke owned mechanics —
never restate selection/claim/lease/dispatch/triage machinery (constraint; the seam hub
`reference/tracker-seam.md` is read at invocation start per plugin norm).

| File | Action | What |
|---|---|---|
| `plugins/work-items/skills/work-loop/SKILL.md` | Create | loop layer (content below) |
| `plugins/work-items/skills/work-loop/evals/evals.json` | Create | per evals schema |
| `plugins/work-items/skills/attend-queue/SKILL.md` | Create | attended lane (content below) |
| `plugins/work-items/skills/attend-queue/evals/evals.json` | Create | per evals schema |
| `plugins/work-items/CHANGELOG.md` + `.claude-plugin/plugin.json` | Modify | version bump + any new `userConfig` keys (adaptive-cap bounds) |
| `scripts/skill-leaf-name-registry.txt` | Modify | register new leaf names (gate) |

`work-loop` delta only: per-cycle intake sweep (mechanical triage via `/work-items:triage`
autonomous lane — #459 CLOSED, autonomous mode exists); work-class gate in autonomy contract
terms (C2 → autonomous; C3 bug-fix-shaped → autonomous, feature-shaped → human-gated (operator
tightening, no justification needed per `admission-policy.md` L38-40); C4/C5/unclassified →
human-gated fail-closed); adaptive item cap (start 2, +1 after 3 consecutive clean, ceiling 3,
−1 on dirty, floor 1; no ramp while rate-limit warning live; QUOTA GUARD (operator-resolved
2026-07-23): frontier-tier items run at concurrency 1 with adaptive ceiling 2 — the general
ceiling 3 applies to non-frontier tiers only; composed-budget statement with
implement-dispatch per-item wave cap — note #573: userConfig cap enforcement not yet wired,
loop-body arithmetic is the interim enforcement); #572 workaround in-template (explicit
`/source-control:worktree` provisioning before dispatch); escalation + telemetry + stop shapes +
budget semantics BY CITATION to `docs/conventions/loop-lane/`; guard integration by citation to
reader contract; exit condition = seam frontier empty ∧ every open issue closed or has active
non-draft PR (GraphQL close-linkage per `adapters/github/README.md` L183-192); dynamic `/loop`
pacing via ScheduleWakeup. Q18 quality-signal surfacing: verify design intent against
`reference/pipeline-shape.md` at build (deferred-question arbiter: this plan).

`attend-queue` delta only: attended poll of (a) `needs-human` role-labeled items with
machine-marked comments, (b) untriaged intake via `/work-items:triage` attention view;
composes `/work-items:triage` + `/planning:interview`; writes answers back as issue comments;
flips items to `agent-ready` role label (resolved, never literal).

Cross-cutting work items (both skills, reviewer findings):

- Inline the guard operable floor (fixed tee path, 90% threshold, staleness rule,
  drain-then-pause) per the convention's inline-floor rule; cite reader contract for provenance.
- Every cross-plugin invocation (`/implementation:implement-dispatch`,
  `/source-control:worktree`, `/planning:interview`) follows the declared-or-guarded rule +
  seam-phrasing convention (gate + documented fallback per site) — same phrasing
  `plugins/work-items/skills/work/SKILL.md` already uses for the same targets.
- Naming (operator-resolved 2026-07-23): the Brief's `hitl-queue` violated the repo
  imperative-verb rule (PLUGIN-PHILOSOPHY L41); renamed to `attend-queue` (verb-first)
  throughout this Plan. `work-loop` is verb-first and passes.
- Exit-condition includes the drain-terminal state (all remaining open items human-gated →
  report + clean stop) per the Phase 2 invariant `[FALLBACK]`.
- Admission hardening (stress-test findings — the C3 classifier is the loop that benefits from
  admitting the item): (a) instantiation-level path/topic HARD GATE — any item touching SHA
  pins, checksums, or ci-workflows CLAUDE.md ground-rule surfaces is human-gated regardless of
  classification (the tool-version-drift advisories surface-read as C2 dep bumps but are
  trust-on-first-download checksum recomputation on security-critical pins, operator-owned per
  ci-workflows CLAUDE.md L52-58); (b) triage author-rule: workflow-bot-authored advisory issues
  → `needs-human` by default (also makes drain exits terminate against automated intake);
  (c) HITL ratifies every C3-autonomous admission for the first full drain (operator-adopted
  2026-07-23; earn-trust posture).

- **Sanity Check:** `/skill-quality:check work-loop` PASS and `/skill-quality:check attend-queue` PASS.
- **Sanity Check:** `bash scripts/validate-plugins.sh` exit 0 (evals schema, leaf-name, catalog gates).
- **Sanity Check:** `grep -c "conventions/loop-lane" plugins/work-items/skills/work-loop/SKILL.md` ≥ 1 and same for `attend-queue` (convention cited, not restated).
- **Sanity Check:** `grep -En "claude-(opus|sonnet|haiku|fable)-[0-9]" plugins/work-items/skills/{work-loop,attend-queue}/SKILL.md` returns empty (no hard-coded model IDs; tiers only).

### Phase 5: ci-workflows tracker binding [DONE]

One PR to ci-workflows. Pre-flight consumer check (contract surface = new tracked file): grep
ci-workflows workflows/scripts for `.work-item-tracker.json` readers — none expected (file is
new; seam tooling is the only reader).

| File | Action | What |
|---|---|---|
| `ci-workflows/.work-item-tracker.json` | Create | `{"schema_version":"1.0","provider":"github","config":{"lease_ttl_hours":24}}` (role_labels omitted → canonical defaults `agent-ready`/`needs-human`/`recurring` with loud warning; defaults match the org label taxonomy) |

Phase-entry checks: (a) `gh label list -R melodic-software/ci-workflows --limit 100` — confirm
`agent-ready`, `needs-human`, `recurring` exist; any missing label → STOP, route a `github-iac`
Pulumi PR first (label set is IaC-owned; never `gh label create`). (b) worker commit signing vs
`required_signatures` ruleset: RESOLVED by live verification during the stress-test — the
`signing` ruleset targets the default branch only, `base` allows squash as the sole merge
method, and GitHub web-flow signs API-created squash commits; worker feature-branch commits sit
outside the ruleset's ref scope. Retain a one-command build-time confirmation as a formality.

- **Sanity Check:** seam `capabilities` verb exit 0 in the ci-workflows checkout, and
  `list-frontier --autonomous` exit 0 returning a JSON envelope (the `health` verb does not
  exist — verified against `CONTRACT.md` verb table).
- **Sanity Check:** two-session claim race — CONCURRENT, not sequential (stress-test finding:
  sequential A-then-B passes even with a wide race window): fire both `claim <id>` calls in
  parallel; exactly one exits 0 and one exits 7, OR the residual is documented (GitHub comment
  listing propagation window — consequence is a duplicate PR, recoverable, caught by
  babysit/HITL; a jittered delay before the lease re-read shrinks it cheaply).

### Phase 6: standards lease-activation PR draft [DONE]

Prepare (never file) the standards change editing `conventions/process/issue-tracker.md` L37
block: record activation of assignee+lease claiming (trigger met: autonomous multi-worker
contention is now real), point at work-items `tools/work-item-tracker/CONTRACT.md` lease
protocol as the mechanism, retire the interim 🔒 marker text. "Never auto-filed" (Brief goal 4)
means: the session pushes the branch + prepares the PR body file and notifies the operator; NO
`gh pr create` is executed by the session or any lane — the operator files and merges.

- **Sanity Check:** branch exists on origin (`git ls-remote --heads origin <branch>` non-empty),
  PR body file present in the topic slice, operator notified in the session report;
  `gh pr list -R melodic-software/standards --head <branch>` returns EMPTY (nothing auto-filed).

> **Close-out disposition (2026-07-24):** branch `docs/activate-lease-claiming` remains
> pushed (head `ce1b9a7c`) and UNFILED — operator filing still outstanding. The verified
> operator one-liner (title from the head commit subject; body = the prepared file's
> content below its memo separator) was handed to the operator 2026-07-24. The body file
> stays in the ci-workflows memory slice until filed; it is machine-local working memory,
> not a tracked artifact.

### Phase 7: attended triage + launch + drain [DONE]

> **Close-out note (2026-07-24):** all four sanity checks recorded; the drain reached an
> operator-stabilized terminal state rather than the strict exit condition. Evidence
> (distilled; gathered 2026-07-24 ~02:10Z):
>
> - Triage: 19/19 attended pass completed 2026-07-23 (prior session, quoted in handoff);
>   later intake drift (new telemetry issues + one intake row) is live-tracker steady
>   state, not plan scope.
> - Exit-condition query: run against live state — NOT met on both conjuncts:
>   `list-frontier --autonomous` returned 4 lane-telemetry issues (a seam defect — the
>   telemetry contract says these never gate; tracked in ci-workflows) plus holdouts
>   #238 (blocked) and #177 (operator-assigned); 20 needs-human; 0 open PRs. The
>   operator stopped all three lanes 2026-07-23 ~23:15-23:20Z with the work-loop rate
>   guard tripped (seven-day 98%, pause to 2026-07-29T03:00Z); remaining items are
>   needs-human/design-gated — the drain-terminal remainder lives in the tracker, not
>   this plan.
> - Guard pause+resume: demonstrated via the plan-authorized synthetic probe (pause at
>   future `resets_at`, resume on back-dated `resets_at`; tee restored byte-identical,
>   sha256 verified). Live trip also observed operationally: work-loop wound down on a
>   real seven-day 98% trip.
> - Attend-queue attention view: real data — 2 escalated + 2 ratify + 1 intake rows in
>   one view.

1. Attended triage pass (operator present) over live untriaged intake via `/work-items:triage`
   attention view — 10 zero-label issues at plan time (Brief said 14; count drifted, use live).
2. Launch three lanes: `work-loop` + `attend-queue` + babysit (babysit-PLAN Phase B3). Launch
   surface `[FALLBACK — confirm or override]`: FIRST cycle attended via interactive `/loop`
   (observability); unattended operation then moves to `claude-ops:lanes` (background named
   sessions) so budget-hit restart-requests reach the operator's one-command relaunch surface
   (`lanes restart` — operator-invoked by its contract; no automatic trigger consumes
   restart-requests today, per the loop-lane convention) — reviewer finding:
   an interactive `/loop` with nobody watching turns a #691 budget-hit into a silent drain
   stall. Until lanes-managed, the stall risk is documented in the lane report.
   PHASE-ENTRY GATE before any lanes migration (stress-test finding: statusline is verified
   interactive-only): verify statusline execution + tee freshness under ONE `--bg`
   lanes-launched probe session. If `--bg` sessions do not tee, either keep one attended anchor
   session alive (one machine, one tee file serves all readers) or defer the migration and
   document unattended operation as reactive-only.
3. Drain to exit condition; final report lists items closed / PR'd / escalated (acceptance
   criterion).

- **Sanity Check:** `gh issue list -R melodic-software/ci-workflows --state open --json number,labels` — zero items lacking triage outcome (every open item labeled or role-labeled).
- **Sanity Check:** exit-condition query — seam `list-frontier --autonomous` returns empty AND
  every remaining open issue has an OPEN close-linked PR (GraphQL probe) OR carries the
  `needs-human` role label (drain-terminal state, per the Phase 2 `[FALLBACK]` invariant).
- **Sanity Check:** guard behavior observed once: pause log entry at ≥90% or explicit
  reactive-only declaration on non-subscription auth; AND one resume observed without operator
  action after a `resets_at` passes (acceptance criterion covers resume, not just pause — if no
  window trips during the drain, record a synthetic tee-file probe: back-dated `resets_at`
  unpauses the loop).
- **Sanity Check:** one live `attend-queue` exercise: attention view lists ≥1 escalated +
  untriaged item in ONE view during the attended pass (acceptance criterion), captured in the
  session report.

Deferred acceptance criterion (explicit, with trigger): "second-repo adoption = one binding PR
\+ three sessions; zero plugin edits" is verifiable only when a second repo adopts — deferred to
that trigger (Brief "Deferred questions": multi-repo scale-out, USER-RESERVED).

## Blast radius

HIGH — new org-wide convention constraining all future lanes; new plugin with hooks + user-scope
statusline wiring (affects every session on the machine); autonomous PR authorship + (via babysit)
merges in an org repo with an org credential; multiple stress-test triggers match (infrastructure,
new conventions/enforcement, cross-plugin architecture, security-sensitive). Formal stress-test
REQUIRED.

## Stress-test summary

Two independent fresh-context rounds (2026-07-23):

- **Plan reviewer** (Step 3): 1 CRITICAL / 11 IMPORTANT / 5 SUGGESTION — all verified against
  sources and fixed in place. Clusters: rate-limit-guard seam doctrine (setup must not mutate
  user settings; operable floors must be inlined, not cited; dead userConfig; atomic tee;
  pause-window semantics) and unattended-termination realism (drain-terminal state, restart
  owner, compaction-durable loop state). Also: `health` verb didn't exist (→ `capabilities`),
  topic-docs contract kinds, naming grammar, Phase 6 filing ambiguity.
- **/devils-advocate** (Step 4): 2 CRITICAL / 3 HIGH / 6 MEDIUM / 2 LOW. Must-fix all applied
  or routed to operator: merge-authority contradiction with the autonomy matrix (→ Open
  Decision 5); babysit autopilot premise invalidated post-launch (→ Open Decision 6, babysit
  PLAN); `--bg` statusline tee verification gate (Phase 7); C3 self-classification hardening
  (Phase 4 path/topic gate + bot-author triage rule + HITL ratification period); quota
  economics surfaced (Open Decision 7). MEDIUMs absorbed as documentation-pass items across
  Phases 2/3/5. LOW #12 (squash-merge satisfies `required_signatures`) was VERIFIED LIVE
  during the stress-test — signing ruleset targets the default branch only, squash is the only
  allowed merge method, GitHub web-flow signs API squash commits; Phase 5's build-time
  confirmation retained as a formality.
- One research-iterate round sufficed; no unresolved CRITICAL/HIGH remains outside the three
  operator decisions.

## Execution shape

Phases are PR-boundary-sequential on the critical path; one genuine parallel opportunity:

| Phase | Surface | Basis |
|---|---|---|
| 1 | main-session | mechanical move + branch setup, trivial |
| 2 | main-session | judgment-heavy convention authoring; playbook-governed |
| 3 | main-session (+ fresh-context review sub-agents) | security-sensitive plugin; test-first scripts |
| 4 | main-session (+ fresh-context review sub-agents) | skill authoring, playbook-governed |
| 5 | main-session | small config PR + live race test needs two sessions (second is a sub-agent/worktree session) |
| 6 | main-session | operator-routed draft |
| 7 | main-session, operator present | attended triage + launch |

Dependency graph: 1 → 2 → {3, 4} → 5 → 7; 6 independent after 5's lease mechanism proves out;
babysit-PLAN B1/B2 parallel-safe with Phase 4 (disjoint plugin dirs), B3 joins at 7.
Wave option after Phase 2 merges: Phase 3 ∥ Phase 4 ∥ babysit B1+B2 (file-disjoint across
`plugins/rate-limit-guard/`, `plugins/work-items/`, `plugins/source-control/`) — material saving
if run as parallel sessions/agents; cost: 3 concurrent build agents multiply tokens. Default
recommendation: sequential 3 → 4 in one session (shared fresh-docs fetches amortize), babysit
B1/B2 in the operator's parallel session per succession note. Sequential fallback: if any
parallel build session reports scope-fence violation or cross-plugin edit, abort it and fold the
work into the main session sequentially.

## Open questions

None — all seven approval-round decisions resolved by the operator 2026-07-23:

1. Telemetry comment home → per-lane tracking issue in the TARGET repo (ci-workflows);
   convention doc records it at Phase 2. (Operator: "rest makes sense".)
2. Naming → renamed `hitl-queue` to `attend-queue` (verb-first) throughout the Plan.
3. Drain-terminal stop shape → adopted (modifies briefed exit condition).
4. Unattended launch surface → adopted: attended `/loop` first cycle, then `claude-ops:lanes`
   gated on the `--bg` tee probe — with the operator's coupling caveat encoded: lanes is a
   supporting, one-directional launcher; loop skills never depend on claude-ops (guarded
   if-installed, `/loop` fallback documented).
5. Merge authority → configurable autonomy ladder; shipped default = human merge for all but
   gate-proven C2-mechanical PRs (any author; bot authorship alone never qualifies, per the
   loop-lane convention); higher rungs (through full autonomy with frontier
   subagents driving to merge) opt-in via tracked-config flip = the matrix's recorded
   human-ratified promotion.
6. Babysit tier for ci-workflows → WORKER (autopilot premise invalidated post-launch).
7. Quota economics → frontier concurrency 1 + adaptive ceiling 2 for frontier items (general
   ceiling 3 for other tiers).

## Handoff to implementation

### User-approval gates

- Phase 5 phase-entry (a): any missing role label → STOP; github-iac Pulumi PR is operator work.
- Phase 6: draft PR routed to operator; no lane files or merges it.
- Phase 7 launch: operator present for triage; lane launch is operator-initiated.
- `[FALLBACK]` rows in the decisions table (presentation) — confirm or override before the
  affected phase starts.

### Execution shape ([EXEC-SHAPE] tagged)

- [EXEC-SHAPE] Sub-topic promotion DECLINED: PR-boundary phases already give independent commit
  boundaries; the two-track split (two PLAN files) is the operative decomposition; a third level
  would scatter the contract. Evidence: topic-docs convention supports multi-file slices.
- [EXEC-SHAPE] #691 implementation shape: budget-hit emits restart-request telemetry + clean loop
  stop; restart is launcher/operator-initiated. Evidence: `claude-ops` lanes SKILL.md verified
  behavior — a running loop cannot `/clear` or restart itself.
- [EXEC-SHAPE] Lane prompts are one-line skill invocations stored via the `claude-ops` lanes
  `prompt_dir` seam (`.work`, #480 forward dependency) — no new prompt storage built.
- [EXEC-SHAPE] Phase ordering 3 → 4 sequential in one session; babysit track parallel per
  succession note. Scope fences = plugin directory boundaries listed per phase.

### Mechanical work

- One PR per phase (2, 3, 4, 5, 6); commit convention per claude-code-plugins norms; topic docs
  ride build branches; phase tags advanced in PLAN.md as each completes.
- Verification checkpoints: repo CI (`ci-status`) green per PR; `skill-quality:check` per new
  skill; fresh-context review sub-agent before each PR (producer ≠ critic).
- Close-out at final PR: `/planning:plan close-out` (publish PLAN to PR, graduate ADR candidates
  — the loop-lane convention itself is the durable doc; prune the contract slice).
```
</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_0169XBydnqC5S6bkHDz1TDwL

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant