feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)#356
Conversation
The T5 convention as tool-agnostic normative text: three-layer data model with the never-estimate rule and its conditional revisit trigger; v1 record schema (join key by pointer to the telemetry contract, six-band ordinal scale, presence rules, partial-by-counterfactual composition); asynchronous attestation lifecycle with requester-less routing; marker-keyed structured-comment floor + native-field branch; bot-authored record integrity with platform-actor attestation evidence; attestation-preserving dedupe and write rules; query-side join with attestation-rate health signal and requester/independent split; telemetry-is-not-return boundary. Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>
… (WP3 phase 2) Capture slice in setup: tracker-class + close-flow discovery, close-triggered unattested-record wire with async attestation request, native-field branch where entitled with the marker-keyed comment floor as universal fallback, work-items comment-seam routing (upsert layered on its append-only surface, never a parallel writer), advisory paths for GUI-only/entitlement-gated surfaces with metered-minutes caveat, attestation-owner routing for requester-less classes, additive capture section in the binding. Close-boundary templates carry the two canonical questions verbatim and the attestation-preserving write and upsert rules. Four new eval cases incl. the never-estimate refusal and refired-close idempotency; v0.3.0 manifest + README + catalog. Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>
The fleet binding doc merged in melodic-software/standards#196. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 14s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 026d67d5a0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Address the review wave: a bare reaction carries no counterfactual or effort band, so only a reply with both values attests (reply missing a value is re-prompted, never upserted); and the work-items tracker binding offers provider-specific comment mechanics, not a race-safe comment seam — the marker upsert and attestation-preserving dedupe are this contract's own obligations on either path. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 38s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 47c0e1f3eb
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…source key Address the review wave: the close trigger now fires only for items carrying the binding-recorded autonomous-work discriminator (role label or bound-automation closure) — the convention's capture scope — with the guardrail matrix named as the refinement path; and the attestation source citation becomes a schema key (attestation_source, normalized reply-event URL) on the contract's record schema and both surfaces. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 48s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: db3d73bcfb
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The telemetry strip rule is the work-item join key's normalization; comment-event URLs identify the event in query or fragment, so the citation stores the platform-served URL unstripped. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 1m 49s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c3ab1f675a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…h the owner Address the review wave: bot closure alone no longer admits an item to return capture — the autonomous-eligible role label is the discriminator; and the attestation upsert requires the reply's platform actor to be the item's accountable human (requester or standing owner), so any other participant's reply never attests. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 1m 57s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 29e2c1f9cf
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Re: latest claude[bot] review (commit 29e2c1f) — classifying the 5 remaining issues:
Fix commit follows in a moment; will reply here with the SHA. |
… schema Address the review wave: the capture trigger now requires BOTH the autonomous-eligible role label AND closure by the bound automation identity (label alone is a pre-work eligibility marker, not proof of autonomous execution); the capture section of the binding gains a concrete key schema (tracker_class, record_surface, automation_identity, routing) so the actor check and requester-less routing have somewhere deterministic to read from. Also: the return-capture template now states the marker and JSON blocks form one tracker comment, the return-accounting convention's dangling "Attestor" line is reflowed and gains a note on the binding-vs-record schema_version token difference, and evals.json adds the native-fields capture-path eval and its missing trailing newline (also the CI hygiene/ci-status failure this fixes). Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 5m 35s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)
|
|
Fixed in
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: efadf01f4c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… fix vendor-token gate Address the follow-up review wave on efadf01: name label-plus-automation- actor closure as a proxy for execution evidence (not proof), deferring a first-class dispatch/execution-provenance signal to the guardrail matrix rather than inventing one ad hoc; key the capture binding's `routing` object by the bound tracker's recurring-schedule row id (or an interview-persisted identifier absent that binding) instead of an undefined class name; and state that an attestation-exempt class's close trigger posts no record and no request at all, per return-accounting.md's no-perpetually-unattested rule. Also fixes the `plugin-gate` contract-check regression efadf01 introduced: the vendor-token ban on autonomy's reference/ docs matched the literal ".claude" path in the schema_version note; reworded without the path. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 28s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dcded56a66
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Address the follow-up review on dcded56: the capture slice only wired the close trigger (posts the unattested record), with no described mechanism for the reply that attests it — trackers like GitHub don't re-fire a close workflow on a later comment. Adds step 3 (WIRE the reply-triggered attestation handler where machine-editable) and a companion "Attestation-reply trigger shape" section in return-capture.md, performing the same marker-keyed upsert from its own trigger surface; advisory where no reply-triggered surface is machine-editable. Eval 9 updated to cover the new wiring step. Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 54s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 648828a7cf
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…apture An automation-performed not-planned closure of a labeled item passed both gate predicates and would post a record asserting autonomous completion of undone work. The capture gate gains a third predicate: the closure outcome must be completed/delivered; not-planned, cancelled, and duplicate closures never capture. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 29s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bfd6620572
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…m the automation An item-author requester_source on a bot-filed item (or a bot standing_owner) could snapshot a machine identity as attestation_owner, letting the automation attest its own record and bypass the never-estimate rule. The contract now requires the resolved owner to be a human platform account distinct from the bound automation identity; a non-human resolution produces no owned record - route to the class's standing owner, else that item's capture stays advisory. Setup validates declared standing owners the same way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 1s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d7e81b6363
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…est recovery The reply-trigger shape's native branch checked creation only; it now requires the audit trail to attribute every subsequent record-field revision to the bound automation (an altered attestation_owner rejects the record before the snapshot is trusted). The native-field attestation_request citation gains the same idempotent recovery as the comment floor: a later automation pass locates its own bot-authored request comment and fills the missing field, or re-posts the request, never touching attestation fields. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 43s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c6f0fec707
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…it-trail path Step 3's concrete handler wiring checked creation only, while the contract and template already require every subsequent record-field revision to be automation-authored; the setup step now carries the same check - a record with any non-automation revision is rejected before the owner snapshot is trusted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 42s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2fd859932b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Any participant can post the public marker text; the reply path's lookup could select the foreign comment (uneditable) and fail a valid attestation while the real bot record stayed unattested. The lookup now filters by the bound automation identity's authorship, matching the record-integrity rule the close path and consumers already apply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 2m 56s —— View job PR Review: feat(autonomy): return-accounting convention, capture slice, live-proven demo (WP3)Todo List
What
|
…ils, routines, runner design pack) (#359) ## Summary Architect-round Plans for the four remaining ai-adoption-ladder work packages, on the merged WP1-WP3 foundations (#330, #343, #356): - **WP4 trigger-dispatch** (`docs/topics/ai-ladder-wp4-trigger-dispatch/PLAN.md`): T1 contract — four signal-surface classes (`tracker-vcs-event`/`temporal`/`agent-internal`/`channel-feed`), six adapter obligations, signal envelope (`signal.*` marker record incl. `signal.work_class` stamped from the security surface), kick + drain dispatch through the one work-item queue entrypoint, `check-signal-envelope.mjs` enforcement surface, scratch-repo demo. - **WP5 guardrails** (`ai-ladder-wp5-guardrails/PLAN.md`): T2+T3+#241 — matrix hub + 4 leaves, two-surface binding split (security binding in settings-as-code home, `executor_class`, admission decision table + classification rules), `check-security-binding.mjs` with eval-mode promotion-ceiling semantics, probe-before-bind isolation slice, free-tier CI positive probe + fail-closed demo. - **WP6 routines** (`ai-ladder-wp6-routines/PLAN.md`): T7 — catalog hub (39 classes) + 10 v1 leaves (dead-code sweep excluded per D4 correction), contract-owned catalog-to-matrix mapping rules, setup slice with detect-diff-reconcile of existing schedulers, novel-class derivation probe. - **WP7 runner design pack** (`ai-ladder-wp7-runner/PLAN.md`): T4 graduation — spine + 8 seam specs, lifecycle state model, two-family stop criteria, terminal-handoff escalation with severity/ack knobs, topology seam map. ZERO build artifacts; build stays trigger-gated; birth decisions USER-RESERVED. Merge order pinned: WP4 before WP5 (absent admission binding fail-closes to human-gated — no ungoverned window), WP6 after both, WP7 last. Review: WP4+WP5 stress-tested by a fresh-context reviewer (14 findings, all folded — headline: work-class classification obligation + `executor_class` on the security binding). WP6+WP7 stress-tested by a second fresh-context pass (10 findings, FIX-THEN-SHIP both, all folded — headline: routine→class mappings land on the WP5 security surface, never the repo-local agent-writable section; D2 gains a hybrid DET+AGT branch, a `not-a-routine` catalog flag, and provenance/structural axes reaching C4/C5; WP7's escalation severity refinement adds the WP5 schema+validator edits it requires as genuinely additive optional keys, C5→L3 cited from the matrix cell, vendor-hosted merge cap restated on the cloud-backend path). No linked issue ## Related - #330 (contract set + WP1-WP3 plans) - #343 (WP2 telemetry), #356 (WP3 return accounting) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 (1M context) <noreply@anthropic.com>
…elope check (WP4) (#372) WP4 trigger-dispatch package (T1), per the merged PLAN (all four phases complete on this branch). No linked issue ## Related - #359 (WP4–WP7 plans, incl. this package's PLAN as merged) - #356 (WP3 return accounting — the marker-record precedent the envelope serialization follows) - #343 (WP2 telemetry — the trace-tree obligation this contract imports) ## What ships - **`plugins/autonomy/reference/trigger-dispatch.md`** (new contract): four signal-surface classes (`tracker-vcs-event` / `temporal` / `agent-internal` / `channel-feed`), two recorded attributes (provenance, transport incl. `push-lifecycle` expiry fail-close + poll-detector backstop), six class-generic adapter obligations (normalize+enqueue only; layered idempotent dedup with create-then-reconcile and a live-duplicate drain guard; provenance + durable raw link; traceparent propagation; admission enforcement at the seam with absent-binding fail-close; closed-loop ack), security-surface work-class stamping with verified `agent-internal` parent provenance and higher-of class resolution, the schema-versioned signal envelope (marker `<!-- autonomy:signal:v1 -->`), kick + drain dispatch with the one-entrypoint invariant and execution-surface attestation, executor surface classes with security-owned `executor_class`. - **Setup trigger/dispatch slice** (`skills/setup/SKILL.md`): discovery-first per-class surface interview, DIY-floor kick + hourly drain as reviewable changes, plan-gated surfaces advisory with cost, execution-surface recording with the attestation caveat stated, fail-closed absent-admission wiring, additive `triggers` binding section (shared `surfaces` map shape that later scheduling sections reuse). - **Templates**: `templates/trigger-adapters.md` (per-class adapter shapes, six obligations inline), `templates/ack-reply.md` (one-line ack, item URL first). - **Enforcement**: `scripts/check-signal-envelope.mjs` — validates the envelope marker record (required keys, enums, W3C traceparent, agent-internal `signal.parent_item` as normalized canonical URL, deterministic raw-link origin branching via the binding's surfaces maps), exit contract 0/1/2 mirroring `check-emission-conformance.mjs`. - Five new setup evals; plugin 0.4.0; README + catalog. ## Conforming-path demo (scratch repo, live) Run on `kyle-sexton/autonomy-demo-scratch` (private throwaway) with the work-items capability's tracker seam installed (`tools/work-item-tracker/` copied per its documented install path) and a repo-local binding carrying the `triggers` section. No admission binding present → the fail-closed path is what the demo proves. No paid service touched. 1. **Stub signal**: label `autonomy:trigger` applied to source issue [#2](kyle-sexton/autonomy-demo-scratch#2); labeled-event id `28172185501` fetched from the issue timeline (surface-native unique event id → `signal.identity`). 2. **Dedup pre-search** for the identity: `[]` (no existing item). 3. **Enqueue via the seam** (`work-item-tracker.sh create-item`): queue item [#3](kyle-sexton/autonomy-demo-scratch#3) created carrying the full envelope — class/transport/provenance, identity, raw link (`https://api.github.com/repos/kyle-sexton/autonomy-demo-scratch/issues/events/28172185501`), traceparent `00-dab5c21d6389b65c950f6b13f14af7e3-de1b63c53960913b-01` (the trigger-hop trace context observed on the enqueue record), NO `signal.work_class` (no classification rules bound → unclassified), label `needs-human`. 4. **Create-then-reconcile re-search**: exactly one item (#3) matches the identity — no duplicate to close. 5. **Ack echo** on the source surface: [comment on #2](kyle-sexton/autonomy-demo-scratch#2 (comment)) — `Queued as kyle-sexton/autonomy-demo-scratch#3 (autonomy: tracker-vcs-event signal)`. 6. **Envelope conformance**: `check-signal-envelope.mjs <item-3-body> --binding .claude/autonomy/binding.json` → `Signal-envelope conformance OK: 1 envelope(s) checked.` exit 0. 7. **Drain claims via the seam lease** (autonomous drain mode, the ONE entrypoint): `work-item-tracker.sh claim github:kyle-sexton/autonomy-demo-scratch#3` → lease acquired (`holder: kyle-sexton`, lease comment `5014153239`), exit 0. 8. **Fail-closed refusal, audited live**: admission check before execution found NO admission binding → autonomous execution REFUSED, [refusal comment on #3](kyle-sexton/autonomy-demo-scratch#3 (comment)), claim released; final state `OPEN`, labels `["needs-human"]`, assignees `[]`. This also exercises the unclassified → human-gated path (no security binding → no `signal.work_class`). No second dispatch path was exercised anywhere: the event kick and the drain both funnel into the seam lease. ## Near-duplicate audit The dispatch entrypoint composes the work-item queue capability's EXISTING autonomous drain mode and race-safe seam lease — no second claim or dispatch mechanism was created anywhere in this package. The adapter templates normalize and enqueue only; the drain-side live-duplicate guard and the create-then-reconcile rule reuse the queue's own primitives. ## Gates - `scripts/validate-plugins.sh` exit 0; `node scripts/validate-plugin-contracts.mjs` exit 0 (vendor deny-list clean over `reference/`); `node scripts/generate-catalog.mjs --check` in sync. - `claude plugin validate plugins/autonomy --strict` pass; skill-quality `check-skill.sh setup` PASS (2 pre-existing WARNs: line soft-target, gotchas surface); `evals.json` valid against the bundled schema (ajv draft2020). - markdownlint 0 errors; typos clean; lychee (offline) 0 errors. Full plugin test suite rides CI (local run exceeds the shell timeout; prior-session finding). - Checker exit contract verified with fixtures: no-args → 2 + usage; valid tracker/temporal-local → 0; temporal without `--binding`, ci-cron origin with `file:` link, unrecorded surface, agent-internal missing `signal.parent_item` → 1 each. ## Boris check The 2→3 cell ("break up your work into loops and routines; let Claude kick off Claude") is instantiated as governed queue + kick; the step-3 trap is honored (audit trail before scale — envelope, ack, lease, refusal are all tracker-resident audit records; no agent-count machinery); the step-1 guardrail is untouched; the telemetry trace tree is the kick-off audit trail (traceparent on the enqueue record). Fresh-docs mandate: GitHub Actions events reference and the Claude Code headless reference re-fetched this session before the SKILL.md edit (cited in the Phase 2 commit). <details> <summary>PLAN.md (as merged, tags advanced)</summary> See `docs/topics/ai-ladder-wp4-trigger-dispatch/PLAN.md` on this branch — all four phase tags [DONE], decisions table and stress-test summary carried in the file. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…rule, snapshot exit-code taxonomy, cross-PR dependency channel (#456) Closes #373. Implements the four enhancements plus one gotcha from issue #373 in the `babysit-prs` skill. The source retrospective the issue cites (`.work/handoffs/babysit-autopilot-retro.md`) is not present in the checkout (`.work/` is gitignored), so every edit was driven by the issue body plus the live skill files; the retro is provenance only. All doc edits are pointer-style (they point at the existing rule/clause, never restate it). ## (a) No-monitor hard STOP at the merge / gate-completion step Added in both `SKILL.md` (Guarded mutations) and `reference/safety.md` (Pinned-Command Degradation): once a PR is proven ready, or its merge is deferred to a human, report and stop — never arm a watch/monitor. Both point at the existing no-background-monitor clause in the Worker Contract (`reference/orchestration.md`) rather than restating it. ## (b) Bare-wrapper pipe rule `reference/safety.md` Guarded Mutation Wrappers now states that the wrapper's JSON must be parsed in a separate step — never `wrapper | python` / `| jq` — because an interpreter-in-pipeline trips the auto-mode safety classifier and blocks the call before the wrapper runs. ## (c) Snapshot exit-code taxonomy `scripts/pr_queue_snapshot.py` previously collapsed every non-fatal error into exit `1`, so a purely-advisory head-ref alias-check error was indistinguishable from a substantive per-PR hydration failure. The split: | code | meaning | state written | | --- | --- | --- | | `0` | valid snapshot, no errors | yes | | `1` | valid snapshot, ≥1 substantive error (per-PR hydration or discovery failure) | yes | | `2` | fatal — exception before a snapshot existed | no | | `3` | valid snapshot, advisory-only (head-ref alias cross-check) errors | yes | Substantive errors take precedence over advisory (a run with both returns `1`). The advisory marker is a single-source constant (`babysit_delta.HEAD_REF_ALIAS_ERROR_MARKER`) used at both message-construction sites (queue path in `babysit_delta.py`, single-PR path in `pr_queue_snapshot.py`) so they cannot drift. Taxonomy documented in the module docstring; new `scripts/tests/test_pr_queue_snapshot.py` covers 0/1/2/3 including the advisory-vs-substantive distinction and precedence. **Consumers checked** (exit codes are a caller contract): every `*.md`, `*.sh`, `*.py`, `engine.test.sh`, orchestration doc, and `bin/` wrapper referencing the snapshot. No live consumer keys on the script's `$?` — the docs (`orchestration.md`, `SKILL.md`, `loop.md`) invoke it and parse the JSON snapshot; `babysit-python.sh` uses `exec` (child exit code passes through unchanged); no `bin/` wrapper or sibling script branches on it. Introducing `3` therefore breaks no caller that treated `1` as "any error". ## (d) Worker→main cross-PR dependency channel New `Cross-PR Dependency Signalling` section in `reference/orchestration.md` formalizing a worker signalling a discovered cross-PR coupling (e.g. medley#1567 ↔ #1597, #224 ↔ #356) back to the main agent, which owns cross-PR ordering. Points at the existing `SendMessage` mechanism (documented there for main→worker), explicitly reversed. ## Gotcha: self-blocking CI check Recorded in the `SKILL.md` Gotchas list: a newly required check whose own fix PR carries that same check cannot be gate-merged and needs a one-time human admin-merge bootstrap. The concurrent-same-identity yield rationale is already covered by the Checkout/Push invariants (head-SHA recheck + worker lease + `foreign_activity` suppressor) and is deliberately not duplicated. ## Verification - `bash engine.test.sh`: `Ran 186 tests … OK`; ruff `All checks passed!`; guarded-wrapper behavior checks all PASS. - markdownlint-cli2 on the four changed `.md` files: 0 issues. - Plugin bumped `0.9.0` → `0.9.1` (`jq -e .version` confirms) with a matching `## [0.9.1]` CHANGELOG entry. - No trailing whitespace / final-newline issues in any changed file. ## Related - #322 (PR-B) — retro source for these enhancements. - #324 — sibling review-trigger reconciliation (F8 posting-guard follow-on territory). - #326, #327 — same babysit-prs safety-hardening cluster. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Summary Effort-end housekeeping for the AI-adoption-ladder effort: all seven work packages are merged, so the per-WP topic slices come out per the close-out convention (publish the PLAN, then prune the slice). - WP4–WP7 PLANs were already spliced into their PR bodies at each close-out (#372, #377, #600, #676). - WP1–WP3 PLANs are now published the same way into #333, #343, and #356 (each merged PR body gained a `PLAN (published at effort-end close-out)` details block), so every slice's content has a durable home on its merged PR before the directory is removed. - The WP map in `docs/topics/ai-adoption-ladder/index.md` now records delivery PRs instead of live slugs. The effort hub itself (index, design evidence, research docs) stays. No content is lost: every deleted PLAN.md is byte-recoverable from git history and readable in its merged PR's details block. ## Related No linked issue (effort-end housekeeping; the WP map issue #239 is already closed). - #239 (WP map: AI adoption ladder — the effort this closes out) - #333 / #343 / #356 (WP1–WP3 delivery PRs, PLANs published into their bodies in this close-out) - #372 / #377 / #600 / #676 (WP4–WP7 delivery PRs, PLANs published at their own close-outs) Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ger row (#698) ## Summary D1 sweep for #353: the WP1–WP7 PLAN out-of-package notes (mined from the merged PR bodies #333/#343/#356/#372/#377/#600/#676) contained 31 deferred notes; every note that named no tracking issue and had no contract-recorded trigger now has a durable home: - **README roadmap** gains four trigger-gated rows: fleet guardrail materializations (WP5), fleet routine stand-up + existing-scheduler reconciliation (WP6), vendor-binding capability templates (WP6), cost enforcement / hard spend caps (WP5, 3→4 trigger). - **Trigger register** gains the second-binding-consumer cross-repo drift check (WP1). - **`reference/return-accounting.md`** records the per-work-class precision-graduation deferral beside its band-stability rule (WP3). - **CHANGELOG.md created** for the autonomy plugin (first entry; 0.1.0–0.7.0 history stays in the WP PR bodies) + version bump to 0.7.1 — starts the CHANGELOG-parity posture #663 gates on. All other mined notes were already covered: tracked issues (#351, #352, #694–#697), contract-recorded triggers (telemetry immutable-ID, return-accounting expansion/revisit), delivered sibling WPs, resolved `/architect` questions, or the user-held dormant triggers (runner build T4, L3 backend, merge serialization, mid-run interrupt, org-enablement, cross-team). Full disposition table lands as a comment on #353 at close. Documentation only — no contract semantics change. ## Related - #685 (effort-end prune that moved the PLAN records into the merged PR bodies) - #663 (CHANGELOG-parity gate this PR's new CHANGELOG.md starts satisfying) - WP delivery PRs mined: #333, #343, #356, #372, #377, #600, #676 Closes #353 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

Summary
WP3 return-accounting package (T5), on the merged WP2 telemetry contract (#343):
plugins/autonomy/reference/return-accounting.md: three-layer data model (machine/deterministic, human-attested, agent/LLM), the two Boris questions verbatim as the human fields (counterfactual would-have-done-anyway + manual-effort band), six-band effort scale<1h / 1-4h / 4h-1d / 1d-1w / 1w-1mo / >1mo(user-locked), capture at the task boundary (work-item close / PR merge) as tracker-resident marker records, autonomous-class scope only, agents never estimate the return fields.Live demo evidence (executed this effort)
Throwaway private repo
kyle-sexton/autonomy-demo-scratchissue #1: unattested marker record written at close; a re-fired close stayed idempotent (count 1, no downgrade); attestation upsert recorded with platform actor + source citation; DuckDB join against the WP2 committed OTLP fixture matched the join key string-identically (autonomy.work_item.url).Near-duplicate audit: the capture slice composes the work-item tracker seam and the WP2 telemetry contract — no second telemetry or tracker surface created.
No linked issue
Related
🤖 Generated with Claude Code
PLAN (published at effort-end close-out)
ai-ladder-wp3-return-accounting
Brief
TLDR
Return-accounting convention (T5): a lightweight, tracker-resident record captured at the task
boundary of autonomous-class work, answering Boris's two return questions (counterfactual +
manual-effort band), joined to machine cost telemetry by the WP2 work-item attribute. No new
capability, no new cost; agents prompt and aggregate but never estimate return fields.
Goal
Any adopting org can capture return data — would-have-done-anyway plus manual-effort band — at
every autonomous-class task boundary, as a tracker-resident record joinable to its cost
telemetry by work-item ID, via guided setup, with zero paid dependencies.
Locked decisions
autonomy.work_item.url, record fieldwork_item_url— never a tracker-short ID; the Plan's Phase 1 schema governs.Five contiguous bands —SUPERSEDED 2026-07-18 (stress-test evidence, user-locked): six contiguous bands —<1h,1–4h,4h–1d,1d–1w,>1w<1h,1-4h,4h-1d,1d-1w,1w-1mo,>1mo— the open>1wtop band erased the largest avoided-effort signal; the Plan's Phase 1 schema is the governing serialization. Deliberate deviation from the T5 draft (<1h / 1–4h / 1–2d / 1w+), which left 4h–1d and 2d–1w unmapped. Org-custom bands rejected: breaks cross-org aggregation (same argument that rejected per-adopter reverse-DNS in WP2 D3).Constraints
machine layer informs the cost side of the join only.
migrations, never per-org variation.
convention-before-capability, no step-skipping, trust before scale.
Acceptance criteria
architect fills mechanics without reopening the contract.
contract required.
return record exists → joinable to that item's cost telemetry by work-item ID at the sink.
telemetry surface.
paid dependencies.
Captured assumptions
table; verbatim capture is durable in design-threads T5. The table doc (now on disk:
docs/topics/ai-adoption-ladder/design/boris-step-and-your-role.txt, source Google Doc linked fromwayfind issues Map: AI adoption ladder — climb 2→3 (plugins fleet) #239–Unify session and CI verification telemetry? #247) carries the related "is this something an engineer would have
done?" variant — the thread wording is canonical for the prompt.
re-anchor copy.
The five-band contiguous scale is a deliberate correction of the T5 draft's gaps, lockedSUPERSEDED 2026-07-18: the contiguousunder the round's recommendation pre-authorization.
scale is now SIX bands per the D4 supersession above; migration authority for any further
band change is a reviewed contract migration, governed by the Plan's Phase 1 schema.
new instrumentation.
Out-of-scope (deferred with triggers)
build lands.
(conditional constraint, revisit then).
Deferred questions
/architect./architect./architect./architect.Plan
Consumes the WP2 join key as-is: attribute
autonomy.work_item.url, value = the item'scanonical web URL (interview-locked in the batched WP2 round). Live-verified tracker facts
grounding the binding catalog (2026-07-18): GitHub now ships native issue fields
(org-managed, issues-only, REST+GraphQL); Jira custom fields on all plans; GitLab work-item
custom fields since 17.11; Linear has no custom fields by design — so the structured comment
remains the universal floor, native fields are the where-supported branch (D3 confirmed, not
reopened). No external effort-band standard exists to cite; bands stay contract-defined (D4).
Prerequisites: WP1 implementation merged; the WP2 package PR merged (this convention doc cites
the telemetry contract — dead-link avoidance). Fresh branch after both.
Phase 1: Return-accounting convention doc [DONE]
plugins/autonomy/reference/return-accounting.mdschema_version(string, from v1 — hook-envelope precedent),work_item_url(the join key; value contract defined by the telemetry contract'sautonomy.work_item.url— pointer, not restated),attested(boolean),counterfactual(yes|no|partial),effort_band(six contiguous ordinal tokens:<1h,1-4h,4h-1d,1d-1w,1w-1mo,>1mo— user-locked 2026-07-18 on stress-test evidence, superseding D4's five-band set: the open>1wtop band erased the largest avoided-effort signal; serialized as those exact strings, ordinal order defined by the contract, never lexical),attested_at(ISO 8601 UTC),attested_by(the attesting human's platform identity, captured from the attestation action),attestor_role(requester|reviewer|maintainer|other). Composition rule:effort_bandanswers the manual-cost question for the WHOLE delivered item regardless ofcounterfactualvalue;partialqualifies the counterfactual only — aggregation derives avoided cost from the pair, the attestor never prorates. Presence rules: an unattested record carriesattested: falsewithcounterfactual/effort_band/attested_at/attested_by/attestor_roleABSENT (never null-imputed); an attested record carries all fields. Record lifecycle — attestation is ASYNCHRONOUS by construction (autonomous-class work has no human at the close boundary): the machine posts the unattested record at close plus an attestation request routed to the accountable human; attestation later upserts the same record toattested: true. A never-attested record stays visible as unattested. Attestation routing for requester-less classes (standing routines, scheduled sweeps): the binding declares a standing attestation owner per class, or marks the class attestation-exempt with its cost reported separately — never a perpetually-unattested default. Record integrity: a conforming record is authored by the deployment's bound automation identity; consumers MUST ignore marker-matching records from any other author; attestor identity derives from the PLATFORM actor of the attestation action — on the comment floor the upsert itself is bot-authored, so the record'sattested_byMUST be copied from, and the record MUST cite, the attestation source event (the human's reply/reaction whose platform actor answered) — the self-declaredattestor_rolefield is descriptive, never the trust anchor. Aggregation guidance: report the attestation rate as a first-class health signal (a collapsed rate invalidates the dataset as promotion evidence) and separate requester-attested from independently-attested rows (self-attested counterfactual is a conflict of interest). Duplicate tolerance: the standalone capture path's find-then-create has an inherent create-create race — dedupe on read is ATTESTATION-PRESERVING: an attested bot-authored record outranks any unattested one; only among equally-attested records does the latest win. Write rule, same property: the close trigger creates the unattested record only when no marker-matching bot-authored record exists — a re-fired or retried close NEVER overwrites or downgrades an existing record's attestation fields. Capture point: task boundary (work-item close / PR merge); scope: autonomous-class work per the T3 class vocabulary. Tracker binding seam: record surface resolves per tracker class — native fields where the class supports them (org-managed issue fields, project-scheme fields, work-item fields), structured comment as the universal floor: hidden marker<!-- autonomy:return-accounting:v1 -->+ one fenced JSON block, marker-keyed upsert (find-then-edit, else create — established in the wider bot ecosystem; NEW to this repo's seams). Prompt: Boris's two questions near-verbatim ("Would you have spent engineering effort on this anyway?" yes/no/partial; "What would it have cost in manual eng-hours?" band), exactly two fields, non-blocking, explicit skip affordance. Join is query-side by the join key only against cost telemetry (resource-scoped on agent-session signals per the telemetry contract); cost values never duplicated into the tracker record (D6). Zero vendor/fleet names; tracker-product specifics live in SKILL.md.Sanity Check:
grep -c 'autonomy.work_item.url' plugins/autonomy/reference/return-accounting.md≥ 1grep -c 'return-accounting:v1' plugins/autonomy/reference/return-accounting.md≥ 1grep -c -- '1w-1mo' plugins/autonomy/reference/return-accounting.md≥ 1 (spot token) and the schema block lists exactly 6 bandsgrep -ci 'attestation rate' plugins/autonomy/reference/return-accounting.md≥ 1 andgrep -ci 'bound automation identity' plugins/autonomy/reference/return-accounting.md≥ 1 (integrity + health-signal clauses present)grep -ci 'estimate' plugins/autonomy/reference/return-accounting.md≥ 1 (never-estimate rule stated)node scripts/validate-plugin-contracts.mjsexit 0; lychee passesPhase 2: Guided-setup capture slice [DONE]
First work item — fresh-docs mandate (repo CLAUDE.md): re-fetch official skills/hooks docs
before editing SKILL.md; re-verify tracker API surfaces cited in templates at implementation.
plugins/autonomy/skills/setup/SKILL.mdplugins/autonomy/skills/setup/evals/evals.jsonplugins/autonomy/skills/setup/templates/(WP2's snippet home)plugins/autonomy/README.mdplugins/autonomy/.claude-plugin/plugin.jsonSanity Check:
/skill-quality:check+validate-evalspass;claude plugin validate --strictexit 0grep -rc 'return-accounting:v1' plugins/autonomy/skills/setup/templates/≥ 1grep -rc 'engineering effort on this anyway' plugins/autonomy/skills/setup/templates/≥ 1 (canonical question verbatim)grep -c 'attested: false' plugins/autonomy/reference/return-accounting.md≥ 1 (async/unattested lifecycle stated)Phase 3: Conforming-path demonstration [DONE]
Acceptance-criterion probe, scratch consumer repo: close a demo work item → the
close-triggered capture posts the unattested marker-keyed record (comment floor) + attestation
request → attest via upsert → join to demo COST telemetry (the COMMITTED WP2 Phase 3 fixture's
agent-session cost metric — durable interface, no WP2 pipeline re-run) by
autonomy.work_item.urlin a DuckDB query. The fixture's bakedautonomy.work_item.urlvalueis SUBSTITUTED with the demo item's canonical URL as demo-input preparation before the join
query (the scratch item's URL cannot match a pre-committed value; substitution touches the
demo copy only, never the committed fixture). Zero paid dependencies.
Sanity Check:
<!-- autonomy:return-accounting:v1 -->with a parseable JSON record (re-running the capture upserts, count stays 1)attested: falsewith the five attestation fields absent; the post-attestation upsert showsattested: truewith all fields present,attested_bymatching the attestation source event's platform actorwork_item_urlwith an agent-session COST metric whose resource attributeautonomy.work_item.urlis string-identicalPhase 4: Gates [DONE]
Same in-repo gate roster as the WP2 package: validate-plugins, run-plugin-tests,
validate-plugin-contracts, markdown/typos/lychee,
claude plugin validate --strict, catalogregen check. Near-duplicate audit statement: capture composes the work-items close-flow seam,
never duplicates it; no capability estimates return fields.
Sanity Check:
node scripts/generate-catalog.mjsin-syncBlast radius
MEDIUM — ~7 files in one plugin, but the record schema is a new public contract consumed by
every adopter and by the WP2 join; new-convention trigger matches. Git-revertible; automated
gates cover shared surfaces.
Stress-test summary
Step 3 fresh-context plan review (shared with WP2): async-attestation lifecycle was
unspecified (no human exists at an autonomous close — fixed: unattested-at-close + async
upsert); marker-upsert correctly reframed as new-to-this-repo layered on the work-items
comment adapter; unattested field presence pinned; join derivation made pointer-only.
Step 4
/devils-advocate(shared): folded — record integrity (bot-authored records only,platform-actor attestor identity, self-declared role never the trust anchor), attestation-rate
health signal + requester/independent split, requester-less class routing rule, standalone
create-create race stated with attestation-preserving dedupe (attested outranks unattested;
latest wins only among equals; a re-fired close never downgrades), six-band set (user-locked,
supersedes D4 five bands), partial×band composition rule, committed WP2 fixture as the demo
interface. Residual accepted: comment-surface forgery is mitigated by authorship filtering,
not eliminated — native-field surfaces inherit platform ACLs and are the stronger branch where
entitled.
Execution shape
Fully sequential 1 → 2 → 3 → 4 — Phase 2 wires what Phase 1 specifies; Phase 3 exercises
Phase 2's snippet against the committed WP2 fixture; Phase 4 gates. Parallel saving
immaterial. Cross-package: this PR merges after the WP2 PR (contract citation + fixture
dependency); both wait on WP1 implementation.
Open questions
work-items flow after this package lands (out of package per D1).
Decisions made (gate-passed)
reference/return-accounting.mdreference/)<!-- autonomy:return-accounting:v1 -->+ fenced JSON record, marker-keyed upsertattestor_roleenumrequester|reviewer|maintainer|other(descriptive, platform actor is the trust anchor)No [FALLBACK] tags — every remaining item traces to the Brief, an interview lock, or a
verified stress-test fold.
Handoff to implementation
User-approval gates
reviewed contract migration — re-enter
/architect review, never inline.divergence (e.g. field entitlements changed) STOPs and re-surfaces.
Execution shape ([EXEC-SHAPE] tagged)
Sequential 1→4, all main-session (table above). PLAN.md phase tags advance in the same commit
as each phase; scratch-repo demo per Phase 3; divergence escalation applies to every phase.
Mechanical work
Commit per phase on the implementation branch (suggest
feat/autonomy-return-accounting);gates re-run in full at Phase 4; commits via the repo's commit conventions; PR body carries
the demo transcript + never-estimate/near-duplicate audit statements + this PLAN in a
<details>block at close-out.