Skip to content

Merge upstream round 4 through b84e0e362fac, preserving fork contracts - #275

Merged
HelloWorldSungin merged 24 commits into
mainfrom
fm/fm-upstream-sync-2026-09-08-round-4
Sep 14, 2026
Merged

Merge upstream round 4 through b84e0e362fac, preserving fork contracts#275
HelloWorldSungin merged 24 commits into
mainfrom
fm/fm-upstream-sync-2026-09-08-round-4

Conversation

@HelloWorldSungin

@HelloWorldSungin HelloWorldSungin commented Sep 14, 2026

Copy link
Copy Markdown
Owner

Final validation is complete at d5fdeb343d9eb905898ae37e6895efffbce2959b: all 239 scripts are covered, all 16 non-compliance checks pass, and only the intentionally red no-mistakes signature check remains.

This merges the pinned 17-change upstream prefix through b84e0e362face25f3dd8945297a3df1320d7668c into the fork parent 668b61f1c89dc89913fd4f288576039682f00770.
The measured parent merge base is 2e65d2e2482247c22db2ba7aeeb055eb7876fef1.
It follows the previous landed round at #271 and carries its divergence witnesses forward.
The tail after the pinned endpoint is excluded.
Firstmate subsequently instructed reconciliation of newly landed fork main 2026a110019a49ff99b68e420d65483d2c32ec0c from #272 through ordinary fork-parent merge 57223759337ff96a523976949aa75383c8672395; its planning and design-relaunch changes compose without a contract contradiction.
Firstmate also authorized a normal merge of fork main d3da3bdf56444538ebd812c41508893c4eb4470c, retaining #274 and its hung-delivery monitoring fix.
The upstream merge commit remains 38ea36ae57e557e07ac05e110c41e9eac1bdc141; the fork update merge is dc76f2ec8d6e78a1d9b4036575c1ed49f4a6d031.
No history was rebased or rewritten; exactly one new merge directly imports the pinned upstream endpoint.

Validation comparison

Revision Full tests Canonical lint Outer Herdr lab teardown
Fork parent 233 scripts, 0 failures, 33 gate skips Pass, pinned extended ShellCheck and actionlint Pass
Upstream parent 189 scripts, 4 failures, 25 gate skips Pass, pinned extended ShellCheck and actionlint Pass
First merge sweep 237 scripts, 1 failure, 31 gate skips One fixture-source SC2031 issue, corrected using the existing isolated-shell pattern Pass
Combined runtime result before test-layout split 238 scripts, 0 failures, 32 gate skips Pass: ShellCheck 0.11.0 extended analysis; actionlint 1.7.12, all 3 workflows Pass
Corrected full local sweep at 0b838c8b843aacd4b2ecc514c2c4098361d7719f 239 scripts, 0 failures, 32 gate skips Pass: pinned extended ShellCheck and actionlint Pass
Final CI at d5fdeb343d9eb905898ae37e6895efffbce2959b 239 scripts, 0 failures, 33 gate skips; exact disjoint 225 portable + 14 real Herdr coverage Pass in CI and full local canonical lint CI Herdr lane passed; all 14 scripts executed

Upstream failures are recorded separately: missing Ruby in fm-nm-test-contract and fm-test-run, the Nix PATH expectation in fm-on, and the installed harness-wrapper expectation in fm-remote-doctor.
The fork already carries the portable parser and environment-aware fixture behavior; all four corresponding fork witnesses passed in the combined sweep.
Tests use isolated fixture homes and named Herdr labs, preserving credential, browser, and prompt execution opt-ins.
A gate skip is disclosed as a skip rather than counted as executed live coverage.

The initial published CI run reached the unchanged ten-minute parallel-lane limit after passing nine scripts, and the combined watcher-triage script reached its unchanged 480-second bound while continuing to pass cases.
Completed CI measurements now balance the same 24 proven-isolated candidates across the two parallel lanes at about seven minutes each.
The 125 watcher-triage cases are partitioned into 89 core and 36 wait/decision cases with exactly preserved definitions in one shared helper; the complete suite now contains 239 scripts.
Both execution limits and all isolation eligibility gates remain unchanged.

Integration findings

Firstmate approved upstream default-on token-free guards and portable-serial membership while preserving protected per-test execution gates.
The inherited fork suppression of unkeyed blocked repaint conflicted with upstream.
Firstmate chose the documented explicitly keyed boundary; focused keyed/unkeyed/default-key and durable-retention witnesses pass.
The real Pi/Herdr hung-settlement regression passed on the combined branch, including named-lab cleanup and monitoring restoration.

Firstmate approved upstream progress identities and once-per-episode secondmate notification with fork semantic classification and 60-second idle / 1800-second unknown thresholds, plus bounded busy suppression.
The same two-checkpoint clock experiment on both exact parents and the merge produced check: rearm-resurface, retained the foreign row, and published no stall notification.
The interrupted setup watcher owed recovery; the fixture now seeds the progress observation through its owner, while separate executable tests retain first-observation and no-loss recovery controls.
The new upstream runner placement fixture now copies the fork timeout helper it requires.
The design marker assertion records the text-and-Enter channel used for exports, separately from literal launch bytes.

Upstream applicability

Upstream change Applicability and reconciliation
d8e2bb3 / kunchenguid/firstmate#3843 Adopt native Windows Bash invocation in Pi and OpenCode; preserve fork hooks and compatible Unix invocation. Retain fork runner timeout, shard count, and timing coverage while adding the Windows regression.
71ec401 / kunchenguid/firstmate#3870 Adopt executable teardown tests for squash-merged rebased branches; keep fork strict run-abort attribution and preserving-refusal cleanup. This fixture does not authorize squashing this sync.
29015a2 / kunchenguid/firstmate#3860 Adopt registered custom-check supervision as an additional cause in the shared predicate, preserving queued wakes as another cause and existing snapshot fields.
64d3905 / kunchenguid/firstmate#3883 Upstream now implements the fork local-root behavior for remote-parented homes. Adopt the shared implementation and retire the divergence with passing local peer-lock/malformed-binding/cycle witnesses. Keep fork fixture cleanup ownership.
cf7e2fa / kunchenguid/firstmate#3872 Adopt board listener repair, structured reconcile selection, evidence-only request handling, and landed-subject filtering. Preserve fork design completion gates and abandoned-child visibility. Reconcile selection must never be recorded as the user's answer.
1533f47 / kunchenguid/firstmate#3885 Adopt origin-less pooled spawn support and identity fallback; keep agy restrictions, authored-brief branch identity, event wiring, and watermark-failure preservation.
5592cb6 / kunchenguid/firstmate#3889 Firstmate approved upstream fm_live_gate: installed token-free guards default on, prompt-submitting guards opt-in, named Herdr isolation preserved. Keep fork-specific live families and fixtures.
6d396da / kunchenguid/firstmate#3891 Adopt task-marker primary-checkout test refusal and launch propagation, preserving isolated fixture operation and fork spawn policy.
d4eb228 / kunchenguid/firstmate#3842 Adopt bounded stale alarms for backlog-held work. Preserve the fork's live declared-wait routing and widening cadence; declarations and backlog hold revisions must reset the corresponding episode rather than share a stale key.
0b9f518 / kunchenguid/firstmate#3871 Adopt extension-registered provider resolution for the Pi supervision branch; preserve fleet model resolution and away-mode standby.
ffd2c89 / kunchenguid/firstmate#3943 Firstmate chose upstream progress identity/reset/once-per-episode behavior combined with fork semantic busy classification and 60s proven-idle / 1800s unknown thresholds. One owner, bounded busy suppression, and executable combined cases all verified. Record the previously unlisted divergence.
36fd955 / kunchenguid/firstmate#3952 Adopt deterministic export-DOM checks and dated Pi evidence while retaining compatible fork verification history/current guarantees; do not claim upstream evidence as a fresh local run.
3af74fe / kunchenguid/firstmate#3950 Adopt actor-specific queue eligibility and stale grant recovery; preserve queued-wake supervision without alarming main about branch-owned rows it cannot drain.
98b37d4 / kunchenguid/firstmate#3305 Adopt /usr/bin/stat for Darwin BSD formats across shared paths; retain Linux behavior and fork-only paths' compatibility.
72bfdd0 / kunchenguid/firstmate#3945 Carry attribution-off policy through every Claude launch in the fork's launch-command owner, retaining explicit runtime/harness bindings and additive event hooks.
891dc51 / kunchenguid/firstmate#3946 Adopt cadence-derived grace in auto-arm, watcher acquisition, and away-mode turn-end checks only. Preserve upstream’s explicit 300s strict-guard scope and the fork shared snapshot/guard accessor, overrides, and invalid-value fallback. Keep quiet allowance separate from watcher grace.
b84e0e3 / kunchenguid/firstmate#3904 Adopt runner owner leases, orphan handling, and launch pacing without weakening PID/PGID identity refusal. Preserve fork Lavish artifact-missing terminal classification and deleted-artifact retirement semantics.

Active divergence survival

Active entry Behavior evidence Verdict
Agy crew adapter tests/fm-agy-adapter.test.sh Pass
Pinned ShellCheck download retry budget tests/fm-lint.test.sh Pass
LLM quota sidecar tests/fm-quota-sidecar.test.sh Pass
Scout completion gate reopened by a firstmate steer tests/fm-send-strict.test.sh Pass
Watcher restart hand-over tests/fm-watcher-lock.test.sh Pass
Watcher stop-signal disposition tests/fm-backend-herdr.test.sh, tests/fm-watcher-lock.test.sh Pass
Run-progress wedge hold tests/fm-watch-triage.test.sh and tests/fm-watch-triage-waits.test.sh, tests/fm-daemon.test.sh (prior-round witness retained). Pass
Watcher live declared-wait routing and self-widening recheck cadence tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh and tests/fm-watch-triage-waits.test.sh Pass
Herdr pre-Enter footer read on a native working baseline tests/fm-backend-herdr.test.sh Pass
Remote job worker descendant reaping tests/fm-remote-job-orphan-reap.test.sh, tests/fm-remote-job.test.sh Pass
Bounded remote job stdin capture tests/fm-remote-job.test.sh Pass
Default per-script bound on every test sweep tests/fm-on.test.sh, tests/fm-test-run.test.sh Pass
Locale-independent test coverage comparisons tests/fm-test-run.test.sh Pass
Queued wakes remain a supervision requirement tests/fm-guard-stale-banner.test.sh, tests/fm-turnend-guard.test.sh, tests/fm-wake-queue.test.sh Pass
No-mistakes run attribution tests/fm-crew-state.test.sh, tests/fm-teardown.test.sh Pass
Definition-of-done owner carries this fork's ready-to-validate handoff tests/fm-trigger-validation.test.sh Pass
A preserving refusal withdraws the pending backlog close tests/fm-backlog-atomicity.test.sh Pass
A watermark-capture failure keeps the published task record tests/fm-spawn-dispatch-profile.test.sh Pass
Fleet snapshot per-task timeout and abandoned child work tests/fm-dashboard-backlog.test.sh, tests/fm-home-summary-refresh.test.sh Pass
Pi and OMP away-mode supervision standby tests/fm-omp-harness.test.sh, tests/fm-pi-watch-extension.test.sh Pass
Keyed decision repaint suppression tests/fm-watch-triage.test.sh and tests/fm-watch-triage-waits.test.sh; keyed, unkeyed, default provenance and durable-retention focused witnesses passed Pass
Secondmate queue-stall semantic thresholds tests/fm-wake-queue.test.sh Pass
Herdr presentation fixture ownership and cleanup tests/fm-backend-herdr-presentation-e2e.test.sh, tests/fm-backend-herdr-recovery-e2e.test.sh, tests/fm-test-fixture-cleanup.test.sh Pass
Fork-local no-mistakes compliance-gate event scope tests/fm-no-mistakes-required-gate.test.sh Pass
Live pull-request body refresh before the compliance gate tests/fm-no-mistakes-required-gate.test.sh Pass
Merge-proof contract: one divergence accepted, one retired tests/fm-pr-merge.test.sh Pass
Upstream tracking mechanism tests/fm-upstream-status.test.sh and unchanged TRACK owner. Pass
Repository-local validation evidence tests/fm-nm-test-contract.test.sh; retained .no-mistakes.yaml store-false setting. Pass
Upstream-read-only posture in shared tracked docs tests/fm-agents-hard-rules.test.sh Pass
GBrain per-home knowledge memory tests/fm-bootstrap.test.sh, tests/fm-brief.test.sh, tests/fm-gbrain-capture.test.sh, tests/fm-gbrain-readonly-e2e.test.sh, tests/fm-recall.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh Executed witnesses pass; opt-in skipped: fm-gbrain-readonly-e2e.test.sh
Fleet dashboard and agent-event instrumentation tests/fm-bearings-snapshot.test.sh, tests/fm-dashboard-browser.test.sh, tests/fm-dashboard-events.test.sh, tests/fm-dashboard-gbrain-ui.test.sh, tests/fm-dashboard-gbrain.test.sh, tests/fm-dashboard.test.sh Executed witnesses pass; opt-in skipped: fm-dashboard-browser.test.sh

Retired local-root divergence is covered by passing tests/fm-spawn-worktree-settle.test.sh and tests/fm-teardown-endpoint-safety.test.sh in the combined full sweep.

Contract-file reasoning

File Reason for the resulting contract
AGENTS.md Add the upstream reconcile-request record and owner pointer; retain all fork hard rules, design role, agy restrictions, private memory, and delegated delivery boundaries. Retain the newly landed fork planning-intake wording and design-profile trigger, together with upstream reconcile routing.
.agents/skills/bearings/SKILL.md Adopt listener proof, structured reconcile selection and landed-subject filtering; retain abandoned child visibility and the fork's complete four-section digest.
.agents/skills/bearings/assets/board-template.html Preserve selected option separately from the note so reconciliation cannot masquerade as a supplied answer.
.agents/skills/captain-hold-lifecycle/SKILL.md Route board-requested evidence reconciliation through its single owner without losing the fork's design completion gate or attributing evidence to the user.
.agents/skills/firstmate-coding-guidelines/SKILL.md Firstmate approved installed token-free guards by default and prompt-submitting guards by opt-in through fm_live_gate; retain fork source-aware lint, one-owner and project conventions.
.agents/skills/harness-adapters/references/harness/opencode.md Document native Windows Bash invocation; retain OpenCode's existing headless/passive limits.
.agents/skills/harness-adapters/references/harness/pi.md Document native Windows helper invocation without changing Unix startup or fork away-mode ownership.
.agents/skills/process-event-sources/SKILL.md Adopt runner lifetime/lease/pacing and PID-reuse uncertainty limits through owner pointers; preserve the fork's missing/deleted-artifact handling and explicit destructive-poll loss limitation.
.github/workflows/ci.yml Adopt installed token-free live guard coverage; retain the fork's eight serial shards, Node floor, and separate real-Herdr lane; refresh the parallel timing comment without changing either job limit.
bin/fm-supervision-lib.sh One supervision-need predicate includes both upstream registered custom checks and fork queued wakes; retain snapshot/guard accessors and quiet allowance.
bin/fm-guard.sh Warn only about rows the current actor can drain; keep pending-queue supervision and expose live branch-held rows as an advisory.
bin/fm-turnend-guard.sh Add registered-check diagnostics alongside queued wakes, and adopt cadence-derived grace only on the upstream away-mode path.
bin/fm-claude-stop-autoarm.sh Use the shared supervision predicate and upstream cadence-derived acquisition grace; export the same resolved grace to its arm child.
bin/fm-wake-lib.sh Adopt local remote-parent root behavior, per-actor grant validity/counts, progress-marker writes, and the shared cadence formula while retaining fork wake/lock safety.
bin/fm-watch.sh Combine progress episodes with the firstmate-selected semantic thresholds; preserve default stop-signal disposition, run-progress holds, caller-owned Herdr wait cleanup, and live declared-wait widening. Add backlog call lifecycle identity without a secondmate hot-path backlog read.
bin/fm-spawn.sh Adopt origin-less clean-copy handling and task-marker propagation for ship, scout, and fork design workers; keep the fork launch library, agy restrictions, authored branch identity, evidence watermarks and event hooks. Retain the new fork shared design-pin owner and pin reuse on relaunch; preserve upstream task markers and fork launch templates.
bin/fm-launch-lib.sh Carry upstream Claude attribution-off settings in the fork's existing launch owner rather than reintroducing an inline second template owner.
bin/fm-test-run.sh Refresh serial timing hints from 200 completed passing CI measurements while retaining earlier maxima and the existing isolation-proof hint; preserve eight shards, eligibility and deadlines. Add upstream live-gate coverage and primary-copy refusal; retain default per-script deadlines, eight shards, C collation, and fork-only test selection; rebalance the same proven candidates using completed CI measurements and classify both watcher-triage units in their existing family.
tests/lib.sh Add live capability gating and clear the worker marker in isolated fixtures; retain dashboard config/store/credential isolation.
docs/architecture.md Describe combined supervision/stall/wait behavior, design-worker isolation and local-root locks; do not reintroduce upstream's unfetched-head parked-run abort, which the ledger explicitly excludes.
docs/configuration.md Keep fork setting owners and thresholds; add upstream reconcile requests, runner lifetime limits, provider registration, task markers, and attribution policy.
docs/turnend-guard.md Cover registered checks and queued wakes together; preserve the explicit distinction between cadence-derived away/auto-arm checks and the unchanged strict-guard default.
docs/captain-hold-lifecycle.md Keep evidence-only reconciliation separate from exact user answers and describe durable request retirement/retry behavior.
docs/watcher-continuity.md Adopt per-actor queue presentation/retirement guarantees while retaining fork away-standby and takeover behavior.
docs/pi-supervision-branch.md Add extension-registered provider support without changing fork routing and away ownership.
docs/calm-mode-feasibility.md Adopt current Pi 0.85.1 export evidence and retain configurable declaration-package verification; remove obsolete claims that CI never installs Pi.
docs/fm-test-portable-shards.md Disclose the cancelled-lane timing source, 12.89-minute balanced estimate and enclosing-job timeout tradeoff. Reflect live guards in portable serial coverage while preserving eight fork shards and retained fork-only hints; refresh both lane tables from the final runner and disclose CI versus local timing sources, including the two split-suite measurements.
docs/scripts.md Point the runner entry at the task-marker refusal without creating another predicate owner.
docs/supervision-protocols/grok.md Include registered custom checks in the existing Grok supervision protocol.
docs/verification/muse.md Update the command gating description; do not claim an unrun live Muse verification.
docs/verification/runtime-backends.md Carry upstream live-gate/Pi evidence with its dates and limits; retain fork runtime and fixture coverage.
docs/verification/process-event-sources.md Carry upstream board/runner evidence and explicit uncertainty/loss limits; retain fork missing-artifact verification.
docs/verification/trace-context.md Describe task-marker propagation separately from optional trace propagation.
docs/fork-divergence.md Record the previously unlisted secondmate semantic thresholds, retire the now-upstream local-root implementation, describe the browser execution gate under portable-serial membership, and explain registered-check/actor counting and backlog-call integration without altering parked branches.
bin/fm-classify-lib.sh Add an explicit-only decision view for repaint suppression while retaining the default complete durable fold and explicit default-key provenance.
bin/fm-push-transition-lib.sh Scope the repaint identity to explicitly keyed decisions as Firstmate selected; preserve acknowledgement and dead-agent recovery.
.pi/extensions/fm-primary-pi-watch.ts Merge the landed fork monitoring fix normally alongside upstream provider work, retaining independent restore and delivery scheduling.
docs/verification/supervision.md Preserve the landed hung-delivery monitoring evidence; the new isolated real Pi/Herdr regression also passed on the combined branch.
bin/backends/herdr.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-backlog-receive.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-bootstrap.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-busy-event.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-config-inherit-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-fleet-snapshot.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-inactive-reconcile.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-lock-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-pending-reply-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-pr-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-remote-file.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-remote-inherit-push.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-remote-inherit.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-remote-job-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-startup-memory-budget-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-supervise-daemon.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-test-isolation-proof.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
bin/fm-x-lib.sh Select /usr/bin/stat for BSD format operations so PATH-installed GNU stat cannot corrupt metadata checks; preserve GNU/Linux and fork behavior.
.pi/extensions/fm-branch-supervision.ts Adopt registered-provider resolution and native Windows helper invocation without changing fork supervision ownership.
.pi/extensions/fm-primary-turnend-guard.ts Use the shared Windows Bash invocation and nullable child streams; retain the fork EPIPE handling and turn-end semantics.
.pi/extensions/lib/fm-operational-input.ts Own platform-specific shell invocation once for sync and async operational-input consumers; preserve fail-closed encoding.
bin/fm-bearings-board.sh Require a live board listener and route structured evidence-only reconcile selections without manufacturing an answer.
bin/fm-captain-hold.sh Own reconcile requests and stable hold lifecycle identities; preserve silent open exit codes and fork completion boundaries.
bin/fm-procevent-extension-capture.pl Capture the structured reconcile choice separately from answer/note data so evidence requests cannot impersonate decisions.
bin/fm-procevent-lavish.sh Add reconcile extraction and a minimum retry-start cadence while retaining artifact-missing and deletion handling.
bin/fm-procevent-lib.sh Own state-root leases, registration-bound launch pacing and runner identity checks; fail closed on PID/PGID uncertainty.
bin/fm-procevent.sh Compose lease refresh, guarded runner lifetime and bounded restarts through the shared owner without touching foreign process groups.
bin/fm-teardown.sh Clarify conservative landed-work proofs and local-root lock scope; preserve strict fork abort attribution and refusal semantics.
bin/fm-wake-drain.sh Count and acknowledge actor-eligible rows while preserving generation-bound durable acknowledgement and recovery.
bin/fm-wake-grant.sh Repair stale actor grants using current ownership evidence without consuming another actor's pending work.
.agents/skills/design-profile/SKILL.md Retain the newly landed fork planning contract in one conversation, with one keyed question, a consequential ADR bar, existing-brief continuity, and external plugin lifecycle ownership.
bin/fm-design-skills-lib.sh Use one shared owner for dispatch-pin adoption and readable-file checks during control and spawn relaunch; refuse missing or mismatched pins before stopping the old worker.
bin/fm-design-skills.sh Retain the fork read-only installed-skill resolver and clarify external lifecycle ownership without installing or substituting a plugin.
docs/examples/crew-dispatch.json Keep the fork example routing description aligned with one-conversation ADR work.
docs/fleet-data-contracts.md Retain the newly landed fork rule that a design relaunch reuses its recorded plugin release and exact pinned paths.
README.md Keep the public task-shape description aligned with the newly landed fork design workflow.
bin/fm-brief.sh Retain the newly landed single-conversation design brief and its existing keyed authority and ADR-only boundary.
bin/fm-control.sh Retain design relaunch identity, required progress note and pre-stop pin validation alongside upstream runtime behavior.
docs/agent-control.md Include design identity and pre-stop pin validation in the existing transactional relaunch contract.

Ledger changes

Record the previously unlisted secondmate semantic-threshold divergence and keyed-decision repaint suppression, with the latter aligned to the explicit-key boundary Firstmate selected.
Retire the local-lock-root divergence now implemented upstream, retaining its behavior witnesses.
Describe registered custom checks alongside queued-wake supervision and actor-specific warnings.
Describe backlog-call identity alongside fork declared-wait routing.
Describe browser execution gating under portable-serial inventory membership.
Record the watcher-triage partition as survival evidence for the unchanged default per-script bound.
Extend the existing design-kind divergence with pointers to the newly landed one-conversation planning owner, dispatch-pin relaunch owner, and executable composition witnesses.
No parked branch is merged, cherry-picked, resurrected, or used as a rebase base: fm/fm-afk-injection-wedge, fm/fm-crew-state-blind-during-fix-round, fm/fm-parked-decision-stale-noise, fm/fm-subagent-model-routing-guard, and fm/fm-vault-drift-check remain excluded.

Publication and landing evidence

Exact pushed head: d5fdeb343d9eb905898ae37e6895efffbce2959b.
GitHub confirms mergeable=true against fork base 2026a110019a49ff99b68e420d65483d2c32ec0c, with the PR open and unmerged.
Its mergeable_state=unstable reflects the intentional compliance failure; all 16 other checks pass.
Final CI run 34809450100 completed successfully on this exact head.
The portable aggregate contains 225 scripts and the separate Herdr artifact contains 14; their union is disjoint and equals all 239 repository test scripts.
Full canonical local lint also passed, using extended ShellCheck 0.11.0 and actionlint 1.7.12 across all three workflows.
The latest follow-up only refreshes timing placement and documentation after the complete local 239-script sweep passed; runner/isolation-proof tests and complete CI verify that final placement.

Every observed failure is attributed:

  • The original-parent failures and first-merge fixture correction are listed in the comparison above.
  • The initial parallel and combined watcher-suite timeouts reflected stale duration estimates and suite growth; the unchanged 24 proven candidates were rebalanced and all 125 watcher cases were partitioned into two serial scripts with exact definitions retained.
  • The backend fixture left FM_HOME ambient while overriding state/data/config; the Treehouse lock owner requires the root home's state/, present locally but absent in clean CI.
    Both exact original parents reproduce this empty-home failure, adding only that directory clears it, and logical/physical project paths resolve to the same identity.
    Dedicated canonical fixture homes and resolver identity/containment checks fix the fixture without weakening the lock.
  • The isolation-proof ordering assertion still expected the former combined watcher suite; it now follows the split wait suite's larger authoritative hint.
  • The later serial lane 5 job timeout followed 16 passing scripts and no failing assertion.
    Completed passing CI measurements project its old placement at 1217543 ms and balance all eight refreshed shards at no more than 773503 ms each.
    Hints retain earlier maxima so gate skips cannot erase native-Windows or slower-run evidence; all script/job deadlines, admission caps, membership, and the 125-case partition remain unchanged.

This is intentionally direct-PR.
The expected red PR must be raised via no-mistakes check reflects the missing no-mistakes signature: that pipeline rebases onto origin/main and would linearize this merge-only round.
Every other required check passes.
Do not squash or rebase this PR.
The configured landing authority must use bin/fm-pr-merge.sh fm-upstream-sync-2026-09-08-round-4 https://github.com/HelloWorldSungin/firstmate/pull/275 -- --merge; the explicit --merge preserves upstream parentage and advances the merge base.

Local receipts: .local/sync-round-4/{fork,upstream,combined,corrected}-tests.json, matching logs and teardown receipts, serial-refresh-lint.log, and ready-complete-ci.json; the final CI artifacts are linked through the run above.
The isolated Pi/Herdr hung-delivery test passed separately with its opt-in enabled; its default-off full-suite entry remains a disclosed skip.
Installed token-free board-listener, harness-liveness, Herdr-version-floor and Pi-branch-responsiveness guards executed successfully.
Protected credential, prompt, and browser cases retained their execution opt-ins, with ordinary-CI skip witnesses passing in tests/fm-live-gate.test.sh.

cr101 and others added 24 commits September 6, 2026 15:49
…#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>
…unchenguid#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>
…nguid#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes kunchenguid#3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner
…nguid#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership
…nchenguid#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks
* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior
…enguid#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass
… is set (kunchenguid#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>
)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
…ranch (kunchenguid#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy
…gress (kunchenguid#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
…idence (kunchenguid#3952)

* fix(tests): make the Calm export-DOM render step retry and report

The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.

What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.

Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.

The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo

* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change

* no-mistakes(review): tests: drop inert chrome flags, report render timeouts

* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…unchenguid#3950)

* fix(bin): make every counted wake queue row presentable or retired

A wake row could be counted as queued while no drain would ever present
it, leaving the operator told to "drain them before anything else" by a
command that printed nothing and offered no acknowledgement.

Two independent paths produced that state.
A row reserved by a live supervision-branch grant is excluded from a main
drain by design, but fm-guard.sh counted the whole queue, so main was
warned about rows only the branch could present, on every guarded command
for as long as the grant was held.
A row that lost its five appended fields or its numeric sequence can never
be claimed, presented, or named by an --ack-through cutoff, yet it still
counted as queued, wedging the queue permanently.

The guard now counts only the rows the calling actor can itself present or
retire, and a main drain retires unusable rows under the queue lock,
reporting them in bounded escaped form before removal so the evidence
survives for the separate row-generating defects. A retirement failure is
reported loudly and never suppresses unrelated consumable work. A main
drain whose remaining rows are all branch-held says so in one bounded line
instead of exiting silently. Grant row-list and owner-record reads move
into fm-wake-lib.sh so the drain, the grant publisher, and the guard share
one implementation.

Ownership is unchanged: a branch drain still touches nothing outside its
grant and never retires a row, and main still cannot present or acknowledge
an active grant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1

* no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass

* no-mistakes(review): add guard advisory for branch-held wake rows

* no-mistakes(document): document per-actor wake counting and unusable-row retirement

* no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing

* fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main

* test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state

* ci: re-trigger after Chrome headless timeout in calm HTML export test

* ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test

* test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat

* no-mistakes(document): Refresh stat and shard docs
…henguid#3945)

The captain's attribution policy (no Co-Authored-By trailer, no
Claude-Session link, no generated-with line) lives in Claude Code's `user`
settings scope. A spawned worker's settings sources are not guaranteed to
load that scope, so a launched worker could write attribution trailers into
its commits and PR bodies regardless of the captain's own configuration.

launch_template()'s claude case now carries the same policy
("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in
its inline --settings JSON, so every claude launch keeps attribution off
independent of which settings scopes end up loaded. Tests assert the policy
on the rendered launch command for both a crewmate and a secondmate spawn.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nchenguid#3946)

* fix(bin): derive the away-mode beacon grace from the poll cadence

fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.

Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.

fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.

Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.

* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh

* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…henguid#3904)

* fix(procevent): bind a source runner to the session that owns it

A process-event source runner is detached into its own process group so a
persistent source survives the turn that armed it. Nothing bounded that
detachment, so a runner could reparent to init and keep its blocking child -
and every process that child spawned - running with nothing left to reap it.
One such runner outlived its home for about a day; the cost was not the runner
but the exec churn of the poll stubs under it, which stalled every fresh
process launch on the host.

Each runner now starts a small guard beside it, in a separate process group,
that re-reads its home's process-event lease and stops the runner's whole
process group once that lease can no longer be proved fresh. Every ordinary
entry point an owning session runs refreshes the lease, and the watcher's
reconcile cycle keeps it fresh in a live home; nothing a runner spawns can
refresh it, so a source cannot certify its own owner. Scope is the owning state
root and one runner generation, never a script or process name, so a live
source in another home is untouched and a live home simply starts a
replacement runner on its next cycle.

The test scaffolding that starts real runners could not reap them either: the
bearings-board and board-render suites tracked their homes in a shell array
appended to inside a command substitution, so the array was always empty and
every listener they started survived the run. Home registration moves to a
`$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path,
now including HUP and QUIT, and the blocking fixture stubs stop themselves at a
bound so an escaped one cannot keep spawning processes indefinitely.

Adds a regression test that reproduces the orphan shape - a reparented listener
with a live descendant tree under it - and proves the whole group and its
process churn stop once its session is gone, that an identical listener in a
home whose session is still there is untouched, and that retirement still
reaches a reparented listener and everything under it.

* no-mistakes(review): Bound source launches and fail closed on guard startup

* no-mistakes(document): Document runner lease and storm containment

* no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite

* test(procevent): pin that an unprovable stop is retried, not abandoned

The owner guard used to call stop_runner_pid and exit unconditionally, so a
stop it could not prove - a descendant still finishing uninterruptible work
outlives even the group signal, and an unreadable process identity proves
nothing - left a still-running expired runner with nothing watching it. That is
the best-effort reaping this mechanism exists to remove, and the fix that made
the guard retry landed without a test holding it in place.

The unprovable attempt is injected through the signal the real path actually
reads: `ps` answers exactly one process-group query for the runner with a group
it does not lead, which is how a stop that cannot be proved is reported, and
every other call is the real command. The test also asserts that the injected
attempt happened, so it cannot pass vacuously if the fixture stops arming.

Fails against the exit-after-one-attempt guard, where the runner survives its
expired lease, and passes once the guard retries on its check cadence.

* docs(procevent): scope the no-self-refresh rule to confused-agent grade

The runner-lease documentation asserted as an absolute that nothing a runner
spawns can refresh the lease, so a source cannot certify its own owner.
That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually
enforces.

The marker holds at confused-agent grade: a runner and its ordinary children
inherit it and skip every refresh, which is exactly the accidental case this
boundary exists for.
A source that deliberately strips the marker from its environment can still
refresh, so adversarial-grade unforgeability is explicitly out of scope and
tracked as separate follow-up design work.

This states the real scope in docs/configuration.md, which owns the operating
contract, and corrects the two matching comments in bin/fm-procevent.sh.
The process-event-sources skill keeps its cross-reference and gains one line
in its never-to-be-claimed list so the overclaim is not reintroduced from the
agent-facing side.
The lease mechanism itself is unchanged.

* no-mistakes(review): Fix process-event lease and launch pacing edge cases

* no-mistakes(review): Scope launch pacing and clarify lease boundaries

* no-mistakes(review): Reap leftover groups and use monotonic launch pacing

* no-mistakes(review): Keep guards alive across runner PID reuse

* no-mistakes(review): Use monotonic leases and simplify launch generation identity

* no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards

* no-mistakes(review): Preserve active pacing state on failed registration

* no-mistakes(review): Reap reused runner groups with registration evidence

* no-mistakes(review): Avoid ambiguous group kills and encode pacing identities

* no-mistakes(review): Abort kill escalation after runner identity reuse

* no-mistakes(review): Gate group signals and prune stale pacing state

* no-mistakes(review): Document bounded PID reuse signaling safety

* no-mistakes(review): Align leaderless group ambiguity guidance

* no-mistakes(review): Expire reboot stamps and preserve publication success

* no-mistakes(review): Bind owner leases to physical state roots

* no-mistakes(document): Clarify process-event lease and pacing contracts

* fix(procevent): drop a platform-dependent post-TERM test assertion

CI ran red on two lanes that the local gate could not see.

Lint failed with SC2034 on two reads in cmd_owner_watchdog that
destructure the state-root identity into five fields while using only the
device and inode.
Local changed-file mode suppresses the cross-file codes that need
--external-sources, so the warning cleared the pre-push lint step and
failed CI's full analysis, exactly as bin/fm-lint.sh's header describes.
The unused fields now read into `_`.

The behavior shard failed on this suite's own post-TERM assertion, which
required the stubbed identity source to be consulted more than once.
Whether that happens is platform-dependent: where the runner leader keeps
waiting on its TERM-ignoring source child, the post-TERM check sees a live
leader whose identity no longer matches, and where the leader dies
promptly it sees a leaderless group carrying the same numeric id.
fm_procevent_pid_state reaches that second verdict without consulting
process identity at all, so the identity source is never read twice and
the count assertion fails through no fault of the behavior.

The case now asserts the invariant both forms share: retirement refuses,
and the ambiguous group is not signalled.
Scoping a mutation to this fixture and making the refusal signal instead
confirms the case still fails, so dropping the count does not leave it
passing vacuously.

* no-mistakes(review): Prevent superseded runners recreating stale pacing stamps

* no-mistakes(document): Document pacing and ambiguity boundaries

* fix(procevent): retire under the recorded identity source and state the home-scoped lease

The reused-group case started its runner with the proc-root override in
place, so the runner recorded a ps-derived identity, then retired it
without that override.
Where /proc exists the retirement read identity from a different source
than the one recorded, the guard correctly refused an identity it could
not confirm, and the case failed on Linux while passing on macOS.
It now retires under the same source, and clearing the stub marker first
turns that cleanup into the complementary assertion: once the ambiguity is
gone, retirement reaps the whole group instead of leaving it behind.

The lease prose claimed a runner is bound to the session that owns it,
while the mechanism binds it to the home.
That gap is what makes a replacement session or an inspection command look
like a defect: any activity in the same home refreshes the lease.
The granularity is deliberate, because a persistent source is meant to
outlive the session that armed it, and binding a runner to that session
would stop the sources this mechanism exists to keep running.
A runner whose source is no longer wanted in a live home is stopped by
reconcile when that source is retired, independently of the lease, so the
lease is the backstop for a home that is gone - the torn-down sandbox this
change bounds - and the residual is recorded as a known limit.

* no-mistakes(review): Rate-limit polls and skip superseded runner launches

* fix(procevent): build the claim-only sweep case as a runnerless owned claim

A superseded generation now observes the registration-identity mismatch,
self-retires, and releases its claim, which is the behavior we want: it
clears its own residue rather than leaving a claim with no runner for the
home sweep to find.

The claim-only sweep case was built by deleting a registration out from
under a live runner, which used to leave that runner in place. It now
makes the runner retire itself, so the sweep raced that exit and retired
one source or two depending on which won. The case failed three runs in
four, alternating between a preflight-count failure and `attempted=1`.

It now builds the state it means to test: kill the runner's group so it
cannot run its own cleanup, assert the owned claim survived that kill, and
only then drop the registration. Coverage is unchanged - a runnerless
owned claim must still be swept - and the result no longer depends on
whether the runner had exited yet. Three consecutive runs pass.

The superseded exit also skipped the runner-marker cleanup the normal path
performs. The marker is written before the launch floor is waited on, and
a home sweep counts a marker with no owned claim as a preflight failure,
so exiting without clearing it would make that home refuse to sweep.

`FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI
runs, though not under the changed-file mode the pre-push gate uses.

* test(procevent): retire a quiet reparented listener instead of racing a storm

Explicit retirement was exercised against the spawn-churning stub, which
made it nondeterministic. Retirement refuses rather than signalling when it
cannot confirm the runner's identity, that identity is read through `ps`,
and the stub's 0.1s spawn loop starves that read often enough that a single
attempt is a race - the suite failed on this case roughly one run in four,
reporting `cannot confirm runner identity; source remains registered`.

The refusal is correct: it is the documented preserve-for-retry contract,
and a separate case already asserts it. So this is a fixture problem, not a
behavior problem.

The storm is still covered where the evidence for it lives. The owner-loss
home keeps the churning stub and still asserts its tick log stops, which is
what proves the churn ended rather than one pid going away. The retirement
home never asserted ticks; it only ever read the descendant pid, so the
spawn loop bought this case nothing while costing it determinism.

It now uses a quiet stub that still reparents and still holds a real
descendant in its process group, so the assertions are unchanged: retiring
the source must reap the reparented listener's whole group and the
descendant under it. Four consecutive runs pass.

* no-mistakes(review): Serialize registration replacement through source child launch

* chore(no-mistakes): require honest test-step scenario marking

The test step recorded scenarios as passing that were only reached through
a stubbed dependency or the executable suite, and its validator refused
them, because `pass` asserts a scenario was verified against the real live
product.

That refusal is correct, so the fix is to mark honestly rather than to
weaken the gate: a scenario driven live stays a pass and cites its live
transcript, while one reached only through a stub or the suite is recorded
as untested with the reason and a pointer to its executable coverage.
Untested scenarios are reported rather than treated as failures, so real
coverage stays visible without claiming verification that did not happen.

The instruction also forbids dropping a scenario to avoid marking it
untested, since that would hide the gap instead of stating it.

* no-mistakes(review): Remove unrelated test scenario policy

* no-mistakes(document): Clarify process-event home lease documentation

* no-mistakes(document): Correct owner guard failure wording
Merge the pinned 17-change upstream prefix while preserving active fork
contracts and excluding parked branches.

Apply Firstmate decisions for live execution gates and semantic secondmate
stall thresholds, and retain upstream progress and wake recovery semantics.
@HelloWorldSungin
HelloWorldSungin merged commit ea6b934 into main Sep 14, 2026
16 of 17 checks passed
@HelloWorldSungin
HelloWorldSungin deleted the fm/fm-upstream-sync-2026-09-08-round-4 branch September 14, 2026 05:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants