Skip to content

feat(omp): batch primary watcher wake notifications - #86

Open
pranaypratyush wants to merge 11 commits into
dnth:mainfrom
pranaypratyush:fm/fm-omp-primary-nextturn-batching
Open

feat(omp): batch primary watcher wake notifications#86
pranaypratyush wants to merge 11 commits into
dnth:mainfrom
pranaypratyush:fm/fm-omp-primary-nextturn-batching

Conversation

@pranaypratyush

Copy link
Copy Markdown

Intent

Implement Part B of issue kunchenguid#3123: OMP primary watcher wakes must use hidden custom nextTurn delivery with triggerTurn, batching one pending notification without changing Part A task-inbox doorbells. Preserve existing durable wake queue and generation-bound acknowledgement ownership. A same-session extension reload must not duplicate the continuation, while a replacement session or process re-notifies the exact unacknowledged durable batch and never retires its queue rows. Do not merge.

What Changed

  • Deliver OMP primary watcher wakes as hidden custom nextTurn messages with triggerTurn, coalescing pending notifications through the shared watcher core.
  • Persist the pending wake batch in an OMP notification claim so same-session extension reloads reuse it and replacement sessions or processes replay unacknowledged batches.
  • Retire notification claims only after fm-wake-drain.sh acknowledges the durable queue rows, and route supervision-branch fallback wakes through the same primary notification path.

Risk Assessment

✅ Low: The change is tightly scoped and its claim, reload, replay, fallback, and mixed-actor acknowledgement paths preserve the required durable-wake invariants on source review.

Testing

Inspected the target change, then exercised the OMP adapter’s end-to-end hidden nextTurn/triggerTurn batching, reload and replacement replay, delivery-failure recovery, and acknowledgement lifecycle; also exercised durable queue generation-bound, actor-scoped acknowledgement/claim retirement. Both focused suites passed, evidence transcripts were captured outside the worktree, and testing left no worktree changes.

Evidence: OMP primary integration transcript

Focused OMP adapter lifecycle transcript: hidden nextTurn delivery, batching, reload/replacement replay, failure recovery, and acknowledgement behavior.

ok - OMP path resolution stays canonical when readlink -f is unavailable
ok - OMP primary identity requires launch-bound Bun and OMP realpaths plus the exact argv boundary
ok - standalone OMP identity requires the launch-bound PID executable
ok - exact-OMP ancestry stops at the innermost foreign harness ancestor
ok - OMP fresh primary lifecycle creates canonical state and atomically replaces a marker symlink without following it
ok - OMP native identity supports physical Bun scripts and known compiled virtual entrypoints
ok - native OMP alone admits a fresh plain checkout and delivers one startup instruction
ok - OMP primary refuses whitespace-bearing identity before marker publication
ok - OMP primary extension binds secondmate doorbells after session readiness
ok - OMP confirms the recovery handling handshake before its hidden next-turn notification
ok - OMP surfaces a refused handling handshake as one hidden notification
ok - OMP batches durable wake rows into replayable hidden next-turn notifications
Evidence: Durable queue integration transcript

Focused durable wake queue transcript: generation-bound acknowledgements and actor-scoped OMP claim retirement.

WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 25 --recovery-generation 394885.1788181401.B9L345
ok - concurrent append plus drain preserves durable records through acknowledgement
ok - signal written while no watcher runs is caught on next run
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 408416.1788181408.v2Hyvd
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no live watcher process holds this home lock (last beat: 0s ago).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  After draining queued wakes, repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
ok - stale wake is queued before suppressor state is advanced
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 410218.1788181409.mv6dLS
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no live watcher process holds this home lock (last beat: 0s ago).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  After draining queued wakes, repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else.
ok - a not-provably-working stale wake is queued before its suppressor is advanced
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 411384.1788181410.d097fc
ok - registered custom check output is queued before cadence suppression
ok - concurrent drains replay until one post-handling acknowledgement consumes records
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 413044.1788181411.Qmys7n
ok - drain collapses obvious duplicate heartbeat and signal records
ok - drain asserts watcher liveness: warns on a lapse, stays silent for a live watcher with a fresh beacon
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 8 --recovery-generation 414230.1788181411.3MwlI4
ok - structural signal enrichment is separate, deduped, home-local, and tier-zero for other wakes
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 13 --recovery-generation 415773.1788181412.CWKruY
ok - bounded reads and per-item/global caps fail open with explicit truncation and omission markers
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 418271.1788181413.Bt0mcU
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 2 --recovery-generation 418271.1788181413.Bt0mcU
ok - slow annotation releases the append lock and a deleted status file fails open
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation existing
ok - wake append publishes atomic recovery evidence before durable rows
ok - wake drain: generation-less legacy wakes are adopted and acknowledged
ok - wake drain: a stale acknowledgement cannot retire or consume a newer recovery episode
ok - wake drain: recovery acknowledgement failures are explicit and retryable
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 2 --recovery-generation 432313.1788181418.L8ukrM
ok - interruptions preserve durable rows until post-handling acknowledgement
ok - recovery-marker transitions reenter safely without breaking exclusion
ok - handling confirmation is bounded without bypassing foreign marker-lock exclusion
ok - a branch-actor scoped ack never swallows an unacked main-owned row, and main's later drain sees exactly what remains
ok - a main drain excludes a branch-granted row and acknowledges only its own presented rows
ok - OMP notification acknowledgement waits for all scoped rows through its cutoff
ok - branch activation rollback succeeds before publication and preserves every active grant

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (6) ✅
  • 🚨 .omp/extensions/fm-primary-omp.ts:398 - This deletes the only retained nextTurn payload before binding the replacement session, and the switch path never replays it. If session A has a pending unacknowledged wake then /new or /resume switches to B before that continuation begins, B has durable queue rows but receives no exact-batch notification. This contradicts the required “replacement session or process re-notifies the exact unacknowledged durable batch.” Preserve the claim across the switch and replay after binding B.
  • 🚨 .omp/extensions/fm-branch-supervision-omp.ts:866 - The branch-failure fallback remains reachable after the primary offers an ordinary wake to the supervision branch: a failed branch sends the main watcher wake as visible-hidden steer, not hidden custom nextTurn. This leaves a concrete OMP primary-wake path contradicting the required “OMP primary watcher wakes must use hidden custom nextTurn delivery with triggerTurn.”

🔧 Fix: Preserve OMP wake claims across switches and fallback
3 issues (2 errors, 1 warning) still open:

  • 🚨 .omp/extensions/fm-primary-omp.ts:299 - The claim deduplicates by notification key, so fallback bypasses the core latch. A main-owned check: wake first queues wake-queue; before its turn begins, an ordinary wake can be accepted by the branch and then fail, emitting branch-fallback. That overwrites the retained claim and submits a second nextTurn, violating the required “batching one pending notification” and losing the first payload for later replay.
  • 🚨 .omp/extensions/fm-primary-omp.ts:320 - Replay identifies a process only by PID and session hash. A replacement process that reuses the former PID and resumes the same session returns here without re-notifying the unacknowledged batch, contradicting the required “replacement session or process re-notifies the exact unacknowledged durable batch.” Persist and compare a per-process boot nonce.
  • ⚠️ .omp/extensions/fm-primary-omp.ts:356 - The fallback event is marked accepted before claim publication and notification submission. If either throws, the branch suppresses the error and skips its direct fallback because accepted was already set, leaving durable rows without a notification until another wake occurs. Accept only after the primary path is safely established, or let submission failure leave the offer unaccepted.

🔧 Fix: Harden OMP wake claim ownership and batching
1 error still open:

  • 🚨 bin/fm-primary-watch-core.ts:721 - notificationTurnStarted() retires the persisted claim at before_agent_start, before the agent has run fm-wake-drain.sh and its generation-bound acknowledgement. Concrete path: a queued hidden nextTurn starts, this unlink succeeds, then the agent switches sessions or the process exits before acknowledging; the durable queue rows remain, but the replacement has no claim for replayWakeNotification() to re-notify. This contradicts the required criterion, “a replacement session or process re-notifies the exact unacknowledged durable batch.” Keep or reconstruct the claim until durable acknowledgement establishes handling.

🔧 Fix: Retain OMP wake claims through handling
2 errors still open:

  • 🚨 .omp/extensions/fm-primary-omp.ts:377 - The claim is only changed to inflight; no durable-acknowledgement path removes or advances it. After fm-wake-drain.sh --ack-through successfully retires all rows, /new or /resume replays the stale claim at replayWakeNotification(), creating a continuation for an already-acknowledged batch. This contradicts the required “replacement session or process re-notifies the exact unacknowledged durable batch.” Reconcile the claim at the generation-bound acknowledgement boundary.
  • 🚨 .omp/extensions/fm-primary-omp.ts:469 - Every before_agent_start clears the pending latch and marks the claim inflight, although it is not scoped to consumption of the custom watcher nextTurn. If another agent turn starts while that continuation remains queued, a later wake overwrites the first claim and schedules a second continuation, violating required “batching one pending notification” and losing the first exact replay payload. Correlate this transition with actual watcher-notification consumption.

🔧 Fix: Reconcile OMP claims with acknowledged watcher delivery
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-wake-drain.sh:43 - This contradicts the required “replacement session or process re-notifies the exact unacknowledged durable batch.” If notification A has begun (inflight), a new wake B arrives during that handling turn, the adapter replaces A's claim with pending B and queues one next-turn continuation. The handler can then drain and acknowledge all A+B rows in its exact --ack-through command, but this condition refuses to publish an acknowledgement for pending B. B is therefore replayed after /new, /resume, or process replacement even though its durable rows were acknowledged. Bind claims to the drain presentation/cutoff and reconcile every claim whose represented rows were actually acknowledged.
  • ⚠️ .omp/extensions/fm-primary-omp.ts:430 - If queueWakeNotification() throws while handling a branch fallback (for example, its claim-file publication fails), wake.accept() is never reached and the exception propagates into the branch's catch, which suppresses it. The direct no-handler fallback is skipped, leaving durable rows without a wake until a later event. Catch the primary-handshake failure locally and attempt the existing direct hidden nextTurn fallback before allowing the branch path to finish.

🔧 Fix: Bind OMP claim retirement to acknowledged queue cutoffs
3 errors still open:

  • 🚨 .omp/extensions/fm-branch-supervision-omp.ts:859 - If the primary handshake’s sendWakeNotification() throws, queueWakeNotification() discards its claim before rethrowing; this direct fallback then schedules nextTurn without restoring a claim. A replacement before acknowledgement finds no claim although durable rows remain, contradicting the required “replacement session or process re-notifies the exact unacknowledged durable batch.” Route fallback retry through claim-preserving delivery.
  • 🚨 bin/fm-wake-drain.sh:40 - Claim retirement only publishes acknowledgement for the main actor. With main-owned check seq 1 and branch-owned signal seq 2, the claim records through=2; main acknowledges seq 1, branch later consumes seq 2 but cannot reconcile the claim, and a replacement replays an already-acknowledged batch. This violates “exact unacknowledged durable batch.”
  • 🚨 .omp/extensions/fm-primary-omp.ts:532 - A same-session extension reload leaves the superseded message_start handler registered. It marks a pending claim inflight first; the active handler then returns here without clearing its coalescing latch. After that active core submits one later notification, its latch remains set, so a wake arriving after that turn’s drain cutoff is silently coalesced without any continuation. This breaks the required one-pending batching/continuity behavior on reload.

🔧 Fix: Harden OMP wake claim continuity
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • git diff --stat ef8ee4955c343fb46dc7d3ebf474bd252c220ac0 bf823ad9c7a1608cdad2d9d106c3c50ff66ce097 and focused test discovery
  • bash tests/fm-omp-primary.test.sh
  • bash tests/fm-wake-queue.test.sh
  • git status --short and focused-path diff check after testing
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Aug 31, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-08-31T13:31:58.542652Z f9df832 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f9df832beb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .omp/extensions/fm-primary-omp.ts Outdated
Comment on lines +327 to +328
const sessionFile = ctx.sessionManager.getSessionFile() || "unknown";
notificationSession = createHash("sha256").update(sessionFile).digest("hex");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use a unique identity when no session file exists

Captain, when getSessionFile() returns undefined—which its API permits—every fileless conversation hashes the same literal "unknown". If /new or another session replacement occurs before a session file exists, replayWakeNotification() mistakes the replacement for the original session and suppresses replay; the old hidden nextTurn is tied to the abandoned conversation, while newly detected durable rows coalesce behind its pending claim, potentially leaving supervision notifications stranded until another process restart. Use a per-session identity that changes on every replacement when no file path is available.

AGENTS.md reference: AGENTS.md:L225-L225

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant