# Durable submit acknowledgement: file-ack P0 + pane fast-confirm P1 > Completion, final answer, steering and background subagents (all CLIs, current): > [../core/coding_cli_turn_signals.md](../core/coding_cli_turn_signals.md). **Status: all five providers implemented 2026-09-19/20. Decision updated 2026-09-20: establish explicit P0 and P1 contracts. Durable correctness and pane-only safety blockers are P0; secondary pane interpretation and terminal experience are P1. Live P0 green for all five providers; live server↔chat e2e PASS for codex/pi/muse/Claude (probe: /tmp/durable-e2e/probe.py). Cursor live P0 passed 2026-09-20 using the RTS deployment key; its server↔chat e2e remains pending. 2026-09-20 fix: durability receipts are now consumed on every ingestion path including durable restore (see "Restore-path receipt consumption" below).** **Current review disposition: changes required (fixes implemented, pending re-review).** The follow-up code review below reproduced three remaining correctness bugs after the FIFO-receipt and append-then-apply fixes; the "Follow-up P1a/P1b/P2 resolutions" section records the fixes with committed regression coverage. Historical live passes do not cover those cases, and the reviewer has not yet re-verified. ### AGY onboarding extension (2026-09-27) AGY is the sixth coding CLI with a provider-side durable live-input receipt. The tmux send returns quickly; the existing server watcher then confirms the exact message against a new type-14 user step in AGY's conversation SQLite. The watcher snapshots the pre-send step index, and repeated identical sends require separate user rows. The real CLI test `TestAgyCLIRealDurableAckContract` passed for two identical retained turns. The isolated full application P0 runner and retained Chat live check passed with AGY 1.2.12 in Gemini API-key mode. This extends the live receipt proof to six providers; it does not revise the earlier historical five-provider run. ### RTS rapid-input follow-up (2026-09-21) Production session `c7d58080-0058-4f6f-a394-feea651f405b` clarified a UI latency case that must not be misclassified as slow durable acknowledgement. Four busy Cursor `/api/query` submissions completed in 15.0–19.6 seconds. The delay occurred before `sent_to_cli`, while Cursor's adapter waited for a safe idle or “Add a follow-up” composer; the asynchronous `store.db` durability watch was not blocking HTTP. The frontend then compounded that legitimate pending-delivery interval by creating later optimistic rows inside its own serialized request lane, making rapid messages appear lost. The frontend fix stages each local user row before that lane, while preserving the existing meanings of the receipts: no tick before provider acceptance, single tick after `sent_to_cli`, double tick after exact durable proof. An accept-now/deliver-later server prototype was rejected and removed because it would violate this document's attempt-once/report-truth decision. Cursor's pending server↔chat e2e must now include rapid identical and distinct sends, assert immediate pending-row visibility, FIFO provider delivery, no premature single tick, and one distinct durable confirmation per accepted send. ### Muse initial-turn false failure (2026-09-22) The workflow builder received the same auto-notification repeatedly while its Muse adapter reported `muse TUI never took in the prompt after submit`. The native `session.jsonl` actually contained `runtime.user_intent.accepted` for the exact text at 09:33:28 IST, after the 09:33:16 turn began. The discovery anchor cut the prompt at 120 **bytes**, splitting the UTF-8 em dash in `completed — status`; JSON matching could never find that invalid fragment. Blind Enter retries during the 60-second intake wait then risked duplicate submissions. This was a false delivery error, not a failed Muse model turn. The Muse initial-turn correction is in `multi-llm-provider-go`: * A pane submit failure is provisional. The turn checks native intake before reporting a delivery error; durable intake overrides a pane mismatch. * The intake wait is observe-only. Neither it nor the fast submitter repeats Enter after an uncertain submit. A missing durable record at the budget boundary remains an *unconfirmed delivery* error, not proof tmux lost bytes. * Discovery uses a valid UTF-8 prefix and a new timestamped `runtime.user_intent.accepted` record containing that text. A file mtime, old identical prompt, assistant echo, or nested subagent log cannot confirm the new send. This initial-turn intake check is distinct from the live-input receipt watcher below. The live path still uses the pane for a fast single tick and the native log for the durable double tick; its pane-failure arbiter was already observe-only. The live watcher now requires an exact `runtime.user_intent.accepted` text match above the send's pre-submit sequence: an assistant echo or the paired queue event cannot masquerade as another accepted send. The invariant for both paths is that pane string matching alone must never decide a user-visible *delivery failure* once submission may have occurred. Focused SDK regressions are green. Broad local runs had unrelated CLI/environment-sensitive failures (Pi wedged-process PID fixture, Claude paste-chip settle, Muse MCP-mount probe, and a Muse resume session dying before prompt readiness). The Muse resume test passed when rerun alone. This change has not yet been re-certified against a live Muse workflow run. ### Cross-provider submit-error audit (2026-09-22) The same P0 rule now covers initial turns in Codex, Pi, Claude Code, Cursor, and Muse. Each adapter snapshots its provider-specific pre-send boundary and, when its fast pane submit check errors, waits **observe-only** for a new provider-owned user-acceptance record before returning that error. A matching rollout user row (Codex), `message_end` user marker (Pi), transcript user row (Claude), `user_query` store row (Cursor), or `user_intent.accepted` row (Muse) overrides the pane error. No match by the provider's budget means *delivery unconfirmed*, not proof that tmux dropped the bytes. Pi's retained new-turn path receives the same fallback. Claude and Cursor live-input sends also now invoke their existing durable watchers when the fast send returns an error; Codex, Pi, and Muse already had live-input arbiters. Provider-specific regression tests cover a pane error with an accepted durable record. Existing automatic Enter-recovery behavior in the non-Muse adapters remains separate from this error-arbitration change: its draft-only safety and duplicate risk still need live review before calling the full transport policy re-certified. The changes are local and have not yet passed live provider/server↔chat re-certification. ### Production dependency parity (2026-09-22) The production agent Dockerfile removed the local sibling-repository replaces and then built the versions pinned in `agent_go/go.mod`. That pin was `multi-llm-provider-go` revision `36f1e19` from 2026-09-19, while local tests used the sibling repository's `main`. Production could therefore omit newer submit/Enter and durable-ack fixes even after those fixes passed locally. The production Docker build now resolves both `mcpagent` and `multi-llm-provider-go` from `main`, matching local development. The refs are explicit Docker build arguments defaulting to `main`; an exact commit may be supplied only for an intentional rollback or bisect. The deterministic Coding CLI workflow verifies those defaults so a pinned-production/local-main split cannot be reintroduced silently. ### Double-submit key transport policy (2026-09-22) Production showed that tmux can accept a submit keystroke while the provider TUI fails to act on it. Every tmux coding-agent adapter now sends two submit keys for the initial submission after writing the draft (`Enter Enter` for Codex, Pi, and Muse; `C-m C-m` for Cursor; `C-e Enter Enter` for Claude Code). The pair is emitted by one `tmux send-keys` command and is covered by an exact key-sequence regression in each provider package (`multi-llm-provider-go` commit `bcaff33`). This redundancy does not elevate terminal text scraping into the source of truth. Provider-owned rollout/transcript/marker/database evidence remains the authoritative durable acknowledgement. Pane matching remains secondary for blocking-state detection and bounded stuck-draft recovery. After the 2026-09-22 main pull, `/api/query` owns steer-vs-next-turn routing. Human submissions now attempt a compatible retained tmux CLI once *before* the occupied-turn queue. A known missing target falls through to the durable next-turn queue; an uncertain send never falls through and risks a duplicate. Claimed queue workers, bot/scheduled/Pulse turns, personal-access-token API callers, explicit new turns, and synthetic notifications do not take this steer-first path. The server's live-delivery deadline is six minutes, beyond the adapters' maximum five-minute durable-ack budget, so its former 15-second cap no longer truncates the fallback. Request cancellation can still end an in-flight HTTP send. Chat's Axios transport has no request timeout (`timeout: 0` in the installed Axios defaults); the Go HTTP server has no write deadline and the checked-in Caddy configs have no explicit response timeout. Other deployed ingress/client disconnect behavior is not verified. A queued turn's `queued_for_turn` receipt is not a provider intake acknowledgment. ## Problem A tmux `send-keys` exit 0 only proves tmux accepted the keystroke. Every CLI adapter therefore confirms submission by string-matching the pane (draft cleared, activity started, queued text visible). When the pane *looks* wrong — redraw races, version drift in spinner/progress strings, ghost/suggestion placeholders, compaction banners, deep scrollback — the adapters return `failed to submit live input to ` and `mcpagent` surfaces that straight to the user instead of queueing (`mcpagent/agent/message_delivery.go`: failed tmux submission is returned to the caller). False negatives become user-visible errors; the natural user reaction (retype and resend) creates duplicates. Meanwhile the pane is also the only fast signal. The provider's own rollout/transcript/marker file is authoritative for *what* the model received but lags (milliseconds idle, unbounded while busy), and it cannot see unsubmitted drafts, modals, trust prompts, or compaction. Neither signal alone is sufficient. ## Purpose and reliability outcome The purpose of this split is to make live-input delivery more reliable by reducing how much correctness depends on interpreting terminal screenshots. It does not make tmux itself inherently more reliable. It gives tmux a smaller, clearer responsibility and uses the provider's durable data for the facts that data can prove. Target ownership: * **Tmux transports interaction:** start and retain the CLI process, serialize input, paste exact bytes, submit control keys, interrupt, and expose the live terminal. * **Provider JSON/markers/transcript/DB prove acceptance:** the exact send-specific durable record produces the double delivery tick and is authoritative for whether the provider received the message. * **Pane inspection handles only pane-only safety states:** trust, login, approval, blocking compaction, and an exact draft visibly stuck in the editor. * **Structured provider data owns conversation state:** final answers, restoration, resume identity, and auditable history should not be reconstructed from hard-wrapped terminal text. This improves platform reliability in two concrete ways: 1. Provider TUI wording, layout, animation, wrapping, or repaint timing can no longer turn a successfully accepted message into a false delivery failure. 2. Shared tmux primitives for session ownership, input ordering, paste, keys, and bounded recovery replace duplicated provider-specific machinery, so one transport fix applies consistently to every CLI. The resulting rule is: tmux answers **"did we transport these bytes to this live process?"**; the provider's durable store answers **"did this exact send become part of the conversation?"** Pane parsing must not silently substitute for the second answer except for the explicitly documented safety blockers that have no durable representation. ## Probe evidence (2026-09-19, codex-cli 0.155.0, gpt-5.6-luna) Hands-on TUI probe in an isolated tmux server, adapter-exact load-buffer/paste-buffer/Enter sequence, scratch workdir: * No rollout file exists before the workspace trust prompt is accepted. Trust stays pane-only. * Rollout rows for a turn: `session_meta` (cwd, cli_version, session_id), `task_started`, `turn_context` (turn_id), `response_item` user row with the **exact message text**, assistant row with `phase=final_answer`, `token_usage_record` + `token_count`, terminal `task_complete {turn_id, last_agent_message, duration_ms}`. * Idle submit to user row in file: ~0.2s warm, ~1s cold (first turn). Short turn to `task_complete`: ~4-5s. * Mid-turn steer: native queue confirmed (`Messages to be submitted after next tool call` + `↳ `). The queued user row landed in the **same turn** (no new `task_started`) **17s after Enter**, folded into the extended turn; one `task_complete` closed it; the model obeyed. * Mid-turn pane reproduced the contract regression case: a `Planning ... (esc to interrupt)` indicator above a ready-looking `Ask Codex to do anything` footer. * Chromes to keep ignoring: update banner, `model: loading` splash, per-turn `done