Repository navigation
Replies: 1 comment
|
Implementation proposal: PR #17902, currently at The current independent behavior run passed 87 controlled cases: 49 native protocol and persistence cases, 24 delegated lifecycle cases, and 14 effect-worker cases. Scoped server type checking, targeted lint, and cumulative AutoReview loop 6 passed. The three earlier inline findings remain resolved. Terminal cleanup now stays durable after dispatch faults without extending the continuation budget, and cancellation of an absent target thread is accepted as a typed, exact-target no-op. Regressions verify retained failure evidence, lease and Stop precedence, lost phase transitions, original task identity, one acknowledged result wake, and continued retries for transient or unrelated read failures. The fixtures simulate transport and command failures; they do not reproduce an upstream classifier incident. CodeRabbit approved the current head, The five Actions runs on the prior head, Maintainer direction requested: the PR description review requires explicit approval of the proposal's direction and scope before merging. Please confirm whether the default-off, two-attempt Codex recovery described above is acceptable, including qualifying custom endpoints and tasks delegated through T3, or identify the scope changes required. No approval is implied by posting this proposal or its implementation. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
Add environment-wide, bounded recovery for confirmed provider stream failures after the provider has exhausted its own retries. Keep the recovery decision outside the agent model so a failed main turn can recover without a person returning to send a continuation.
Problem to solve
A provider can terminate a turn while the conversation remains reusable. The agent cannot schedule its own next turn after its stream has stopped. Completed delegated results can remain available while the parent stays idle indefinitely, especially when the last child completion already triggered the failed parent turn.
In one inspected non-security workflow on T3 Code
0.0.46-nightly.20261010.2908with Codex CLI0.162.1, five retries carriedresponseStreamDisconnectedand a retryable transport classification. The terminal error becameotherwith the same stream-disconnection message, and the main run stayed idle for approximately five and a half hours. A user follow-up started a successful new turn on the same conversation. The request used a custom local Responses endpoint; retained records do not establish whether the upstream service or intermediary closed the streams. Cyber safety-buffering notices appeared during the successful later turn, so they do not establish that a classifier caused the original stop.Proposed behavior
othercode alone nor message-text matching qualifies a run.Acceptance criteria
otherremain distinguishable from an unrelated unknown failure. A changed failure, explicit policy error, or missing authoritative evidence prevents recovery.Affected area
Environment preferences, provider failure classification, server-owned continuation scheduling, and root and delegated run lifecycle. Existing thread records and environment operations should expose the outcome consistently to local and remote clients.
Non-goals
Automatic continuation after policy refusals, guardrail bypass, recovery inferred from inactivity, prompt replay, model or account fallback, changes to live user databases, general retries across every provider, and capacity or usage-limit recovery.
Alternatives considered
A small external watchdog could prototype the same deterministic rules but would require separate authentication and durable state. A scheduled supervisory agent adds model usage and can itself fail. Codex lifecycle hooks can help with normal completion checks, but fatal stream-error coverage is not established. Bounded subagents and saved checkpoints reduce repeated work without restarting a failed main agent.
Supporting context
The accepted connection-loss proposal defines a narrower initial boundary: native Codex on the built-in OpenAI connection. This proposal requests direction on the current V2 lifecycle, preserved error-chain evidence, custom routes, and root and delegated task ownership. The separate capacity-recovery discussion covers overload and is outside this proposal's scope. Neither related item establishes approval of this broader direction.
OpenAI's app-server documentation exposes terminal turn errors. Its recovery guidance distinguishes bounded transient retries from policy errors and requires checking completed actions before retrying. The proposal preserves that distinction.
This is a feature proposal for maintainer direction and scope. Posting it does not establish maintainer approval of an implementation.
Triage assessment
All reactions