You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Opt-in automatic recovery when Codex ends a turn because the model is at capacity
#14687
Update (2026-10-05). Since this was posted, #2829 merged the V2 orchestrator into main. That changes three things:
Capacity errors are now surfaced. A terminal serverOverloaded becomes a failed turn with a "Provider error" row carrying Codex's message, and the sidebar shows Failed (CodexAdapterV2.ts L3914, L4923, session-logic.ts L365). Codex's in-turn retries show as "Retrying provider (n/m)". What remains is only the recovery question below.
Decision 6 (which line) is moot: V2 is main. main already has the opt-in autoResumeLimitedThreads (default off, settings.ts L1276) and UsageLimitRecoveryWorker, which a capacity recovery could follow. Nothing on main recovers from serverOverloaded.
Observed once locally: a Codex gpt-6-astra turn worked for 77 minutes, ended with code: "serverOverloaded", and the thread then sat idle for about 3½ hours until I returned. It was 1 occurrence in about 5,500 runs, so it's rare, but expensive when it happens.
Problem
Codex retries capacity errors itself. While it retries, the error carries willRetry: true and T3 shows "Reconnecting… N/5". If capacity does not return, Codex ends the turn with codexErrorInfo: "serverOverloaded" ("Selected model is at capacity. Please try a different model."). T3 then shows a terminal provider error and does nothing else. Unattended or remotely driven work stays stopped until someone comes back and types "continue".
Current main (a3abb52). A non-retrying Codex error becomes a generic provider_error; nothing distinguishes serverOverloaded (CodexAdapter.ts L2139-L2153). The typed code already exists in the protocol (schema.gen.ts L17867).
These shaped the proposal. None of them approves it.
feat(threads): opt in to automatic resume after connection loss #13740 (accepted) adds opt-in Resume after connection loss. Its triage defines a shape I've copied: off by default, native Codex, typed codexErrorInfo, only after willRetry ends, an empty-prompt continuation, and a server-owned notice. The same triage says not to arm that setting for serverOverloaded. Capacity therefore needs its own decision.
Setting. A new environment setting, Retry when the model is at capacity, placed beside the restart and connection-loss toggles. Off by default. It applies only to native Codex; other providers keep today's behavior.
Trigger. Arm only when a native Codex turn ends failed with codexErrorInfo === "serverOverloaded" and willRetry unset. Never arm from message text, stderr, a bare 503, usage or rate limits, auth, policy or context errors.
Continuation. Send continuation: true with an empty prompt on the same Codex thread. Do not replay the prompt or add a user row. Failed attempts leave nothing in the transcript. Send no model override, so the thread's current model applies and T3 never switches models.
Backoff. Bounded and jittered. For example: 3 attempts at about 1, 4 and 15 minutes, each drawn from 50–100% of its base. Codex sends no typed retry hint today, so nothing is parsed from error text.
Notice. One server-owned thread notice ("Model at capacity. Retrying in ~N min, attempt k/3") with Stop retrying, on web and mobile. Stop retrying is a service method, so agents can reach it too.
Precedence. A pending retry is cancelled by any of:
Stop or Stop retrying
a new user message or turn
settle, snooze, archive or delete
turning the setting off
an open approval or question
A retry never unsnoozes or unsettles a thread.
One owner. Capacity, connection-loss and restart recovery never all fire for the same failure; only one sends. A pending wait is dropped on server restart, and its notice with it.
Out of scope: Claude and other adapters, which have no promptless continuation; generic retry across adapters; model or account fallback; prompt replay; rate or usage limits; any change to restart or connection-loss behavior.
Decisions requested
This is a new workflow. Being opt-in does not make it a configuration tweak, so I'm asking for direction and scope before writing a PR.
Workflow. May T3 start a Codex turn on its own after a terminal serverOverloaded failure, as an opt-in setting that defaults to off?
Transcript. Should the retry be a promptless native continuation with no user row, as proposed? Or should it follow the V2 usage-limit pattern of a server-authored "Continue where you left off."?
Model. Should the retry use the thread's current model selection at dispatch, with no T3-initiated switch?
Setting. Should this be its own toggle, or one switch shared with "Resume after connection loss" that keeps per-reason classification?
Timing. Are the attempt count, delays and jitter acceptable?
Restart. Is dropping a pending wait on server restart acceptable?
Verification plan if approved
Tests. Focused tests use a scripted Codex app-server failure, because a real overload can't be triggered on demand:
the arm and exclusion matrix
jitter and attempt budget with TestClock
every cancellation path
snooze checks on parsed instants
mutual exclusion with restart and connection-loss recovery
Media. Before and after screenshots of the notice on web and mobile, plus a short recording of the real client going from notice to continuation, and of Stop retrying. All labelled as driven by the scripted failure.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
Codex retries capacity errors itself. While it retries, the error carries
willRetry: trueand T3 shows "Reconnecting… N/5". If capacity does not return, Codex ends the turn withcodexErrorInfo: "serverOverloaded"("Selected model is at capacity. Please try a different model."). T3 then shows a terminal provider error and does nothing else. Unattended or remotely driven work stays stopped until someone comes back and types "continue".Evidence:
turn/start(comment). Discussion [Feature]: Surface Codex model capacity errors in T3 Code #11851 covers surfacing these errors. Closed fix(v2): retry runs that fail because the provider is busy #12692 describes a Codex run that worked for 77 minutes before ending this way. No raw payload is attached here.main(a3abb52). A non-retrying Codexerrorbecomes a genericprovider_error; nothing distinguishesserverOverloaded(CodexAdapter.ts L2139-L2153). The typed code already exists in the protocol (schema.gen.ts L17867).Related work: precedent, not approval
These shaped the proposal. None of them approves it.
codexErrorInfo, only afterwillRetryends, an empty-prompt continuation, and a server-owned notice. The same triage says not to arm that setting forserverOverloaded. Capacity therefore needs its own decision.t3code/codex-turn-mapping, feat(v2): resume limited threads when usage resets #12686, tracked in [Feature]: Recognize USAGE LIMIT and auto-ping when it's back #6796) is opt-in throughautoResumeLimitedThreads, which defaults to off. It sends a server-authored "Continue where you left off." when a usage window resets. It covers only usage limits; V2 has no capacity failure class.Proposed outcome
Setting. A new environment setting, Retry when the model is at capacity, placed beside the restart and connection-loss toggles. Off by default. It applies only to native Codex; other providers keep today's behavior.
Trigger. Arm only when a native Codex turn ends
failedwithcodexErrorInfo === "serverOverloaded"andwillRetryunset. Never arm from message text, stderr, a bare 503, usage or rate limits, auth, policy or context errors.Continuation. Send
continuation: truewith an empty prompt on the same Codex thread. Do not replay the prompt or add a user row. Failed attempts leave nothing in the transcript. Send no model override, so the thread's current model applies and T3 never switches models.Backoff. Bounded and jittered. For example: 3 attempts at about 1, 4 and 15 minutes, each drawn from 50–100% of its base. Codex sends no typed retry hint today, so nothing is parsed from error text.
Notice. One server-owned thread notice ("Model at capacity. Retrying in ~N min, attempt k/3") with Stop retrying, on web and mobile. Stop retrying is a service method, so agents can reach it too.
Precedence. A pending retry is cancelled by any of:
A retry never unsnoozes or unsettles a thread.
One owner. Capacity, connection-loss and restart recovery never all fire for the same failure; only one sends. A pending wait is dropped on server restart, and its notice with it.
Out of scope: Claude and other adapters, which have no promptless continuation; generic retry across adapters; model or account fallback; prompt replay; rate or usage limits; any change to restart or connection-loss behavior.
Decisions requested
This is a new workflow. Being opt-in does not make it a configuration tweak, so I'm asking for direction and scope before writing a PR.
serverOverloadedfailure, as an opt-in setting that defaults to off?main.Verification plan if approved
TestClockAll reactions