You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(threads): summarize recoverable progress after a failed turn
#15982
Offer a user-requested progress recovery summary after a provider turn fails. Help the user identify work already done, artifacts that may be usable, and checks still outstanding before deciding how to continue. Existing transcript preservation, portable handoffs, and child-task result retrieval remain useful foundations.
Problem to solve
A long task can fail after producing useful intermediate work but before delivering its final answer. The terminal provider error explains that execution stopped, but the user still has to reconstruct what survived and what remains unfinished from messages, tool activity, child results, and the filesystem.
In one inspected native Codex run, eight research lanes completed and 15 Markdown drafts were staged. Independent reviews and a sequence-plan validation result were retained. The provider then reported five reconnection attempts followed by stream disconnected before completion: stream closed before response.completed and a failed turn. No final assistant report followed. Read-only inspection afterward found all 15 staged files still present and two reproduction-input objections still unresolved in the affected drafts. The recoverable outcome was substantial, but the draft set was not ready to claim as complete.
The useful recovery question is therefore what work can be kept and what must happen next. A generic continuation prompt or a preserved transcript leaves the user or next agent to assemble that answer.
Proposed behavior
Provide an explicit way to inspect recoverable progress for a failed turn. Present the last request, recorded work completed, known artifact references, completed child-task outcomes, and outstanding review, validation, or user-input requirements. Distinguish directly recorded results from inferred progress and from checks that have not been performed.
An artifact path mentioned by a tool is a reference, not proof that the artifact currently exists or is correct. If the recovery flow inspects a referenced artifact, state what was checked. If it does not inspect it, mark its existence and readiness as unverified. A passed review or tool command must not imply that the whole task is finished, especially when later objections or missing delivery steps are recorded.
Let the user carry the recovery summary into a deliberate continuation using the existing supported continuation or handoff paths. Preserve the original request and unfinished closure work. Starting execution must remain a separate user action; requesting a summary must not replay the original task, restart child tasks, or create a new conversation.
Acceptance criteria
Given a failed turn with completed tool work, completed child results, artifact references, and unfinished checks, the user can request one consolidated recovery summary without reading raw provider logs or a SQLite database.
The summary identifies the last request and separates recorded completed work, partial work, failed work, and unknown state. It attributes child conclusions and retains concrete unresolved review objections.
Referenced artifacts appear with their inspection status. Missing, changed, unreadable, or unchecked artifacts cannot be presented as verified deliverables. Files outside the recorded task references are not scanned merely to build this summary.
In the motivating case, recovery keeps the 15 staged drafts and existing research as potential inputs, identifies the two unresolved reproduction-input objections and remaining delivery checks, and does not declare the draft set complete.
A user-directed continuation receives the original goal, useful progress references, and remaining checks. It does not instruct the agent to rerun already completed mutations blindly. Missing provider context or unsupported continuation is disclosed, with any available portable handoff identified accurately.
Summary generation starts no implementation, continuation, child-task replay, or new thread. Stop, newer work, and pending approvals or questions retain their existing meaning. Repeated summary requests create no duplicate task execution.
Existing native continuation, portable transcript handoffs, child-result retrieval, and provider retry behavior retain their current guarantees. When no useful progress is available, the summary states that limitation instead of inventing work.
Affected area
Failed-turn recovery in the thread workflow, including the user interface and the context passed to an explicitly chosen continuation.
Non-goals
Diagnosing the underlying provider outage, asserting that T3 caused the stream failure, automatic retries, automatic forks, changing compaction policy, restoring a provider's hidden reasoning or in-flight tool state, and completing the original task merely by generating a summary.
Alternatives considered
Reporting the stream failure as a T3 bug would require evidence of a violated T3 contract or an app-level cause. The inspected incident establishes a provider failure, not that attribution.
Automatic resume after connection loss is already covered by accepted issue #13740. That execution policy does not replace progress inspection for a user who wants to assess unfinished work before continuing.
Manual transcript inspection and existing portable handoffs preserve valuable context. A focused recovery summary reduces the work required to identify which artifacts and checks are usable. This is the preferred outcome; it leaves storage, summary generation, and continuation implementation open.
Supporting context
The incident evidence came from read-only inspection of a persisted T3 thread, its native Codex protocol events, and the specific staged files referenced by that thread. No outage reproduction, failure injection, or recovery attempt was run. The observed provider was native Codex using GPT-6.1 Sol at high reasoning. The installed desktop version read during this investigation was 0.0.46-nightly.20261005.2667; the investigation did not independently establish which build was active when the failed turn began.
The final provider notification reported willRetry: false after Reconnecting... 1/5 through 5/5, then turn/completed with status failed and duration 3642432 milliseconds. Its last reported token usage was 547902 against a modelContextWindow of 828400. These observations do not establish context exhaustion or the underlying cause of the disconnect. Private thread identifiers, paths, project details, and raw transcripts are omitted.
At inspected upstream commit cf3e714b0f58e29c8fa8660db50d2187e3263b65, ChatView.tsx offers its manual Resume path for interrupted turns and qualifying usage-limit failures. The Resume handler sends "Continue where you left off." This source inspection concerns current main, not a claim that the installed build has exactly the same implementation.
Portable handoffs already preserve selected requests, answers, partial failed work, and references for reading more history. Their documented form is a transcript handoff rather than an agent-written recovery summary. App-owned child results are also already available through task_status. The proposed capability should use and accurately explain those existing records.
Impact basis: This proposes a user-visible recovery aid. The inspected provider failure does not establish a violated T3 contract or a T3-caused outage.
Workaround status: not-applicable
Workaround basis: Defect impact is non-applicable. Manual inspection, ordinary continuation, portable handoffs, and child-result retrieval are existing alternatives; this request makes the recoverable progress and remaining checks explicit.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Offer a user-requested progress recovery summary after a provider turn fails. Help the user identify work already done, artifacts that may be usable, and checks still outstanding before deciding how to continue. Existing transcript preservation, portable handoffs, and child-task result retrieval remain useful foundations.
Problem to solve
A long task can fail after producing useful intermediate work but before delivering its final answer. The terminal provider error explains that execution stopped, but the user still has to reconstruct what survived and what remains unfinished from messages, tool activity, child results, and the filesystem.
In one inspected native Codex run, eight research lanes completed and 15 Markdown drafts were staged. Independent reviews and a sequence-plan validation result were retained. The provider then reported five reconnection attempts followed by
stream disconnected before completion: stream closed before response.completedand a failed turn. No final assistant report followed. Read-only inspection afterward found all 15 staged files still present and two reproduction-input objections still unresolved in the affected drafts. The recoverable outcome was substantial, but the draft set was not ready to claim as complete.The useful recovery question is therefore what work can be kept and what must happen next. A generic continuation prompt or a preserved transcript leaves the user or next agent to assemble that answer.
Proposed behavior
Provide an explicit way to inspect recoverable progress for a failed turn. Present the last request, recorded work completed, known artifact references, completed child-task outcomes, and outstanding review, validation, or user-input requirements. Distinguish directly recorded results from inferred progress and from checks that have not been performed.
An artifact path mentioned by a tool is a reference, not proof that the artifact currently exists or is correct. If the recovery flow inspects a referenced artifact, state what was checked. If it does not inspect it, mark its existence and readiness as unverified. A passed review or tool command must not imply that the whole task is finished, especially when later objections or missing delivery steps are recorded.
Let the user carry the recovery summary into a deliberate continuation using the existing supported continuation or handoff paths. Preserve the original request and unfinished closure work. Starting execution must remain a separate user action; requesting a summary must not replay the original task, restart child tasks, or create a new conversation.
Acceptance criteria
Affected area
Failed-turn recovery in the thread workflow, including the user interface and the context passed to an explicitly chosen continuation.
Non-goals
Diagnosing the underlying provider outage, asserting that T3 caused the stream failure, automatic retries, automatic forks, changing compaction policy, restoring a provider's hidden reasoning or in-flight tool state, and completing the original task merely by generating a summary.
Alternatives considered
Reporting the stream failure as a T3 bug would require evidence of a violated T3 contract or an app-level cause. The inspected incident establishes a provider failure, not that attribution.
Automatic resume after connection loss is already covered by accepted issue #13740. That execution policy does not replace progress inspection for a user who wants to assess unfinished work before continuing.
Manual transcript inspection and existing portable handoffs preserve valuable context. A focused recovery summary reduces the work required to identify which artifacts and checks are usable. This is the preferred outcome; it leaves storage, summary generation, and continuation implementation open.
Supporting context
The incident evidence came from read-only inspection of a persisted T3 thread, its native Codex protocol events, and the specific staged files referenced by that thread. No outage reproduction, failure injection, or recovery attempt was run. The observed provider was native Codex using GPT-6.1 Sol at high reasoning. The installed desktop version read during this investigation was
0.0.46-nightly.20261005.2667; the investigation did not independently establish which build was active when the failed turn began.The final provider notification reported
willRetry: falseafterReconnecting... 1/5through5/5, thenturn/completedwith statusfailedand duration3642432milliseconds. Its last reported token usage was547902against amodelContextWindowof828400. These observations do not establish context exhaustion or the underlying cause of the disconnect. Private thread identifiers, paths, project details, and raw transcripts are omitted.At inspected upstream commit cf3e714b0f58e29c8fa8660db50d2187e3263b65, ChatView.tsx offers its manual Resume path for interrupted turns and qualifying usage-limit failures. The Resume handler sends "Continue where you left off." This source inspection concerns current main, not a claim that the installed build has exactly the same implementation.
Portable handoffs already preserve selected requests, answers, partial failed work, and references for reading more history. Their documented form is a transcript handoff rather than an agent-written recovery summary. App-owned child results are also already available through
task_status. The proposed capability should use and accurately explain those existing records.Related proposals are automatic resume after connection loss, quota-cutoff distillation into a new thread, and provider continuity and compaction provenance. This request focuses on a user-requested account of recoverable progress after a failed turn, without requiring automatic execution or a new thread.
Triage assessment
All reactions