[safe-output-health] Safe Output Health Monitor — 2026-09-22 #62550
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-09-23T05:00:09.378Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Executive Summary
Audited 227 cached agentic workflow runs from the last 24 hours, focused exclusively on safe output job health (the consolidated
safe_outputsjob that executescreate_discussion,create_issue,add_comment,add_labels,approve_workflow_run,update_pull_request, etc. via its single "Process Safe Outputs" step). 2 of 227 runs (0.88%) had a failingsafe_outputsjob — a 99.12% job success rate. Both failures show partial success: some items in the batch completed (a discussion was created; two PRs were updated and commented on), while specific item types failed, consistent with this system's well-established per-item-isolation behavior. One of today's two failures (add_labelsin Auto-Triage Issues) is a brand-new signature never seen in the 13 prior daily audits on record since 2026-08-22; the other (approve_workflow_runin PR Sous Chef) is the 13th occurrence of a long-standing, already root-caused recurring issue whose fix remains unshipped.Job Statistics
safe_outputsjob failuresafe_outputsjob success rateadd_labels, Auto-Triage Issues)approve_workflow_run, PR Sous Chef — 13th occurrence)Note: the
safe_outputsjob is a single consolidated job per run (not split per safe-output type); all safe-output writes for a run happen inside its one "Process Safe Outputs" step, and failures are isolated per-item rather than aborting the whole batch.Error Clusters
Cluster 1 — Auto-Triage Issues /
add_labels(NEW, 1 occurrence)codex(openai/gpt-5.3-codex), 2026-09-22T00:56:29Zcreate_discussionsucceeded (discussion #62505 created)add_labels(1 call) did not appear in the run's successful-items list — this is what flipped the job tofailureadd_labelsfailure; raw error text was not retrievable this audit.Cluster 2 — PR Sous Chef /
approve_workflow_run(RECURRING, 13th occurrence)pi(copilot/gpt-5.4), 2026-09-22T02:55:50Zupdate_pull_requestx2 andadd_commentx2 succeeded (PR #62490, PR #62484)approve_workflow_run(6 calls) all failed — matches the long-confirmed recurring signature (fork-PR PAT permission gap, or a protected-files decline miscategorized as a hard failure); exact sub-cause unconfirmed today (raw error text unavailable)create_issue(1 call) also did not produce a result — a new data point for this workflow, possibly related to a priorcreate_discussionhandler bug seen 2026-09-02Root Cause Analysis
approve_workflow_run(PR Sous Chef): Root-caused since 2026-08-26 across two confirmed sub-causes — (a) the safe_outputs job's PAT lacks permission to approve fork-PR workflow runs ("Resource not accessible by personal access token"), a genuine scope bug; (b) the tool correctly declines to approve PRs touching protected files, but the decline is surfaced via the same hard-failure code path as genuine errors, a miscategorization. 13 occurrences over a month, fix still unshipped.add_labels(Auto-Triage Issues): Not yet root-caused — first occurrence. The fact that a siblingcreate_discussioncall in the same batch succeeded confirms per-item isolation is working, but the specific cause of the label failure is unknown pending raw error text.create_issue(PR Sous Chef): Not yet root-caused — first occurrence for this workflow. Could be a genuine handler bug (structurally similar to the 2026-09-02create_discussion"Cannot read properties of undefined" TypeError) or a benign no-op from dedup logic that simply doesn't populate a result URL.safe-output-errors.json/workflow-logs/*_safe_outputs.txt) was not retrievable even when the underlying log-fetch call succeeded — it returned only the same structured summary already available locally, not the console text needed to confirm exact error messages. This is now the single biggest blocker to quickly root-causing new failure signatures.Recommendations
Immediate actions:
approve_workflow_run's protected-files decline as a non-failure outcome instead of a hard job failure (proposed 2026-08-26, still open after 13 occurrences).New investigations to open:
3. Retrieve raw error text (next occurrence or targeted log pull) for the
add_labelsfailure in Auto-Triage Issues run 35673928865 to determine root cause.4. Retrieve raw error text for the
create_issuefailure in PR Sous Chef run 35681273619, and check whether it shares code with the previously-confirmedcreate_discussionundefined-object bug.Process improvement:
5. Stabilize raw per-item error-text retrieval for safe_outputs job failures — it currently succeeds inconsistently across audits, which slows down root-cause confirmation for every new signature.
Work Item Plans
add_labelsfailure (Auto-Triage Issues)create_issueno-result failure (PR Sous Chef)Historical Context
This is the 14th audit in this series (daily since 2026-08-22, with a few gaps). Success rates have ranged 97.4%–100% across that history, with two fully clean days (2026-09-04, 2026-09-11). The dominant recurring issue throughout has been PR Sous Chef's
approve_workflow_run(13 total occurrences now), alongside a similar-shapedpush_to_pull_request_branchallowed-files decline in Design Decision Gate (6 occurrences, last seen 2026-09-05) — both are policy-correct declines miscategorized as hard failures, and both fixes remain unshipped a month later. Several other lower-frequency signatures remain open with insufficient data (Smoke Project bad-credentials, Smoke Issues Jira/Linear credentials, submit_pull_request_review non-PR-context).Metrics & KPIs
approve_workflow_runprotected-files/permission failures — 13 occurrences, open since 2026-08-22 (31 days)add_labels,create_issue)Next Steps
approve_workflow_run— this remains the single most consistently recurring issue in the monitor's history.add_labels,create_issue) before proposing a specific code fix.Automated report generated by the Safe Output Health Monitor. Findings are persisted to
/tmp/gh-aw/cache-memory/safe-output-health/for trend tracking across future audits.Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
api.anthropic.comTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions