[agent-job-health] Agent Job Health Report — 2026-09-26 (24h window) #63694
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-09-29T23:56:54.962Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Warning
Fleet-wide agent-job failure rate is 17.0% this window — roughly double the late-August baseline (~7-9%) and essentially unchanged from three days ago (17.97% on 2026-09-23), meaning the elevated rate has now persisted across at least two monitoring windows without any tracked remediation. The single largest cluster (Execute GitHub Copilot CLI step failures, 10 workflows) was already flagged as novel and untracked on 2026-09-23 and remains untracked today.
Summary
audit.jsonrecords exist for this window, but 13 lack completeaw_info.json/run_summary.jsondetail — likely runs cancelled before summary generation; this report is based on the 206 runs with complete data)agent-job failures — see "Non-agent-job failures" note below)safe_outputsjob failures, out of this monitor's scope) — its volume dilutes the run-weighted average while masking that a full third of distinct workflows hit at least one agent-job failure.not_planned, auto-expired) and don't match the current signatures precisely.agent-job-health/history.jsonl): 2026-08-23: 9.22% → 2026-08-28: 7.41% → 2026-09-23: 17.97% → 2026-09-26 (today): 17.0%. The jump from August to September (~+9pts) has now held for 3+ days; day-over-day change vs. 09-23 is only -1pt (below the 10pt regression threshold) but the sustained elevation itself is the signal worth acting on.Failure Rate by Step
Tracked Failures
None. No failure signature in this window is covered by an existing open issue. (Related but closed/expired trackers: #54186, #55413 — see Recommendations.)
Novel Failure Clusters
1.
Execute GitHub Copilot CLIstep failure — 14 runs, 10 workflows, engine: copilot (largest cluster — see linked issue)2.
Redact secrets in logsstep failure — 6 runs, 6 workflows, 4 different engines3.
Execute Pi CLIstep failure — 6 runs, 3 workflows, engine: pi4.
Execute Codex CLIstep failure — 4 runs, 1 workflow: Avenger (100% failure rate this window — all 4 of its runs failed)5.
Execute Aider CLIstep failure — 2 runs, 2 workflows: Daily Go Test Stubs — Aider, Daily Code Debt Cleanup — Aider6. Singleton failures (1 run each):
Execute Claude Code CLI(Daily Harness Experiment Proposer),Download workflow logs for PR analysis(Copilot Agent Prompt Clustering Analysis),Download container images(Daily Cache Strategy Analyzer — infra/registry step, not a model step)Non-agent-job failures (out of scope, noted for context only)
Of the 46 runs whose overall Actions run failed, 11 were not agent-job failures:
safe_outputspost-processing job (agent itself succeeded) — all PR Sous Chef; owned by the safe-output-job monitor, not this report.activation(agent job never ran / was skipped): Smoke Issues, Daily Code Metrics and Trend Tracking Agent, Copilot Agent PR Analysis, Daily CLI Performance Agent.conclusion, at a step called "Handle agent failure" — even though theagentjob itself reported success: Test Quality Sentinel (36252091697), PR Code Quality Reviewer (36274158434), Daily Project Performance Summary Generator (36270482565). This is a confusing signal (a job named after handling agent failure, failing itself, while the agent job it's supposedly reacting to succeeded) — flagged in Recommendations.Schedule Heartbeat
64 schedule-triggered workflows found (excluding
shared/includes and workflow_dispatch-only workflows). 23 were checked this window (all sub-daily/6h-or-less, all multi-day 2-3x/week, all weekly, plus an 8-workflow spot-check sample of dailies), prioritized by cadence to catch the highest-risk gaps first.No blind spots detected in the checked sample — every checked workflow had a most-recent run (of any status) well within its expected cadence + slack window. The remaining 41 daily-cadence workflows were not individually checked this cycle; no anomaly signal emerged from the checked sample to suggest a systemic scheduler problem, but full coverage is recommended on a future run.
Per-workflow breakdown (workflows with ≥1 agent-job failure this window)
For contrast, the fleet's highest-volume workflow, PR Sous Chef (pi engine), ran 54 times this window (26% of all runs) with zero agent-job failures — all 4 of its overall-failed runs were
safe_outputs-stage failures, outside this report's scope.Recommendations
Execute GitHub Copilot CLIfailures (10 workflows, 14 runs) — the largest and most persistent cluster, already seen untracked on 2026-09-23. Cross-workflow spread with a single engine suggests a shared root cause (auth/rate-limit/binary issue on the Copilot CLI path) rather than per-workflow bugs. Tracked via the issue linked below.Redact secrets in logsfailures (6 runs, 6 workflows, 4 engines) — happens after the model CLI already succeeded, so this looks like a shared post-processing/log-redaction script bug blocking otherwise-successful runs from completing cleanly. Worth a dedicated look even though it's currently below the single-issue threshold.conclusion/ "Handle agent failure" job semantics — 3 runs this window show this downstream job failing while theagentjob itself reports success (Test Quality Sentinel, PR Code Quality Reviewer, Daily Project Performance Summary Generator). This is a confusing signal that could cause other monitors or dashboards to double-count or misattribute agent failures; worth a tooling fix or at least documentation.not_plannedbefore a fix landed. Consider a longer-lived tracking mechanism for clusters that recur across multiple reporting windows.All reactions