You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Full ~24h audit (444 runs, last cycle was 2026-07-06 — a 45-day audit-cadence gap). Fleet success rate: 80.6% raw / 80.7% adjusted (excl. intentional-failure guardrail workflows Daily Credit Limit Test and Daily Max Ai Credits Test — only 3 runs matched, so the exclusion barely moves the needle this cycle). The dominant new signal is a fleet-wide codex-engine outage (0/18, 100% fail) — the rest of the fleet is broadly on the established chronic baseline, with two other notable escalations below.
Key metrics
Engine
Runs
Success rate
codex
18
0%⚠️
copilot
217
85.7%
claude
65
87.7%
pi
116
93.1%
0 missing-tool / missing-data / MCP-failure reports this cycle.
Success rate sits at 80.6%, a step down from the 86.25% recorded on 2026-07-06 — driven almost entirely by the new codex outage rather than a broad-based decline. The 45-day gap between audit points is visible as empty space in the trend; day-to-day granularity is lost for that stretch, so this reading is a two-point comparison, not a trend line.
Daily AIC (16,508.6) is broadly in line with the pre-gap baseline (~29.6K on the last full day observed, prior to several partial-window readings), used here in preference to raw tokens since the fleet's TokenUsage aggregation gap (tracked as token-usage-reporting-gap) still affects reliability of the raw token metric. No single-workflow cost spike stands out this cycle.
Findings
Collection phase — outliers and hard failures
Codex engine, fleet-wide outage (NEW/CRITICAL): 18/18 codex runs failed, spanning 10 workflows (Daily Cache Strategy Analyzer, AI Moderator, Daily Fact, Changeset Generator, Smoke Codex, Daily Compiler Threat Spec Optimizer, Daily AW Cross-Repo Compile Check, Daily Firewall Logs Collector and Reporter, Weekly Workflow Analysis, Copilot Session Insights). All fail 0-turn at "Execute Codex CLI" with null token/AIC — e.g. §32330666822 (AI Moderator). Every AI Moderator invocation in the window failed — zero effective moderation coverage for 24h.
Daily Rendering Scripts Verifier — npm firewall block storm:§32346060360 generated 2,436,385 blocked requests to registry.npmjs.org in 36.8 minutes (99.98% of all requests), hit its 30-min timeout, and failed via driver_exit. Root cause: the workflow's tools.bash allowlist includes npm*/npx *, but allowed_domains is restricted to *.grafana.net, *.sentry.io, defaults — no npm registry host.
Ponytail Reviewer: 11/23 (47.8%) fail, 100% at "Execute GitHub Copilot CLI" 0-turn, e.g. §32314732184.
Daily Go Test Parallelizer: 5/8 (62.5%) fail, same signature, also the fleet's most expensive failing runs (270–360K tokens burned before the CLI step reds).
Auto-Triage Issues: 3/4 fail on pi/copilot-gpt-5.4 (experimental), e.g. §32317125168 — a relapse after a brief full recovery on 2026-07-04.
Design Decision Gate 🏗️: 3/4 fail at 15–17min with real token usage (driver_exit) — matches the longrun-fail signature, new addition to that family.
Data-quality gap: "Daily Max Ai Credits Test"'s guardrail-trip run was not tagged intentional_failure=true (unlike its sibling "Daily Credit Limit Test"), requiring a name-regex workaround in this cycle's rollup.
Clustering phase — known issues vs. novel anomalies
codex-gh-aw-binary-not-found-for-mcp — escalated to CRITICAL, from one chronic workflow to a total engine outage (recurrence 6).
claude-agent-job-fail-longrun — new severe instance (npm firewall storm) + Design Decision Gate added to affected list.
copilot-sdk-driver-failures — recurrence 25; Ponytail Reviewer and Daily Go Test Parallelizer now lead this family since Smoke CI no longer exists under that name (it appears restructured into ~13 per-engine "Smoke X" workflows) — smoke-ci-copilot-cli-100pct-fail-on-push is marked RESOLVED-OBSOLETE.
safe-output-partial-failure-intolerance — PR Sous Chef (5/71 fail at "Process Safe Outputs"); engine attribution corrected to pi (all 71 runs), not copilot as previously logged.
pi-gpt54-0tok-agentjob-fail — flipped from RECOVERED-WATCH back to OPEN (Auto-Triage Issues relapse).
New entry: daily-max-ai-credits-test-mistagged (LOW severity, data-quality).
Recommendation phase — smallest actionable fix set
P0 — Fix the codex MCP-helper binary path fleet-wide. Same fix proposed 2026-06-15 for one workflow, never applied; now blocks 10 workflows and all codex-engine coverage including AI Moderator.
HIGH — Add registry.npmjs.org to allowed_domains (or add retry backoff) for any workflow whose tools.bash allowlist includes npm*/npx *. Audit other workflows for the same gap before it repeats elsewhere.
HIGH — Instrument "Execute GitHub Copilot CLI" stderr/exit-code. Long-open ask (recurrence 25); Ponytail Reviewer and Daily Go Test Parallelizer are now the clearest reproduction cases.
HIGH — Re-pin Auto-Triage Issues off pi/gpt-5.4 to a stable model; the 07-04 recovery did not hold.
MEDIUM — Make safe_outputs job failures non-fatal when the agent step already succeeded (PR Sous Chef, pi engine).
LOW — Fix intentional_failure auto-tagging for the AI-credit guardrail path.
Next actions
Codex outage (#1) is the clear top priority — it's a full engine outage affecting 10 workflows including moderation coverage, not a routine chronic-workflow recurrence. Repo memory (known-issues.json, workflow-trends.json, recommendations.json, anomalies.json, metrics-summary.json, audit-history.jsonl) has been updated with all findings above, cross-referenced against prior cycles, and validated within storage limits (75 KB / 7 files).
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Overview
Full ~24h audit (444 runs, last cycle was 2026-07-06 — a 45-day audit-cadence gap). Fleet success rate: 80.6% raw / 80.7% adjusted (excl. intentional-failure guardrail workflows Daily Credit Limit Test and Daily Max Ai Credits Test — only 3 runs matched, so the exclusion barely moves the needle this cycle). The dominant new signal is a fleet-wide codex-engine outage (0/18, 100% fail) — the rest of the fleet is broadly on the established chronic baseline, with two other notable escalations below.
Key metrics
Charts
Success rate sits at 80.6%, a step down from the 86.25% recorded on 2026-07-06 — driven almost entirely by the new codex outage rather than a broad-based decline. The 45-day gap between audit points is visible as empty space in the trend; day-to-day granularity is lost for that stretch, so this reading is a two-point comparison, not a trend line.
Daily AIC (16,508.6) is broadly in line with the pre-gap baseline (~29.6K on the last full day observed, prior to several partial-window readings), used here in preference to raw tokens since the fleet's
TokenUsageaggregation gap (tracked astoken-usage-reporting-gap) still affects reliability of the raw token metric. No single-workflow cost spike stands out this cycle.Findings
Collection phase — outliers and hard failures
registry.npmjs.orgin 36.8 minutes (99.98% of all requests), hit its 30-min timeout, and failed viadriver_exit. Root cause: the workflow'stools.bashallowlist includesnpm*/npx *, butallowed_domainsis restricted to*.grafana.net,*.sentry.io,defaults— no npm registry host.intentional_failure=true(unlike its sibling "Daily Credit Limit Test"), requiring a name-regex workaround in this cycle's rollup.Clustering phase — known issues vs. novel anomalies
codex-gh-aw-binary-not-found-for-mcp— escalated to CRITICAL, from one chronic workflow to a total engine outage (recurrence 6).claude-agent-job-fail-longrun— new severe instance (npm firewall storm) + Design Decision Gate added to affected list.copilot-sdk-driver-failures— recurrence 25; Ponytail Reviewer and Daily Go Test Parallelizer now lead this family sinceSmoke CIno longer exists under that name (it appears restructured into ~13 per-engine "Smoke X" workflows) —smoke-ci-copilot-cli-100pct-fail-on-pushis marked RESOLVED-OBSOLETE.safe-output-partial-failure-intolerance— PR Sous Chef (5/71 fail at "Process Safe Outputs"); engine attribution corrected to pi (all 71 runs), not copilot as previously logged.pi-gpt54-0tok-agentjob-fail— flipped from RECOVERED-WATCH back to OPEN (Auto-Triage Issues relapse).daily-max-ai-credits-test-mistagged(LOW severity, data-quality).Recommendation phase — smallest actionable fix set
registry.npmjs.orgtoallowed_domains(or add retry backoff) for any workflow whosetools.bashallowlist includesnpm*/npx *. Audit other workflows for the same gap before it repeats elsewhere.intentional_failureauto-tagging for the AI-credit guardrail path.Next actions
Codex outage (#1) is the clear top priority — it's a full engine outage affecting 10 workflows including moderation coverage, not a routine chronic-workflow recurrence. Repo memory (
known-issues.json,workflow-trends.json,recommendations.json,anomalies.json,metrics-summary.json,audit-history.jsonl) has been updated with all findings above, cross-referenced against prior cycles, and validated within storage limits (75 KB / 7 files).References:
All reactions