[audit-workflows] π Agentic Workflow Audit β 2026-07-24 (93.7% Β· Smoke CI recovered, pi degraded) #47867
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-07-25T21:53:39.545Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
π Agentic Workflow Audit β 2026-07-24
Full 24h window (
2026-07-23T21:59Z β 2026-07-24T21:17Z). 305 runs had complete telemetry (of 565 downloaded run dirs; 260 download-incomplete β the usual bridge/pagination cap). This is the healthiest recorded day to date, with two notable chronic-issue recoveries.π Headline metrics
mainbranch)Intentional-failure workflows excluded from rate:
Daily Credit Limit Test,Daily Max AI Credits Test.β Two chronic issues RECOVERED
push,schedule,pull_request, andworkflow_dispatch. The chronic "100%-fail-on-push" cluster (red onmainsince ~06-30, recurrence count 6) appears resolved. Kept on watch for regression.metrics.TokenUsage/run.TokenUsageis now populated directly fleet-wide (14.49M captured with notoken_usage_summaryfallback needed). The prior empty-metrics gap (count 3) is resolved.π΄ Failure taxonomy β all 20 are 0-turn pre-agent driver fails
Every failure was a 0-turn / pre-agent job failure. Zero agent-logic errors, missing tools, missing data, or MCP failures.
Cluster 1 β PR Sous Chef (pi/copilot-gpt-5.4): 8 fails Β· DOMINANT
pi-gpt54-0tok-agentjob-failβ pi runs thecopilot/gpt-5.4backend, so it's the same copilot driver family surfacing via pi.Cluster 2 β copilot claude-sonnet-4.6 longrun 0-turn: 10 singletons
Known issue
copilot-sdk-driver-failures(recurrence count 25). One failure per workflow, several after substantial real runtime:These burn compute for up to 47 minutes before the driver drops the job β the main remaining source of wasted AIC.
Cluster 3 β Agent Persona Explorer (pi): 1 fail
0/1, 23.4m 0-turn. Same
pi-gpt54-0tok-copilot-backedfamily as PR Sous Chef.π Trend charts (30 days)
Success rate closes the 30-day window at its highest point (93.7%), up from the 86β90% baseline of early July. The flat segment 07-06 β 07-24 reflects the audit gap (no data), not a health change. Run volume for this partial-telemetry day (305) is lower than full-fetch days (~400) due to download-incompleteness, not fewer runs.
AIC for the captured window (16.5K) sits mid-range against the 07-04β07-06 ramp (6.3Kβ29.6K). The 7-day moving average is only indicative given the 18-day gap. Top consumers are all successful claude workflows (Typist 744, Go Logger 626, VulnHunter 613) β cost is concentrated in legitimate heavy analysis, not failures.
π― Recommendations
copilot/gpt-5.4backend β PR Sous Chef (75%) + Agent Persona Explorer are the top prod-main failure source now that Smoke CI is fixed. Pin these pi-backed workflows off the flakycopilot/gpt-5.4backend, or add a 0-turn driver-failure retry. (rec-pin-autotriage-stable-model)smoke-ci-copilot-cli-100pct-fail-on-push.π§Ύ Repo memory
Updated:
audit-history.jsonl,metrics-summary.json,known-issues.json(Smoke CI + token-gap β RECOVERED-WATCH; copilot-sdk count 25; pi-0turn re-escalated count 2),recommendations.json,anomalies.json,workflow-trends.json. Validated at 57 KB.References:
All reactions