[audit-workflows] π©Ί Agentic Workflow Audit β 2026-07-22 (96.2%, healthiest in weeks; Avenger + token-reporting recovered) #47409
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by Agentic Workflow Audit Agent. A newer discussion is available at Discussion #47668. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
π©Ί Agentic Workflow Audit β 2026-07-22
Window: 2026-07-22 08:32β21:12 UTC (~12.7h partial) Β· 314 runs sampled (of ~366 in 24h) Β· 103 distinct workflows
First audit in 16 days (last: 2026-07-06) β the 07-07...07-21 span has no audit coverage.
π Key metrics
main)Per-engine: copilot 198/206 (96.1%) Β· claude 54/57 (94.7%) Β· pi 41/42 (97.6%) Β· codex 7/7 Β· antigravity 1/1 Β· gemini 1/1
β Recoveries (chronic issues that improved)
err-config / no-structured-logs, recurrence 18). First window where it both fired and fully succeeded. Prior "green" days were pauses; this is genuine passing traffic. Watching 2 more audits before closing.token_usage_summaryand per-runrun.TokenUsage), after a null-metric gap spanning ~06-19...07-05.gh-aw binary-not-found for MCPsince 06-15). Single sample.π΄ Failure taxonomy (12 = 11 real + 1 intentional)
All failures are 0-turn (agent never completed a turn) β no agent-logic failures.
A. Execute-CLI 0-turn driver failures β 8 (cross-engine, all singletons)
Fail step:
Execute {GitHub Copilot\|Claude Code} CLIin theagentjob. Same signature as the chroniccopilot-sdk-driver-failuresfamily β but diffuse and low-volume now (no single-workflow cluster), down from ~71% of all fleet fails during 06-30...07-06.B. Safe-output partial-intolerance β 3 (agent green, safe_outputs red)
All three:
agentjob SUCCESS,safe_outputsjob FAILED at Process Safe Outputs. The classicsafe-output-partial-failure-intolerancesignature β one bad safe-output item reds the whole job and discards good agent work.Intentional (excluded): Daily Max Ai Credits Test (copilot, main) β credit-guardrail stress test, designed to fail.
π Trend (last 12 audited windows)
Windows are partial (bridge time-caps) and 07-07...07-21 has no coverage; compare rates, not absolute run counts. Success rate is the clear high point of the series, and the token metric recovered.
π― Recommended actions
matplotlib+seaborn+pandas+numpyinto the audit runner image/venv, or allowlist PyPI for the install step, or switch to a pure-Python SVG renderer. Blocks the required PNG trend charts.safe_outputsjob (partial-apply + warn) so successful agent work (PR CQR, PR Sous Chef) is not reddened.π§ Repo memory updated
audit-history.jsonl(+1) Β·metrics-summary.json(07-22 window) Β·known-issues.json(3 recoveries, 2 recurrences) Β·anomalies.json(Avenger de-escalated, safe-output anomaly added) Β·recommendations.json(+2) Β·workflow-trends.json(snapshot).References:
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
awmgmcpgSee Network Configuration for more information.
All reactions