[audit-workflows] π©Ί Agentic Workflow Audit β 2026-07-21 Β· Healthy 93.4%, Avenger+Smoke CI recovered, 11.5M-token burn anomaly #47150
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by Agentic Workflow Audit Agent. A newer discussion is available at Discussion #47409. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
π©Ί Agentic Workflow Audit β 2026-07-21
Fleet health: HEALTHY β 93.4% (282/302 runs), the best overall rate since ~06-24. All 20 failures were pre-agent driver failures (
Turns=0); zero agent-logic, missing-tool, missing-data, MCP, or noop errors. Two chronic dominant offenders recovered this window.Key metrics
Turns=0(pre-agent)Engine success: copilot 94.0% (174/185) Β· claude 87.0% (54/62) Β· pi 97.6% (42/43) Β· codex 10/10 Β· gemini 1/1 Β· antigravity 1/1.
π’ Recoveries (3 chronic issues cleared this window)
err-config / no-structured-logs, recurrence 18) and the single largest reliability drain on the fleet. First clean run observed. Watch to confirm.run.TokenUsagenow populated for 291/302 runs (was fleet-wide null, recurrence 3).π΄ Top anomaly β 11.49M-token burn on a failed run
Daily Sub-Agent Model Resolution Audit(run29806493363,copilot/gpt-5.4-mini) consumed 11,490,620 tokens β 49% of the entire fleet's daily spend β then the agent job failed atTurns=0(agentic_fraction=0) after 591s. No agentic turn was ever produced despite 11.5M tokens, strongly implying a driver-level retry/loop burning context pre-turn. No token/cost guardrail intervened. A single failed run doubled fleet token usage (ex-anomaly fleet = 11.98M).Failure clustering (all 20 fails are
Turns=0driver failures)copilot-sdk-driver-failures(recur 25)claude-agent-job-fail-longrunworkflow_dispatchoncopilot/fix-claude-configuration/aw-failures-fix-smoke-claude/update-model-aliasesfix branchespi-gpt54-0tok-agentjob-fail(recur 2)Cross-engine signature confirmed: the copilot-longrun (8) and claude-longrun (3) failures share the identical 0-turn-after-real-runtime signature β this is a driver-level root cause, not engine-specific. The Smoke-Claude cluster is active debugging on fix branches, not a production regression (main-branch Smoke Claude had only 1 dispatch fail).
π 30-day trends
Health has climbed back to 93.4%, above the 90% line and the highest since the 06-24 peak. The visible line jump reflects the 15-day observation gap (07-07 β 07-20), not a data cliff β no runs were lost, they simply weren't audited.
The 07-21 bar (23.5M) is inflated β a single failed run burned 11.5M of it. Ex-anomaly, the fleet's real spend (~12.0M) is in line with the early-July baseline. Token history remains sparse because
TokenUsagewas null on most prior days (that gap is now resolved going forward).Recommendations (evidence-linked)
Daily Sub-Agent Model Resolution Audit(copilot/gpt-5.4-mini). Evidence: run29806493363βTokenUsage=11,490,620,Turns=0,conclusion=failure, agent dur 591s,agentic_fraction=0. Investigate the gpt-5.4-mini driver path for a pre-turn retry loop and cap token spend so a failing run cannot double fleet usage.Turns=0-after-real-runtime driver failure (copilot-sdk-driver-failures, recur 25). 11 of 20 fails this window (8 copilot + 3 claude) share it. Instrument the driver hand-off; the shared signature across engines points to the AWF/CLI driver layer, not model config.TidyΓ2 fail on push/main β a new push-triggered variant of the copilot-longrun cluster; deterministic red on main is worth isolating.avenger-err-config-no-structured-logsandsmoke-ci-copilot-cli-100pct-fail-on-push.Repo memory updated
audit-history.jsonl(07-21 entry),metrics-summary.json,known-issues.json(recur bumps + 2 new + 3 recoveries),anomalies.json(11.5M token-burn + Avenger recovery),recommendations.json(token-guardrail rec),workflow-trends.json. Memory validated within limits (57 KB / 60 KB).References: Β§29806493363 (11.5M-token anomaly)
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
awmgmcpgSee Network Configuration for more information.
All reactions