Agent Performance Report - Week of 2026-08-22 #54798
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-08-23T13:04:31.775Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Executive Summary
metrics/latest.json, 24h window 2026-08-21→22)overall_success_rate: 42.3%metrics/latest.jsonwas collected via GitHub API fallback (agentic-workflows logs tool timed out on every attempt). Token/cost data unavailable (null). Safe-output attribution only covers items carrying thegh-aw-workflow-idmarker.q,ci,cgo) with executed success rates under 30%Running Copilot Code Review(98.9%),pr-sous-chef(97.3%),auto-close-parent-issues(100%),Running Copilot cloud agent(94%)q(0.8% success),ci(23.4%),cgo(29.3%),ai-moderator(47.6% of executed runs, still net-negative),agentic_commands(45.5%, but 550 runs = huge blast radius)Performance Rankings
Top Performing Agents 🏆 (executed ≥10 runs, success rate ranked)
Agents Needing Improvement 📉
qworkflow (Success: 2/261 total runs executed = 0.8%, avg runtime 9.9s)qworkflow is still at 0.8% success today, so the merge did not resolve the gating — the root cause tracked in stale alerts is stale/wrong and needs re-diagnosis, not "wait for merge."qworkflow trigger/condition logic directly (do not rely on Add shared prompt quality gate for plateaued agent-review workflows #43527 lore). Check itson:gating conditions and recent run logs for the actual skip/failure reason.ci(CI workflow) (25/107 executed = 23.4% success, avg 214s)cgo(55/188 executed = 29.3% success, avg 1751s — longest average runtime of any workflow sampled)cgowas "stabilizing" (1/5 AR on Jul 8) — current data shows a regression to 70.7% failure rate among executed runs, and very high runtime (~29 min avg). This may indicate a recurrence of the CI regression flagged in July, or a new issue.ai-moderator(10/21 executed = 47.6%, but total volume 279 with only 21 executed — meaning ~92% of triggers never executute, i.e., gated/skipped, not just failing)agentic_commands(250/550 executed... — actually 550 total, 250 successful, 0 recorded failed, success_rate 0.455, meaning ~300 runs are unaccounted for/skipped) — high blast radius given 550 total invocations/day.Zero-Executed / Gated Workflows (not scored as failures per methodology)
squad: 423 total, 0 executed (100% skipped/gated) — likely slash-command workflow with narrow trigger conditions (squad-implement-workersimilarly 106 total / 0 executed).deployment-incident-monitor: 93 total, 0 executed.workflow-generator: 8 total, 0 executed.Behavioral Patterns
q: pattern =under-creation(near-total non-execution, 0.8% success) — persists >6 weeks past the fix PR merge date recorded in memory, indicating the prior root-cause attribution was incorrect.cgo: pattern =inconsistency(memory shows oscillation between "escalated 100% AR" → "stabilizing 1/5 AR" → now 70.7% failure again over successive runs).ai-moderator: pattern =under-creation(92% of triggers never execute) — consistent with previously-reported Codex engine 404s.design-decision-gate: prior memory pattern wasscope-creep/full-AR deprecation candidate; current data shows no matching pattern (95.5% success) — flag as a stale alert requiring memory correction, not a live pattern.shared-alerts.md.Coverage / Ecosystem Notes
create_issue(52) overcreate_pull_request(15), with zeroadd_comment/create_discussionoutputs recorded — worth checking whether comment/discussion safe-output types are under-utilized or simply unattributed (per the collection note, PRs without thegh-aw-workflow-idmarker aren't attributed).Recommendations
High Priority
qworkflow gating (0.8% success, 261 runs) — the previously-cited fix (PR Add shared prompt quality gate for plateaued agent-review workflows #43527) merged 6+ weeks ago without resolving the issue; treat prior root-cause note as stale.ai-moderatorengine switch status — confirm whether Codex→Copilot engine change was actually applied; if not, apply it now.cgoregression (29.3% success, longest avg runtime ~29min) — escalate to Workflow Health Manager, do not rely on "stabilizing" note from prior cycle.Medium Priority
design-decision-gate— current 95.5% success contradicts the "100% AR / deprecation candidate" label carried inshared-alerts.md; update tracking to avoid wrongly deprecating a now-healthy agent.agentic_commandsunaccounted-for runs (~300/550, 55%) to distinguish intentional gating from silent failures.Low Priority
squad,squad-implement-worker,deployment-incident-monitor,workflow-generatorzero-executed patterns are intentional trigger gating, not broken conditions.Trends (vs. daily snapshots)
overall_success_rate: 95.0% (Aug 17) → 84.4% (Aug 20) → 83.6% (Aug 21) → 42.3% (Aug 22, today) — a sharp decline, though today's figure reflects the API-fallback collection method noted above and may not be directly comparable; flagging for cross-validation once the logs tool is available again.active_workflows: 62 (Aug 17) → 66 (Aug 20) → 154 (Aug 21) → 247 (Aug 22) — rising activity volume.total_safe_outputs: 368 → 319 → 718 → 67 — large swing, likely reflecting the attribution-coverage caveat in today's collection note rather than a true output collapse.Actions Taken This Run
qworkflow re-diagnosis (stale root-cause correction).agent-performance-latest.mdandshared-alerts.mdin shared memory with corrected findings (see Do-Not-Re-File list below — do not re-fileai-moderator,cgo,design-decision-gateper existing IDs, but do note thedesign-decision-gatestatus correction).Next Steps
cgo/cireliability investigation.ai-moderator's Codex→Copilot engine switch (recommended repeatedly since July) has shipped; escalate if not.design-decision-gateclassification in shared memory to reflect recovery.All reactions