You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Window evaluated: last 24 full hours (UTC), 2026-09-12T23:19Z → 2026-09-13T23:19Z
Total runs analyzed: 405 (with complete engine/job telemetry); 167 additional runs had no aw_info.json/run_summary.json (agent job never executed — e.g. activation-only or skipped triggers) and are excluded from the detection comparison below
Detection-enabled runs: 364 (89.9%)
Regular runs (gh-aw-detection: false): 41 (10.1%)
Misconfigured workflows found: 0
Note
No misconfigured gh-aw-detection workflows were found in this window. All 41 "regular" runs are one-off smoke-test workflows (e.g. Smoke Pydantic AI, Smoke Aider, Smoke OpenCode) intentionally exercising engines with detection off; none exceeded the >3-run threshold, none showed a failing detection job while enabled, and no workflow flip-flopped its gh-aw-detection setting within the window.
Comparison Chart
Regular runs skew heavily toward one-off engine smoke tests (low n, low observed success rate), while detection runs represent the bulk of production scheduled/triggered workflows. Token usage could not be compared — every run in this log snapshot reports TokenUsage: 0 in run_summary.json, which looks like a metrics-pipeline gap upstream of this analysis rather than an actual detection-vs-regular difference (worth a separate look, see Recommendations).
Misconfigured Workflows
No misconfigured workflows detected in this window.
View All Run Metrics (405 runs across 173 workflows)
Workflow
Total Runs
Detection Runs
Regular Runs
Success Rate
Avg Tokens
Engine(s)
PR Sous Chef
80
80
0
31.2%
0
pi
Issue Monster
46
46
0
21.7%
0
pi
Daily Go Test Parallelizer
12
12
0
91.7%
0
codex
Design Decision Gate 🏗️
11
11
0
100.0%
0
pi
Impeccable Skills Reviewer
11
11
0
100.0%
0
copilot
Matt Pocock Skills Reviewer
11
11
0
100.0%
0
copilot
PR Code Quality Reviewer
11
11
0
0.0%
0
copilot
Ponytail Reviewer
11
11
0
100.0%
0
codex
Test Quality Sentinel
11
11
0
100.0%
0
copilot
Avenger
10
10
0
90.0%
0
codex
Auto-Triage Issues
7
7
0
100.0%
0
codex
Contribution Check
5
5
0
100.0%
0
copilot
Code Scanning Fixer
4
4
0
0.0%
0
copilot
Deep Report
4
4
0
100.0%
0
claude
PR Triage Agent
4
4
0
0.0%
0
copilot
[aw] Failure Investigator (6h)
4
4
0
75.0%
0
claude
Release
3
3
0
100.0%
0
copilot
Smoke Copilot
3
0
3
0.0%
0
copilot
AI Moderator
2
2
0
0.0%
0
pi
Daily Credit Limit Test
2
2
0
50.0%
0
pi
Agent Container Smoke Test
1
0
1
0.0%
0
pi
Agent Job Health Monitor
1
1
0
0.0%
0
claude
Agent Performance Analyzer - Meta-Orchestrator
1
1
0
100.0%
0
copilot
Agent Persona Explorer
1
1
0
100.0%
0
codex
Agentic Workflow Audit Agent
1
1
0
100.0%
0
codex
Artifacts Usage Report
1
1
0
100.0%
0
codex
CLI Version Checker
1
1
0
100.0%
0
pi
Claude Code User Documentation Review
1
1
0
100.0%
0
claude
Code Simplifier
1
1
0
0.0%
0
copilot
Constraint Solving — Problem of the Day
1
1
0
100.0%
0
copilot
Copilot Agent PR Analysis
1
1
0
0.0%
0
claude
Copilot Agent Prompt Clustering Analysis
1
1
0
0.0%
0
claude
Copilot CLI Deep Research Agent
1
1
0
100.0%
0
copilot
Copilot PR Prompt Pattern Analysis
1
1
0
100.0%
0
copilot
Copilot Session Insights
1
1
0
100.0%
0
claude
Daily A/B Testing Advisor
1
1
0
0.0%
0
codex
Daily AWF Spec Compiler Surfacing Review
1
1
0
100.0%
0
codex
Daily Ambient Context Optimizer
1
1
0
100.0%
0
copilot
Daily Assign Issue To User
1
1
0
100.0%
0
copilot
Daily AstroStyleLite Markdown Spellcheck
1
1
0
100.0%
0
claude
Daily CLI Performance Agent
1
1
0
100.0%
0
codex
Daily Cache Strategy Analyzer
1
1
0
100.0%
0
codex
Daily Caveman Optimizer
1
1
0
100.0%
0
claude
Daily Cli Tools Tester
1
1
0
100.0%
0
codex
Daily Code Debt Cleanup — Aider
1
1
0
0.0%
0
aider
Daily Code Metrics and Trend Tracking Agent
1
1
0
0.0%
0
copilot
Daily Community Attribution Updater
1
1
0
0.0%
0
copilot
Daily Compiler Quality Check
1
1
0
100.0%
0
copilot
Daily Compiler Threat Spec Optimizer
1
1
0
100.0%
0
copilot
Daily Container Image Security Scan
1
1
0
0.0%
0
copilot
Daily Documentation Diagram
1
1
0
100.0%
0
codex
Daily Documentation Healer
1
1
0
100.0%
0
claude
Daily Documentation Updater
1
1
0
0.0%
0
codex
Daily Evals Feature Report
1
1
0
100.0%
0
codex
Daily Formal Spec Verifier
1
1
0
100.0%
0
copilot
Daily GitHub Docs SEO Optimizer
1
0
1
100.0%
0
copilot
Daily Go Function Namer
1
1
0
100.0%
0
pi
Daily Go Test Stubs — Aider
1
1
0
0.0%
0
aider
Daily Grader Audit
1
1
0
100.0%
0
claude
Daily Graft Intelligence
1
1
0
100.0%
0
copilot
Daily Issues Report Generator
1
1
0
100.0%
0
copilot
Daily Malicious Code Scan Agent
1
1
0
100.0%
0
copilot
Daily Max Ai Credits Test
1
1
0
100.0%
0
codex
Daily Model Inventory Checker
1
1
0
100.0%
0
copilot
Daily Observability Report for AWF Firewall and MCP Gateway
Detection-run volume has climbed steadily from ~250-430/day over the past ~3 weeks while regular (opted-out) runs stay near-zero most days, spiking only on days with heavier smoke-test activity (e.g. 2026-09-11: 12 regular runs). Detection success rate has been trending down recently (100% on 2026-09-10 → 35.1% on 2026-09-12 → 60.2% today), which tracks a broader rise in agent job failures across the fleet, not a detection-specific regression.
Recommendations
Investigate the zero-token metrics gap. Every run in this snapshot reports TokenUsage: 0 regardless of detection status, which breaks any token-cost comparison between regular and detection runs. Check whether the log-download/processing pipeline that produces run_summary.json is dropping TokenUsage before these reports are generated.
Watch the declining detection success rate. Detection-run success dropped to 35.1% on 2026-09-12 before recovering to 60.2% today — worth confirming this tracks a known incident (e.g. PR Sous Chef at 31.2% and Issue Monster at 21.7% success over this window are large contributors) rather than a detection-job regression.
No action needed on gh-aw-detection configuration — current opt-outs are all intentional, low-volume smoke tests. Re-run this check after 7 days of continuous logs to properly evaluate Rule 1 (>3 runs/7 days), since this window only covers 24h.
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
api.anthropic.com
To allow these domains, add them to the network.allowed list in your workflow frontmatter:
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
aw_info.json/run_summary.json(agent job never executed — e.g. activation-only or skipped triggers) and are excluded from the detection comparison belowgh-aw-detection: false): 41 (10.1%)Note
No misconfigured
gh-aw-detectionworkflows were found in this window. All 41 "regular" runs are one-off smoke-test workflows (e.g.Smoke Pydantic AI,Smoke Aider,Smoke OpenCode) intentionally exercising engines with detection off; none exceeded the >3-run threshold, none showed a failingdetectionjob while enabled, and no workflow flip-flopped itsgh-aw-detectionsetting within the window.Comparison Chart
Regular runs skew heavily toward one-off engine smoke tests (low n, low observed success rate), while detection runs represent the bulk of production scheduled/triggered workflows. Token usage could not be compared — every run in this log snapshot reports
TokenUsage: 0inrun_summary.json, which looks like a metrics-pipeline gap upstream of this analysis rather than an actual detection-vs-regular difference (worth a separate look, see Recommendations).Misconfigured Workflows
No misconfigured workflows detected in this window.
View All Run Metrics (405 runs across 173 workflows)
View Historical Trend
Detection-run volume has climbed steadily from ~250-430/day over the past ~3 weeks while regular (opted-out) runs stay near-zero most days, spiking only on days with heavier smoke-test activity (e.g. 2026-09-11: 12 regular runs). Detection success rate has been trending down recently (100% on 2026-09-10 → 35.1% on 2026-09-12 → 60.2% today), which tracks a broader rise in
agentjob failures across the fleet, not a detection-specific regression.Recommendations
TokenUsage: 0regardless of detection status, which breaks any token-cost comparison between regular and detection runs. Check whether the log-download/processing pipeline that producesrun_summary.jsonis droppingTokenUsagebefore these reports are generated.PR Sous Chefat 31.2% andIssue Monsterat 21.7% success over this window are large contributors) rather than a detection-job regression.gh-aw-detectionconfiguration — current opt-outs are all intentional, low-volume smoke tests. Re-run this check after 7 days of continuous logs to properly evaluate Rule 1 (>3 runs/7 days), since this window only covers 24h.Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
api.anthropic.comTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions