Skip to content

[aw-failures] [aw] Failure Investigator (6h) - Issue Group #48898

Description

@github-actions

[aw] Failure Investigator (6h)

Parent issue for grouping related issues from [aw] Failure Investigator (6h).

Sub-issues are automatically linked below (max 64 per parent).

Workflow: [aw] Failure Investigator (6h)

  • expires on Aug 6, 2026, 5:35 AM UTC-08:00


Failure sweep — 2026-08-07 (last 6h)

Fix the Copilot Task-tool policy gate first — it is the only P0 that is 100% reproducible on every retry and blocks any Copilot-engine workflow using subagents. Rebuild the sandbox images second — the CVE gate has been red for 5+ straight days and the fixes already exist upstream.

Cluster table

Severity Cluster Workflows Runs Status
P0 Copilot CLI Task-tool No model available PR Code Quality Reviewer §31140894504, §31147686965 New sub-issue filed
P0 Container image critical CVEs, 5+ days red Daily Container Image Security Scan §31152189112 + 4 prior daily runs New sub-issue filed
P1 Claude Code CLI hits Maximum LLM invocations exceeded (30/30) Step Name Alignment §31147854623 Not filed — dropped by the 2-issue cap this cycle, flagging here so it is not lost
P2 curl: (35) Recv failure: Connection reset by peer downloading AWF checksums jsweep - JavaScript Unbloater §31147511863 Isolated network flake, no action taken
P2 permission_denied / CAPIError 422 at Execute GitHub Copilot CLI The Great Escapi §31149984070 Likely by design — workflow intentionally probes firewall escape paths; blocked network calls are the expected outcome, no action taken

Evidence highlights

  • audit-diff of §31147686965 vs successful baseline §31146244896 shows api.githubcopilot.com:443 at 64 allowed requests in the baseline vs 0 in the failure — the Task-tool subagent call never reaches the network, confirming a client-side policy/entitlement block, not a firewall or connectivity issue.
  • Container scan critical findings span 6 of 7 images (arxiv, ast-grep, context7, grafana, memory, serena) — see the new sub-issue for the full CVE breakdown.
  • Step Name Alignment's failure is unrelated to its actual task (step-name checks); it burned its 30-invocation budget on rate-limited retries (error_status:429) before the Maximum LLM invocations exceeded (30 / 30) guard killed it non-retryably. Worth a follow-up issue if this recurs.

Existing tracking correlation

Checked all 5 open agentic-workflows issues (#49583, #49446, #50687, #48838, #50819) — none cover any of today's clusters; no closures warranted this cycle.

Sub-issues created this cycle

  1. Fix Copilot CLI Task-tool subagent calls — 100% failure with No model available
  2. Rebuild scanned container images — critical CVEs have failed the security gate 5+ days straight> Generated by 🔍 [aw] Failure Investigator (6h) · agent · 213.3 AIC · ⊞ 5.2K ·

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions