Skip to content

[agentrx-optimizer] Daily Workflow Optimization - 2026-08-03 #49918

Description

@github-actions

Executive Summary

AgentRx analyzed the last 24h of gh-aw agent run telemetry (14 completed runs, all success, 0 errors/warnings/missing-tools reported at the safe-outputs layer). Because no runs failed outright, the highest-impact finding is a cost bottleneck, not a failure: the Issue Arborist workflow's single "Fetch issues" step drives one LLM request that consumes 821,982 tokens — 66.9% of the entire 14-run fleet's token usage in the sample window, and 26.3x the average of every other run. The AgentRx invariant-checking stages (static/dynamic/check/judge) could not complete in this sandbox because they require an authenticated LLM endpoint (see Artifacts below), so this recommendation is grounded directly in raw run telemetry (run_summary.json / agentic_assessments) rather than a completed AgentRx judge classification.

AgentRx Evidence

  • Critical step: "Fetch issues" step, Issue Arborist workflow (§30789333991), .github/workflows/issue-arborist.md:49-54
  • Failure category: token-heavy context payload (single-request context bloat, not an error/failure)
  • Frequency / impact: 1 run in the 24h sample, but it alone accounts for 821,982 of 1,228,121 total fleet tokens (66.9%); 802,207 of those are input tokens delivered in a single LLM request (total_requests: 1), of which 749,056 were cache-read — meaning ~53K tokens of genuinely fresh input are paid every single day just to re-load the same shape of data, on top of a 749K-token cache the model must still process.
  • Representative run IDs: §30789333991 (Issue Arborist, 821,982 tokens) vs. fleet median ≈ 27,897 tokens (The Great Escapi) and fleet average of all other 13 runs ≈ 31,241 tokens.
AgentRx Artifacts

IR stage: Failed. AgentRx ships three built-in domain converters (tau/retail, magentic, flash); none recognize gh-aw's run-summary trajectory shape, so the pipeline fell back to its LLM-based IR converter (llm_ir). That fallback calls the copilot endpoint by default, which requires gh copilot OAuth/PAT authentication (COPILOT_GITHUB_TOKEN/GH_TOKEN) not available in this sandboxed run — azure and trapi endpoints are similarly unconfigured (no AGENT_VERIFY_ENDPOINT / TRAPI vars). Result: state.json shows "completed_stages": [] for this run — no static, dynamic, check, judge, or report artifacts exist.

Packaging blocker discovered along the way: the AgentRx PyPI/git install ships only the agentrx* Python package per its pyproject.toml ([tool.setuptools.packages.find] include = ["agentrx*"]); the actual pipeline entrypoint run.py lives at the repo root and is never included in the wheel, so pip install (github.com/redacted) alone leaves the agentrx console-script non-functional (Error: run.py not found). Worked around by fetching run.pydirectly from the upstream repo for this run — reported viamissing_tool`.

Substitute evidence used instead: gh-aw's own per-run run_summary.json already includes an agentic_assessments field (a lightweight built-in heuristic, not AgentRx's judge) that independently flagged:

  • Issue Arborist / Issue Monster (both issue_response domain): model_downgrade_available (low) — smaller model may suffice. Note: token_usage_summary.by_model shows Issue Arborist already ran on gpt-5.4-mini, so this specific low-severity hint is largely moot; the real cost driver is payload size, not model tier.
  • Safe Output Health Monitor (105,668 tokens, 32m36s, 0 tracked turns/tools/write-actions): resource_heavy_for_domain (high)
  • Static Analysis Report (37,389 tokens, 31m43s, 0 tracked turns/tools): resource_heavy_for_domain (high)
  • ESLint Refiner (47,392 tokens, 16m47s): resource_heavy_for_domain (medium)

All three resource_heavy_for_domain runs report turns=0 tool_types=0 write_actions=0 alongside 16–33 minute durations — the run-level telemetry can't explain where that wall-clock time or token spend goes, which is itself a telemetry gap (see classification table).

Violation classification (derived manually from raw telemetry in place of AgentRx's check/judge stages, which did not complete):

Violation Evidence Fix Type Rationale
Issue Arborist: 821,982 tokens / 802,207 input tokens in 1 LLM request §30789333991; .github/workflows/issue-arborist.md:49-54 fetches number,title,author,createdAt,state,url,body,labels,updatedAt,closedAt,milestone,assignees for up to 100 open issues, then instructs the agent (lines 116-119) to cat/jq the whole file into context Reducing token-heavy context payloads Full issue bodies for 100 issues are loaded every run, but only the small subset that ends up in a linking decision needs full text — 66.9% of fleet-wide tokens for one daily workflow
Safe Output Health Monitor: high resource_heavy_for_domain, 0 tracked turns/tools over 32m36s §30786933137 agentic_assessments Adding missing telemetry attributes for better triage Duration and token spend can't be attributed to specific turns/tools, blocking further diagnosis
Static Analysis Report: high resource_heavy_for_domain, 0 tracked turns/tools over 31m43s §30788331908 agentic_assessments Adding missing telemetry attributes for better triage Same telemetry gap

Known limitations: No AgentRx static_invariants.json, dynamic_invariants/, checker_results/, or judge_output/ exist for this run — the pipeline never advanced past ir. Nothing above should be read as an AgentRx invariant violation or LLM-judge root-cause classification; it is a direct read of gh-aw's own run summaries.

Recommended Optimization

Change: In .github/workflows/issue-arborist.md, split the "Fetch issues" step into two passes instead of one full-body dump:

  1. Fetch a lean field set for all ~100 candidate issues (number,title,labels,createdAt,updatedAt,milestone,assignees — no body) to do the clustering/relationship pass.
  2. Only fetch body for the smaller shortlist of issues the agent flags as plausible cluster members (e.g., via a second gh issue view (n) --json body per shortlisted issue, or a follow-up gh issue list scoped by number).

Why highest impact: This single step is responsible for two-thirds of all tokens burned across the fleet in the sampled 24h window (821,982 of 1,228,121) — 26x any other run. body text (often multiple paragraphs) is the single largest per-issue field and is fetched for all 100 issues even though most won't end up in any cluster; trimming it from the bulk fetch should cut this run's dominant token cost by a large margin while leaving the clustering heuristic's inputs (title/labels/timestamps) intact.

Where to implement: .github/workflows/issue-arborist.md:49-54 (the gh issue list ... --json ... fetch) and the accompanying prompt guidance at lines 113-119 that tells the agent to load the whole file.

Validation Plan

  • Re-run logs/audit on the next Issue Arborist scheduled run and compare token_usage / token_usage_summary.total_input_tokens against this run's 821,982 / 802,207 baseline.
  • Success metric: total tokens for Issue Arborist drop materially (target: bring it back in line with the other 13 runs' ~5K–105K range) while evals.sub_issues_linked_or_noop and issues_analyzed outcomes stay unchanged (no loss of linking quality).
  • Re-attempt the AgentRx check/judge stages once an authenticated LLM endpoint (copilot login or Azure/TRAPI env vars) is available in the runner, so future runs get a real invariant-based classification instead of this telemetry-only substitute.

References

  • §30789333991 — Issue Arborist (821,982 tokens, primary evidence)
  • §30786933137 — Safe Output Health Monitor (high resource_heavy_for_domain)
  • §30788331908 — Static Analysis Report (high resource_heavy_for_domain)

Generated by ⚡ Daily AgentRx Trace Optimizer · agent · 136.2 AIC · ⌖ 18 AIC · ⊞ 7.4K ·

  • expires on Aug 9, 2026, 10:47 PM UTC-08:00

Metadata

Metadata

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions