[prompt-clustering] Prompt Clustering Analysis - 2026-09-24 #63149
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by Copilot Agent Prompt Clustering Analysis. A newer discussion is available at Discussion #63411. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
Analysis Period: Last 30 days (2026-08-25 to 2026-09-24)
Total Tasks Analyzed: 722 copilot-agent PRs (usable prompt text for all 722)
Clusters Identified: 7 (k chosen via elbow method on TF-IDF features)
Overall Merge Rate: 80.7% (583 merged / 135 closed / 4 open)
Full Analysis Report
General Insights
Cluster Analysis
Cluster 3: General Workflow Fixes & Feature Additions
.github/workflows/*.mdagentic workflows in this repo. Healthy merge rate, typical of well-scoped, incremental changes.</details>tags in generated markdown #63069, Make dev-mode actions folder checkout shallow #63035, Harden docker-sbx install: remove curl|sudo sh root install #62991, Add engine-independent tool profile configuration #62951, Report gateway steering events in audit output #62943Cluster 6: Core Engineering / Internals Hardening
Cluster 4: Copilot/Codex Model & Engine Configuration
Cluster 0: Operational-Value Grading Initiative — OUTLIER
<workflow>" or "Record blocked operational-value study for<workflow>". See root-cause section below — this is a recurring, previously-flagged pattern (also an outlier in the 2026-09-15 report).gh aw logs#60635, Record blocked operational-value study for daily-choice-test #58559, Add operational-value grader for daily-caveman-optimizer #58557, Add operational-value grader for cache strategy remediation #58555, Record blocked operational-value study for daily-byok-ollama-test #58554Cluster 5: Safe-Outputs Pipeline & Workflow Features
safe-outputsMCP server and its consumers (diagnostics, ignore-missing-branch handling, steering-issue refactors). Healthy, typical merge rate.Cluster 2: Infra & Tooling Utilities
pkg/consolewasm/native parity, rate-limit reporting, target-only checkout. Smallest average diff size among healthy clusters but highest review/comment density, suggesting careful infra review.Cluster 1: CI-Failure Auto-Fix Tasks
<name>" — the automated CI-failure-response workflow. Small cluster but reliably merges, indicating this auto-triage flow is working as intended.Success Rate by Cluster
Full Data Table (sample of 20 most recent PRs per cluster, largest clusters abbreviated)
</details>tags in generated markdowngh aw logsFull per-PR cluster assignments (722 rows) were computed but are omitted here for length; ask if a CSV export is wanted.
Key Findings
Lowest-Merge Cluster Root Cause
Supporting evidence:
PR Triage Agenttag:Batch: operational-value-grading,Action: batch_review— confirming these were pre-flagged as a single batch for disposition together, not reviewed individually on merit.aw-valueskill tooperational-value-designerand update grader guidance references #55829 ("Renameaw-valueskill tooperational-value-designer...") did merge — suggesting the per-workflow grader-PR approach was superseded by a redesigned, consolidated skill, and the individually-generated PRs were pruned as no-longer-needed rather than rejected for defects.Recommendations
operational-value-designerskill redesign (per Renameaw-valueskill tooperational-value-designerand update grader guidance references #55829) is finalized — this is currently the single largest source of wasted agent runs (~60/month, ~8.7% of all copilot volume) with a 4.8% success rate.Generated by Prompt Clustering Analysis (Run: 35983267147)
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
api.anthropic.comTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions