[prompt-clustering] Prompt Clustering Analysis - 2026-09-29 #64239
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by Copilot Agent Prompt Clustering Analysis. A newer discussion is available at Discussion #64459. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
Analysis Period: Last 30 days (2026-08-30 → 2026-09-29)
Total Tasks Analyzed: 558 copilot-authored PRs
Clusters Identified: 3 (elbow method capped k at 3; a 4th–7th split was tried but added no separable structure beyond the dominant outlier)
Overall Success Rate: 76.3% (426 merged / 558 total)
Full Analysis Report
General Insights
Cluster Analysis
Cluster 1: Routine fixes, graders & CI maintenance
Cluster 2: Substantial feature & infrastructure work
Cluster 3: Bulk auto-generated "operational-value grading" PRs — OUTLIER
<workflow name>" / "Record blocked operational-value study for<workflow name>") repeated ~58 times, one per workflow in the repo. See root-cause section below.Success Rate by Cluster
Full Data Table (sample)
Representative PRs per cluster (24 shown)
Key Findings
Lowest-Merge Cluster Root Cause
What happened: On 2026-09-04, between 10:42 and 15:32 UTC, ~58 near-identical PRs were opened — one per existing agentic workflow in the repo — each adding an "operational-value grader" or recording a "blocked operational-value study" for that workflow. Every one of these PRs:
Action: batch_review | Batch: operational-value-grading— confirming these were triaged and closed as a batch by an automated agent, not rejected individually on their merits.batch_review)This is not a quality problem with individual PRs — it's a single experimental/meta task (generating one probe PR per workflow to test "operational-value grading" instrumentation) that was, by design or by an automated triage policy, closed in bulk. It should be excluded from any headline "PR success rate" metric, since it represents one coordinated event rather than 58 independent task failures.
Recommendations
Data Limitations
gh-awCLI binary was not present in this environment andghis not authenticated, soaw_info.jsonturn counts could not be collected. This report is based on PR metadata (title, body, comments, reviews, file counts) only.Generated by Prompt Clustering Analysis (Run: 36551654081)
All reactions