[prompt-clustering] Prompt Clustering Analysis - 2026-09-10 #59939
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-09-11T10:09:39.955Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
Analysis Period: 2026-08-17 to 2026-09-10 (last 1,000 copilot-authored PRs, ~24 days — the search API cap was hit before reaching a full 30 days back)
Total Tasks Analyzed: 1,000
Clusters Identified: 6
Overall Success Rate: 81.0% (810 merged / 188 closed / 2 open)
Full Analysis Report
General Insights
Cluster Analysis
Cluster 0: General workflow & GitHub Actions fixes
gh-awworkflow engine — logs command, safe-outputs, enclave/MCP plumbing.gh aw logs#59859, Improve logs command cache diagnostics and Drain3 training #59823, Clarify workflow source 404s in log consumers #59820, Fix invalid empty GitHub server guard policy for enclave-only GitHub tools #59816Cluster 1: Failing CI job auto-fixes
[WIP] Fix failing GitHub Actions job <name>PRs — targeted, single-purpose, and merge cleanly every time.Cluster 2: "PR Sous Chef" bot automation
/souschefiterations before merge.permissions.contents: noneskips default checkout #59743, Retry transient slash-command provenance checks #59313, Cache GitHub Actions job metadata in logs output #59039, Register replace_label handler in safe-output collect job dispatch map #58917, Add deterministic body footer templates to safe outputs #58841Cluster 3: Operational-value grader sweep (OUTLIER)
copilot/operational-value-study-paper-v1*) tasking the agent to design an "operational-value grader" for dozens of other workflows one PR at a time. See root-cause section below.Cluster 4: Firewall / network / audit logging
Cluster 5: Model/engine inventory & version management
Success Rate by Cluster
Full Data Table (sample of 30 most recent PRs)
gh aw logsaw-valueskill tooperational-value-designer(Full 1,000-row dataset was processed for clustering; this table shows a representative sample. See
cluster-summary.jsonin cache for the complete per-cluster PR lists.)Key Findings
Lowest-Merge Cluster Root Cause
59 of 61 PRs in this cluster have zero review comments and zero formal reviews — they were closed without any human ever weighing in. Titles are explicit about the outcome: "Record blocked operational-value study for X (no defensible direct metric)", "Block operational-value grader design for Y". Branch names (
copilot/operational-value-study-paper-v1[-again]) confirm this is one coordinated sweep, not organic task variety: the agent was dispatched once per target workflow to attempt designing a deterministic "operational value" grading metric, and in the large majority of cases concluded the metric wasn't definable and self-closed the PR rather than merging speculative grading logic. Only 2 PRs escaped this pattern, and neither is a grader itself — one documents workflow design intent (#57005), the other renames the skill driving the sweep (#55829).Recommendations
Generated by Prompt Clustering Analysis (Run: 34462352293)
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
api.anthropic.comTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions