[copilot-session-insights] Daily Copilot Agent Session Analysis — 2026-09-27 #63802
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by Copilot Session Insights. A newer discussion is available at Discussion #63958. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🤖 Copilot Agent Session Analysis — 2026-09-27
Executive Summary
{run_id}-conversation.txt) were completely absent for the 35th consecutive recorded day —find .../logs -name "*-conversation.txt"returned 0 files against 50 listed sessions. This blocks turn-by-turn tool-usage, token-efficiency, loop-detection, and prompt-quality analysis (the "true behavioral analysis" this workflow is designed to produce). Everything below is derived from run metadata only (timestamps, conclusion, workflow name, branch) plus livegh apiqueries against PRs and Actions runs. Flagged viamissing_data.Key Metrics
Success Factors ✅
Failure Signals⚠️
failureat 23:41:45Z (1s later, 3 runs invalidated); (b) PR Report MCP payload sizes in audit and aggregate analysis #63687 merged 00:34:21Z → Squad Implement Worker wentaction_requiredat 00:34:27Z (6s later, 1 run invalidated). Event (b) is the 4th confirmed instance of the "merge → Squad Implement Worker +6–7s" sub-signature (first seen 09-19/09-24, recurred 3× on 09-26 alone) — this now looks like a routine per-merge side effect, not a rare edge case.copilot/dynamic-checkouts-github-actionproduced 20/50 (40%) of all sessions but only 20.0% success — the heaviest gate-noise branch in today's window, tied to PR Support dynamic checkout sets from GitHub Actions expressions #63241 which has been open since 09-24.Prompt Quality Analysis 📝
Per-Prompt Breakdown
No prompt text or task-description content is available in run metadata — session-list entries expose only workflow name, branch, timestamps, and conclusion. Meaningful prompt-quality scoring (high/medium/low characteristics, example prompts) requires the conversation transcripts described in the data-gap notice above and cannot be produced from this window's inputs. This has now been true for 35 consecutive recorded days.
Orphaned Branch Escalation Alerts 🚨
Summary
Escalation Candidate Details
Escalation Candidates
✅ No orphaned branches exceed the escalation threshold today. Only 2 in-progress Actions runs existed repo-wide in the trailing 6 hours, both on
main(this workflow + Tidy) — no branch approached the ≥5-gate threshold.CI Waste Estimate
Notable Observations
Loop Detection and Session Diagnostics
Loop Detection
Tool Usage
tool.execution_*log lines were available.Branch / Workflow Concentration (metadata proxy for behavior)
copilot/dynamic-checkouts-github-action20/50 (40%, 20.0% success),copilot/task-...87cd569315/50 (30%, 73.3% success, merged as PR Report MCP payload sizes in audit and aggregate analysis #63687),copilot/task-...ca3f08859/50 (18%, 55.6% success, merged as PR Track skill usage in audit artifacts #63688),copilot/preserve-review-provenance-marker6/50 (12%, 66.7% success).Context Issues
Experimental Analysis
Standard run only (roll=95/100, threshold <30 for experimental) — no experimental strategy this run.
📈 Session Trends Analysis
Completion Patterns
Today's 48.0% completion rate is the 2nd-highest reading of the last two weeks (after 09-24's 44.0%) and sits well above the ~31% 30-day mean, continuing the sawtooth pattern this workflow has shown for over a month rather than any sustained trend in either direction. The failed/action-required series remains persistently high in absolute count even on good days, reflecting the branch-level stuck-gate pattern (Squad/Agentic Commands/Doc Build-Deploy) rather than genuine task failures.
Duration & Efficiency
Today's average duration (4.46 min) sits in the low range with no cascade- or status-resync-driven spike, unlike the 75.8-minute outlier around 09-20 and the 40-minute one around 09-16. The bar overlay for "sessions with loops" specified in the standard chart template was omitted here rather than fabricated, since loop detection requires conversation-log data that has been unavailable for 35 consecutive days.
Actionable Recommendations
For Users Writing Task Descriptions
No prompt text is visible in this data source, so prompt-level recommendations cannot be generated this run. General guidance from prior runs' findings still applies: reference specific files/paths and state expected outcomes explicitly.
For System Improvements
copilot-session-data-fetchmodule: Highest-impact fix. This is now a 35-consecutive-day-old gap that prevents the workflow from doing the turn-by-turn behavioral analysis it was built for (tool usage, token efficiency, loop detection, error recovery). Every other finding in this report is a metadata-only proxy.Squad/Agentic Commands/Doc Build - Deploygate workflows: These have shown ~0% success for weeks straight while true-agentic and code-scanning workflows sit near 100% — worth confirming whether this is expected gating behavior (blocked pending review) or a genuine regression.push_event_reattributionexperiment found push-clustered completion (66.7%) more than double the raw session rate (30.0%) — promoting this as a standing secondary metric would give a clearer picture than the noisy raw rate.For Tool Development
missing_data) whenlogs/is empty would save analysis time — currently this is discovered fresh each day via manualfind.Historical Trends and Statistical Summary
Trends Over Time
Statistical Summary
Next Steps
Analysis generated automatically on 2026-09-27
Run ID: 36301452080
Workflow: Copilot Session Insights
All reactions