[experiments] Daily Experiment Report — 2026-09-17 #61560
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #61760. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-17
47 experiment workflow branches are active in
github/gh-aw, comprising 51 individual A/B experiments across 48 workflows. None have decision-quality evidence sufficient for aPROMOTE/REJECTverdict yet — core decisions are 24 READY (evidence gates hit but metric pipeline incomplete) and 27 COLLECTING (still accumulating samples). No promotions or rejections today; every deterministic decision remainsEXTENDorINCONCLUSIVE.⚡ Quick Stats
prompt_style·ci-coachH0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
📈 View Detailed Statistics
Sample Sizes & Progress
detailedconciseCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationtone_style·typistH0: no change in engagement. H1: conversational tone increases discussion views+reactions+comments by 20%+ while maintaining analysis quality
📈 View Detailed Statistics
Sample Sizes & Progress
conversationalformalCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparison📊 Summary
View Full Experiments Table
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
guardrail_unsupportedverdicts (ci-coach,daily-fact,daily-security-red-team):run_success_rate/empty_output_rateguardrail metrics have no supported observation source — wire these into the metrics pipeline so READY experiments can reach a real PROMOTE/REJECT verdict instead of perpetually re-extending.unsupported_multi_variant(14 experiments —aw-failure-investigator,copilot-agent-analysis,daily-caveman-optimizer,daily-code-metrics,daily-compiler-quality,daily-doc-healer,daily-doc-updater,daily-issues-report,daily-semgrep-scan,deep-report,dependabot-go-checker,plan,smoke-copilot-sub-agents, and the staledaily-sub-agent-optimizerbranch): 3+ variant designs currently cannot produce a core decision at all — prioritize adding a supported ≥3-arm test (e.g. Kruskal-Wallis + pairwise Bonferroni) to core analysis..mdsource (smoke-antigravity,dependabot-campaign,daily-sub-agent-optimizer,smoke-pi,daily-function-namer) — these accumulate runs but can never reach a decision since the workflow definition was removed/renamed.Top 3 eval actions
output_format_adherenceeval (used bycopilot-agent-analysis,daily-code-metrics,daily-compiler-quality,daily-issues-report,daily-semgrep-scan) should be tightened to also emit a machine-readable per-run score usable as a supported multi-variant metric input, not just a qualitative grade.execution-durationgrader-backed experiments (daily-rendering-scripts-verifier,pr-sous-chef) are technically COLLECTING but showtotal_runs: 1— verify the grader is actually firing on scheduled runs, not just manual triggers.discussion_engagement_score(used by 4 workflows) to a normalized 0-1 eval score so cross-experiment comparison of engagement-driven experiments becomes possible without re-deriving reaction/comment counts per workflow.Decision-pipeline gaps for next PR
guardrail_unsupportedandinsufficient_observationsaccount for 37/51 decisions — the core analyzer has no path from "guardrail declared in frontmatter" to "guardrail backed by a real per-run observation." This is the single largest blocker to any PROMOTE/REJECT verdict in the repo today.smoke-copilot*(caveman × subagent_model) — cells are extremely sparse (n=1–4) at therecent_runswindow of 10; this needs to become a normalized core signal fed by the full 30-run window rather than left as a reporting-layer helper.All reactions