[experiments] Daily Experiment Report — 2026-09-14 #60792
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #61070. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-14
51 active A/B experiments were analysed across 48 workflows in
github/gh-aw. 23 have reachedREADYsample-size thresholds; the core deterministic decision layer returned 0PROMOTE, 37EXTEND, 0REJECT, and 14INCONCLUSIVE— the platform's automatic decision logic currently only supports strict two-variant comparisons with an observed primary-metric signal, so mostREADYmulti-variant or metric-unsupported experiments correctly stayEXTEND/INCONCLUSIVErather than false-promoting.⚡ Quick Stats
smokecopilotfamily:caveman×subagent_model)Data-pipeline gap (affects every experiment): none of the 51 experiments have a working primary-metric observation feed into the core decision engine yet — reason codes are exclusively
insufficient_observations(25),unsupported_multi_variant(14),insufficient_samples(9), andguardrail_unsupported(3). Assignment/balance tracking (chi-square,is_balanced) works correctly for all experiments; the missing link is wiring each workflow's declaredmetric:/guardrail_metrics:into the artifact the analyzer reads. This is the single highest-leverage engineering fix (see Step 10).prompt_compression·agent-performance-analyzer.mdH0: no change in effective_tokens. H1: caveman reduces tokens by ≥20% while maintaining quality ≥90%
📈 View Detailed Statistics
Sample Sizes & Progress
cavemanverboseCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_variant·aw-failure-investigator.mdH0: no change in output_length_chars across tone variants. H1: assertive tone produces shorter, more actionable outputs than clinical or narrative, with equivalent or better sub-issue quality.
📈 View Detailed Statistics
Sample Sizes & Progress
assertiveclinicalnarrativeCore decision: INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsprompt_style·ci-coach.mdH0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationreasoning_depth·daily-fact.mdH0: no change in discussion engagement rate. H1: multi_candidate produces more novel verses with higher reaction counts (expected +20% reactions).
📈 View Detailed Statistics
Sample Sizes & Progress
multi_candidatesingle_passCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "empty_output_rate" is not backed by a supported metric observationprompt_style·daily-news.mdH0: no change in output quality. H1: concise prompt reduces token usage by ≥20% with no significant drop in output completeness score
📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·daily-security-red-team.mdH0: no change in finding quality. H1: iterative reduces false-positive rate by >=20% at <=30% token overhead
📈 View Detailed Statistics
Sample Sizes & Progress
iterativesingle_passCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationprompt_style·issue-arborist.mdH0: no change in links_created. H1: detailed instructions produce ≥15% more correct links per run
📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_style·typist.mdH0: no change in engagement. H1: conversational tone increases discussion views+reactions+comments by 20%+ while maintaining analysis quality
📈 View Detailed Statistics
Sample Sizes & Progress
conversationalformalCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonFactorial interaction:
caveman×subagent_model(smoke-copilotfamily)smokecopilot,smokecopilotaoaiapikey, andsmokecopilotaoaientraeach run two simultaneous experiments (cavemanandsubagent_model). Interaction cells fromsmoke-copilot's last 30 runs (assignment-matched):interaction_p_value= 0.519 · all 4 cells haven < min_samples (20)→interaction_risk_status: SPARSE_CELL_RISK.Both individual experiments (
caveman,subagent_model) are currentlyEXTEND(notPROMOTE), so this safety hold does not change any report action today — noted here so it's tracked ahead of the day the 2×2 design reachesREADYon both axes.📊 Summary
View Full Experiments Table (51 experiments across 48 workflows)
prompt_compressioncavemanverbosesub_agent_strategybatchper_scenariosub_agent_strategysingle_agentsub_agentsaudit_decompositionsingle_agentphased_sub_agentstone_variantassertiveclinicalprompt_styleconcisedetailedtone_variantneutralurgentprompt_styledetailedconciseoutput_formatprosestesub_agent_strategysingle_agentsub_agentsdetail_levelbriefcomprehensiveprompt_styleconcisedetailedmodel_sizegpt-5.3-codexgpt-5.3-codex-sparkmodel_sizeagentclaude-haiku-4.5output_formatexecutive_summaryfull_detailprompt_styleconciseverboseoutput_formatconcisedetailedmodel_sizeagentclaude-haiku-4.5model_sizeagentclaude-haiku-4.5reasoning_depthsingle_passmulti_candidatemodel_sizen/an/aoutput_formatcollapsibleinlineprompt_styleconcisedetailedremove_redundant_context_v1n/an/alog_fetch_strategyeagerlazyreasoning_depthsingle_passiterativesemgrep_output_formatbullet_listprosetimeout_settingdefaultrelaxedcaveman_modenoyesoutput_formatannotated_briefexecutive_briefsummary_detailbriefdetailedprompt_styleconcisedetailedtool_verbosityfull_bashminimal_toolsetprompt_styleconcisedetailedreasoning_depthbaselinedeepremove_redundant_context_v1n/an/asub_agent_strategysingle_agentsub_agentscavemannoyessubagent_modellargesmallcavemannoyessubagent_modellargesmallcavemannoyessubagent_modellargesmallsub_agent_strategydelegated_sequentialinline_strictsub_agent_strategysingle_agentsub_agentssub_agent_decompositionparallel_sub_agentssingle_agentprompt_style_testn/an/asub_agent_strategysingle_agentsub_agentsmodel_sizen/an/atone_styleconversationalformalprefetch_strategyeagerlazy🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
insufficient_observationseven though sample counts are healthy (e.g.awfailureinvestigatorat 415 runs,smokecopilotcaveman at 359 runs). Each workflow declares ametric:field, but the analyzer has no artifact path reading it back. This is the top blocker to ever reachingPROMOTE/REJECT.unsupported_multi_variant(e.g.awfailureinvestigatortone_variant with 3 variants,dailycodemetrics,deepreport). These have large, balanced samples (is_balanced: true, p-values > 0.8) and are otherwise fully ready.guardrail_metricsevaluation for the 3guardrail_unsupportedexperiments (cicoach,dailyfact,dailysecurityredteam) — each declares a guardrail (e.g.run_success_rate >= 0.85) that is never evaluated, soEXTENDis issued purely as a safety default even when samples are sufficient.Top 3 eval actions
daily-factanddaily-security-red-teamworkflow health before trusting any experiment metric. Both show 0% success rate on their last 10 completed runs — no eval or grader questions will produce meaningful signal while the underlying workflow is failing outright.cicoach'sprompt_styleexperiment —detailedshows 67% success vsconcise100% (n=6 vs n=4); before any decision, an eval question isolating whether failures correlate with prompt length vs. unrelated infra flakiness would sharpen the signal.issue-arboristprompt_styleeval questions around link-quality, sincedetailed(n=6, 33% success) underperformsconcise(n=4, 50% success) on raw completion rate — the declared metric (links_created) doesn't currently capture whether low-success runs are silently producing zero links.Decision-pipeline gaps for next PR
insufficient_observations(25 experiments) andguardrail_unsupported(3 experiments) both stem from the same root cause: no artifact writes the declaredmetric:/guardrail_metrics:values back to the experiment state thatgh aw experiments analyzereads. This should become a normalized "observation ingestion" step.unsupported_multi_variant(14 experiments) — the core decision engine needs either pairwise comparison support (control vs. each candidate) or an explicit N-way ANOVA/Kruskal-Wallis path before these can leaveINCONCLUSIVE.smokecopilotfamilycaveman×subagent_model) are currently computed only in this report, ad hoc. They should become a normalized core analysis signal soSPARSE_CELL_RISKcan automatically gatePROMOTEdecisions once any 2×2 design reaches readiness on both axes.All reactions