[experiments] Daily Experiment Report — 2026-09-16 #61311
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #61560. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-16
51 A/B experiments analysed across 48 workflows in
github/gh-aw. All decisions areEXTEND(37) orINCONCLUSIVE(14) — no experiment currently has enough decision-quality evidence toPROMOTEorREJECT. 24 experiments have reachedmin_samples(READY), but native metric observations are missing for most of them.⚡ Quick Stats
issue:field)🟢 Ready for Analysis (24 experiments)
All variants reached
min_samples, but 24/24 still resolve toEXTEND/INCONCLUSIVE— 20 lack native metric observations (insufficient_observationsorguardrail_unsupported), 4 are multi-variant (unsupported_multi_variant, decision engine currently supports exactly 2 variants).prompt_compression·agent-performance-analyzer.md📈 Statistics — hypothesis, samples, balance test
H0: no change in aic. H1: caveman reduces tokens by ≥20% while maintaining quality ≥90%
cavemanverboseCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·agent-persona-explorer.md📈 Statistics — hypothesis, samples, balance test
H0: no change in aic or duration. H1: batch reduces tokens by ≥20% and duration by ≥15% without quality loss
batchper_scenarioCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_variant·aw-failure-investigator.md📈 Statistics — hypothesis, samples, balance test
H0: no change in output_length_chars across tone variants. H1: assertive tone produces shorter, more actionable outputs than clinical or narrative, with equivalent or better sub-issue quality.
assertiveclinicalnarrativeCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstone_variant·breaking-change-checker.md📈 Statistics — hypothesis, samples, balance test
H0: no change in issue_engagement_rate. H1: urgent increases issue_engagement_rate by >=7 percentage points without degrading run_success_rate.
neutralurgentCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·ci-coach.md📈 Statistics — hypothesis, samples, balance test
H0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
concisedetailedCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationsub_agent_strategy·daily-agentrx-trace-optimizer.md📈 Statistics — hypothesis, samples, balance test
H0: no change in issue quality or run success rate. H1: sub_agents variant yields higher evidence completeness score with equal or lower token cost
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·daily-astrostylelite-markdown-spellcheck.md📈 Statistics — hypothesis, samples, balance test
Concise prompt reduces token consumption ≥20% without degrading fix precision. H0: no difference in fix rate.
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·daily-community-attribution.md📈 Statistics — hypothesis, samples, balance test
(not specified)
conciseverboseCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·daily-fact.md📈 Statistics — hypothesis, samples, balance test
H0: no change in discussion engagement rate. H1: multi_candidate produces more novel verses with higher reaction counts (expected +20% reactions).
multi_candidatesingle_passCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "empty_output_rate" is not backed by a supported metric observationprompt_style·daily-news.md📈 Statistics — hypothesis, samples, balance test
H0: no change in output quality. H1: concise prompt reduces token usage by ≥20% with no significant drop in output completeness score
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·daily-security-red-team.md📈 Statistics — hypothesis, samples, balance test
H0: no change in finding quality. H1: iterative reduces false-positive rate by >=20% at <=30% token overhead
iterativesingle_passCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationoutput_format·deep-report.md📈 Statistics — hypothesis, samples, balance test
H0: no change in discussion engagement or token cost. H1: executive_brief reduces token usage by ≥20% without reducing engagement; annotated_brief improves actionability; ste improves clarity while reducing token usage.
annotated_briefexecutive_brieffull_briefingsteCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstool_verbosity·gpclean.md📈 Statistics — hypothesis, samples, balance test
H0: no change in token consumption. H1: minimal toolset reduces tokens by 10-15% while maintaining issue quality (detection accuracy + alternative research depth)
full_bashminimal_toolsetCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·issue-arborist.md📈 Statistics — hypothesis, samples, balance test
H0: no change in links_created. H1: detailed instructions produce ≥15% more correct links per run
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·smokeantigravity📈 Statistics — hypothesis, samples, balance test
(not specified)
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smoke-copilot.md📈 Statistics — hypothesis, samples, balance test
(not specified)
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smoke-copilot.md📈 Statistics — hypothesis, samples, balance test
(not specified)
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smoke-copilot-aoai-apikey.md📈 Statistics — hypothesis, samples, balance test
(not specified)
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smoke-copilot-aoai-apikey.md📈 Statistics — hypothesis, samples, balance test
(not specified)
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smoke-copilot-aoai-entra.md📈 Statistics — hypothesis, samples, balance test
(not specified)
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smoke-copilot-aoai-entra.md📈 Statistics — hypothesis, samples, balance test
(not specified)
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·smoke-gemini.md📈 Statistics — hypothesis, samples, balance test
H0: no change in aic. H1: sub_agents reduces tokens by >=20%
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_decomposition·smokepi📈 Statistics — hypothesis, samples, balance test
(not specified)
parallel_sub_agentssingle_agentCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_style·typist.md📈 Statistics — hypothesis, samples, balance test
H0: no change in engagement. H1: conversational tone increases discussion views+reactions+comments by 20%+ while maintaining analysis quality
conversationalformalCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparison🟡 Collecting Samples (27 experiments)
Below the
min_samplesthreshold for at least one variant. Progress bars show current/target counts.sub_agent_strategysingle_agent=17/20,sub_agents=21/20insufficient_samples)audit_decompositionphased_sub_agents=41/264,single_agent=32/264insufficient_samples)prompt_styleconcise=3/20,detailed=6/20insufficient_samples)output_formatprose=34/30,ste=3/30,structured=43/30unsupported_multi_variant)detail_levelbrief=11/20,comprehensive=6/20insufficient_samples)model_sizegpt-5.3-codex=5/20,gpt-5.3-codex-spark=5/20insufficient_samples)model_sizeagent=1/20,claude-haiku-4.5=61/20,claude-sonnet-4.6=34/20,claude-sonnet-5=6/20,small-agent=2/20unsupported_multi_variant)output_formatexecutive_summary=50/20,full_detail=57/20,ste=13/20unsupported_multi_variant)output_formatconcise=51/20,detailed=59/20,ste=12/20unsupported_multi_variant)model_sizeagent=2/20,claude-haiku-4.5=53/20,claude-sonnet-4.6=43/20,claude-sonnet-5=5/20,small-agent=1/20unsupported_multi_variant)model_sizeagent=1/20,claude-haiku-4.5=44/20,claude-sonnet-4.6=47/20,claude-sonnet-5=7/20,small-agent=2/20unsupported_multi_variant)model_sizeinsufficient_observations)output_formatcollapsible=56/20,inline=60/20,ste=17/20unsupported_multi_variant)remove_redundant_context_v1insufficient_observations)log_fetch_strategyinsufficient_observations)semgrep_output_formatbullet_list=28/30,prose=27/30,structured_sections=22/30unsupported_multi_variant)timeout_settingdefault=2/20,relaxed=1/20,tight=1/20unsupported_multi_variant)caveman_modeno=5/20,yes=9/20insufficient_samples)summary_detailbrief=6/20,detailed=6/20insufficient_samples)prompt_styleconcise=17/20,detailed=17/20,step_by_step=13/20unsupported_multi_variant)reasoning_depthbaseline=4/20,deep=2/20,shallow=1/20unsupported_multi_variant)remove_redundant_context_v1insufficient_observations)sub_agent_strategydelegated_sequential=15/30,inline_strict=12/30,single_agent_control=16/30unsupported_multi_variant)prompt_style_testinsufficient_observations)sub_agent_strategyinsufficient_observations)model_sizeinsufficient_observations)prefetch_strategyeager=5/20,lazy=9/20insufficient_samples)📊 Summary
View Full Experiments Table (all 51)
prompt_compressioncavemanverbosesub_agent_strategybatchper_scenariotone_variantassertiveclinical / narrativetone_variantneutralurgentprompt_styleconcisedetailedsub_agent_strategysingle_agentsub_agentsprompt_styleconcisedetailedprompt_styleconciseverbosereasoning_depthmulti_candidatesingle_passprompt_styleconcisedetailedreasoning_depthiterativesingle_passoutput_formatannotated_briefexecutive_brief / full_briefing / stetool_verbosityfull_bashminimal_toolsetprompt_styleconcisedetailedsub_agent_strategysingle_agentsub_agentscavemannoyessubagent_modellargesmallcavemannoyessubagent_modellargesmallcavemannoyessubagent_modellargesmallsub_agent_strategysingle_agentsub_agentssub_agent_decompositionparallel_sub_agentssingle_agenttone_styleconversationalformalsub_agent_strategysingle_agentsub_agentsaudit_decompositionphased_sub_agentssingle_agentprompt_styleconcisedetailedoutput_formatproseste / structureddetail_levelbriefcomprehensivemodel_sizegpt-5.3-codexgpt-5.3-codex-sparkmodel_sizeagentclaude-haiku-4.5 / claude-sonnet-4.6 / claude-sonnet-5 / small-agentoutput_formatexecutive_summaryfull_detail / steoutput_formatconcisedetailed / stemodel_sizeagentclaude-haiku-4.5 / claude-sonnet-4.6 / claude-sonnet-5 / small-agentmodel_sizeagentclaude-haiku-4.5 / claude-sonnet-4.6 / claude-sonnet-5 / small-agentmodel_sizen/an/aoutput_formatcollapsibleinline / steremove_redundant_context_v1n/an/alog_fetch_strategyn/an/asemgrep_output_formatbullet_listprose / structured_sectionstimeout_settingdefaultrelaxed / tightcaveman_modenoyessummary_detailbriefdetailedprompt_styleconcisedetailed / step_by_stepreasoning_depthbaselinedeep / shallowremove_redundant_context_v1n/an/asub_agent_strategydelegated_sequentialinline_strict / single_agent_controlprompt_style_testn/an/asub_agent_strategyn/an/amodel_sizen/an/aprefetch_strategyeagerlazy🔧 Self-Tuning Continuation Plan
View continuation plan
Top 3 experiment actions
READYexperiments blocked byinsufficient_observations(20 experiments, e.g.agent-performance-analyzer.prompt_compression,daily-news.prompt_style,issue-arborist.prompt_style) — these already have balanced samples ≥min_samplesbut the declaredmetric:field (e.g.aic,token_count) is never written into the git branch state or an eval/grader artifact the core analyzer can read. Add apick_experiment/grader hook that records the primary metric per run so these can resolve toPROMOTE/REJECTinstead of parking indefinitely atEXTEND.cicoach.prompt_style,daily-fact.reasoning_depth,daily-security-red-team.reasoning_depth— all three areREADYbut blocked byguardrail_unsupported; their configuredguardrail_metricsreference fields the current pipeline can't evaluate per-run. Wire these guardrails into the grader/eval observation source before the next report.deep-report.output_formatand 3 otherunsupported_multi_variantexperiments —deep-reporthas a statistically significant imbalance (chi_square=21.4, p=0.0001) worth investigating, but the core decision layer only supports exactly-two-variant comparisons today. Prioritize adding a multi-variant ANOVA/Kruskal-Wallis path so 4-variant experiments like this stop parking atINCONCLUSIVE.Top 3 eval actions
aic/token-cost metrics used byagent-performance-analyzer,agent-persona-explorer, andsmoke-copilot*experiments — these prompts currently produce no scoreable observations feeding the primary metric, so eval graders should assert token/aic capture explicitly per run.empty_output_rateandissue_creation_success_rateguardrails referenced bydeep-report.output_format, so guardrail pass/fail becomes computable instead of defaulting to "not configured".daily-fact.reasoning_depthanddaily-security-red-team.reasoning_depth, add eval grading for the declared guardrail metrics soguardrail_unsupportedconverts to a real pass/fail signal.Decision-pipeline gaps for the next PR
READYexperiments are stuck oninsufficient_observations/guardrail_unsupported— the single largest blocker to converging on anyPROMOTE/REJECTdecision in this repository today.unsupported_multi_variant, 14 experiments) is the second-largest gap;deep-report.output_formatshows a significant balance-test signal that the core analyzer cannot yet act on.issue:tracking field, so Step 8/9 (issue notifications and lifecycle labels) had nothing to act on this run — consider recommending tracking issues be added for experiments nearingREADYstatus.PROMOTEdecision with sparse cells; this remains a reporting-only safety net until a genuine multi-experimentPROMOTEcase arises.All reactions