[experiments] Daily Experiment Report — 2026-08-25 #55715
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #55983. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-08-25
47 experiments analysed across 44 workflows. 23 are ready for analysis. Core decisions: 0 PROMOTE, 32 EXTEND, 0 REJECT, 15 INCONCLUSIVE. No promotions or rejections yet — all active experiments are still accumulating evidence or lack a supported outcome metric to compute effect size.
⚡ Quick Stats
prompt_compression·agentperformanceanalyzer📈 View Sample Progress
cavemanverboseCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·agentpersonaexplorer📈 View Sample Progress
batchper_scenarioCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·architectureguardian📈 View Sample Progress
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesaudit_decomposition·auditworkflowsH0: no change in run_success_rate. H1: phased_sub_agents improves run_success_rate by at least 15% relative while keeping runtime within +10%.
📈 View Sample Progress
phased_sub_agentssingle_agentCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplestone_variant·awfailureinvestigator📈 View Sample Progress
assertiveclinicalnarrativeCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsprompt_style·blogauditor📈 View Sample Progress
concisedetailedCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplestone_variant·breakingchangechecker📈 View Sample Progress
neutralurgentCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesprompt_style·cicoachH0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
📈 View Sample Progress
concisedetailedCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationoutput_format·copilotagentanalysisH0: no change in ai_credits_used. H1: prose or ste format reduces ai_credits_used by >=15% while keeping empty_discussion_rate <=5%
📈 View Sample Progress
prosestestructuredCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantssub_agent_strategy·dailyagentrxtraceoptimizer📈 View Sample Progress
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisondetail_level·dailyarchitecturediagram📈 View Sample Progress
briefcomprehensiveCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesprompt_style·dailyastrostylelitemarkdownspellcheck📈 View Sample Progress
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonmodel_size·dailycachestrategyanalyzerH0: no change in issue creation rate or run success rate. H1: gpt-5.4-mini reduces AI Credits while keeping run success rate >=0.90.
📈 View Sample Progress
agentgpt-5-codexgpt-5-minigpt-5.4gpt-5.4-minismall-agentCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsmodel_size·dailycavemanoptimizer📈 View Sample Progress
agentclaude-haiku-4.5claude-sonnet-4.6small-agentCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsoutput_format·dailycodemetrics📈 View Sample Progress
executive_summaryfull_detailsteCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsprompt_style·dailycommunityattribution📈 View Sample Progress
conciseverboseCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonoutput_format·dailycompilerquality📈 View Sample Progress
concisedetailedsteCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsmodel_size·dailydochealer📈 View Sample Progress
agentclaude-haiku-4.5claude-sonnet-4.6small-agentCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsmodel_size·dailydocupdater📈 View Sample Progress
agentclaude-haiku-4.5claude-sonnet-4.6small-agentCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsreasoning_depth·dailyfactH0: no change in discussion engagement rate. H1: multi_candidate produces more novel verses with higher reaction counts (expected +20% reactions).
📈 View Sample Progress
multi_candidatesingle_passCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesmodel_size·dailyfunctionnamer📈 View Sample Progress
Core decision: 🟡 EXTEND (
insufficient_observations) — at least two variants are required for a decisionoutput_format·dailyissuesreport📈 View Sample Progress
collapsibleinlinesteCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsprompt_style·dailynews📈 View Sample Progress
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonlog_fetch_strategy·dailysafeoutputoptimizerH0: no change in run_duration_ms. H1: eager reduces run duration by >=15% by avoiding MCP log-fetch turns
📈 View Sample Progress
eagerlazyCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "issue_creation_success_rate" is not backed by a supported metric observationreasoning_depth·dailysecurityredteamH0: no change in finding quality. H1: iterative reduces false-positive rate by >=20% at <=30% token overhead
📈 View Sample Progress
iterativesingle_passCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationsemgrep_output_format·dailysemgrepscanH0: no change in alert creation rate across formats. H1: structured_sections produces ≥15% more alerts successfully created vs. baseline bullet_list; ste improves completeness via clearer, simpler language.
📈 View Sample Progress
bullet_listprosestructured_sectionsCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstimeout_setting·dailysubagentoptimizer📈 View Sample Progress
defaultrelaxedtightCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantscaveman_mode·dataflowprdiscussiondatasetH0: no change in input_token_count, retention_rate, or run_success_rate. H1: caveman prompt cuts input tokens ≥30% with retention_rate ≥0.80 and run_success_rate ≥0.90
📈 View Sample Progress
noyesCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesoutput_format·deepreport📈 View Sample Progress
annotated_briefexecutive_brieffull_briefingsteCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantssummary_detail·dependabotcampaign📈 View Sample Progress
briefdetailedCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samplesprompt_style·dependabotgochecker📈 View Sample Progress
concisedetailedstep_by_stepCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstool_verbosity·gpclean📈 View Sample Progress
full_bashminimal_toolsetCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·issuearborist📈 View Sample Progress
concisedetailedCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·plan📈 View Sample Progress
baselinedeepshallowCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantssub_agent_strategy·smokeantigravity📈 View Sample Progress
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilot📈 View Sample Progress
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilot📈 View Sample Progress
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilotaoaiapikey📈 View Sample Progress
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilotaoaiapikey📈 View Sample Progress
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilotaoaientra📈 View Sample Progress
noyesCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilotaoaientra📈 View Sample Progress
largesmallCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·smokecopilotsubagentsH0: no change in pass_rate between inline_strict and delegated_sequential. H1: delegated_sequential improves pass_rate by >= 0.15 absolute versus inline_strict. Note: single_agent_control is a synthetic negative baseline (always FAIL by design) excluded from H1 comparisons.
📈 View Sample Progress
delegated_sequentialinline_strictsingle_agent_controlCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantssub_agent_strategy·smokegemini📈 View Sample Progress
single_agentsub_agentsCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_decomposition·smokepi📈 View Sample Progress
parallel_sub_agentssingle_agentCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonmodel_size·testqualitysentinel📈 View Sample Progress
claude-haiku-4.5claude-sonnet-4.6Core decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_style·typist📈 View Sample Progress
conversationalformalCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprefetch_strategy·weeklyblogpostwriter📈 View Sample Progress
eagerlazyCore decision: 🟡 EXTEND (
insufficient_samples) — one or more variants have fewer usable observations than min_samples📊 Summary
View Full Experiments Table
prompt_compressionagentperformanceanalyzercaveman,verboseinsufficient_observationssub_agent_strategyagentpersonaexplorerbatch,per_scenarioinsufficient_observationssub_agent_strategyarchitectureguardiansingle_agent,sub_agentsinsufficient_samplesaudit_decompositionauditworkflowsphased_sub_agents,single_agentinsufficient_samplestone_variantawfailureinvestigatorassertive,clinical,narrativeunsupported_multi_variantprompt_styleblogauditorconcise,detailedinsufficient_samplestone_variantbreakingchangecheckerneutral,urgentinsufficient_samplesprompt_stylecicoachconcise,detailedguardrail_unsupportedoutput_formatcopilotagentanalysisprose,ste,structuredunsupported_multi_variantsub_agent_strategydailyagentrxtraceoptimizersingle_agent,sub_agentsinsufficient_observationsdetail_leveldailyarchitecturediagrambrief,comprehensiveinsufficient_samplesprompt_styledailyastrostylelitemarkdownspellcheckconcise,detailedinsufficient_observationsmodel_sizedailycachestrategyanalyzeragent,gpt-5-codex,gpt-5-mini,gpt-5.4,gpt-5.4-mini,small-agentunsupported_multi_variantmodel_sizedailycavemanoptimizeragent,claude-haiku-4.5,claude-sonnet-4.6,small-agentunsupported_multi_variantoutput_formatdailycodemetricsexecutive_summary,full_detail,steunsupported_multi_variantprompt_styledailycommunityattributionconcise,verboseinsufficient_observationsoutput_formatdailycompilerqualityconcise,detailed,steunsupported_multi_variantmodel_sizedailydochealeragent,claude-haiku-4.5,claude-sonnet-4.6,small-agentunsupported_multi_variantmodel_sizedailydocupdateragent,claude-haiku-4.5,claude-sonnet-4.6,small-agentunsupported_multi_variantreasoning_depthdailyfactmulti_candidate,single_passinsufficient_samplesmodel_sizedailyfunctionnamerinsufficient_observationsoutput_formatdailyissuesreportcollapsible,inline,steunsupported_multi_variantprompt_styledailynewsconcise,detailedinsufficient_observationslog_fetch_strategydailysafeoutputoptimizereager,lazyguardrail_unsupportedreasoning_depthdailysecurityredteamiterative,single_passguardrail_unsupportedsemgrep_output_formatdailysemgrepscanbullet_list,prose,structured_sectionsunsupported_multi_varianttimeout_settingdailysubagentoptimizerdefault,relaxed,tightunsupported_multi_variantcaveman_modedataflowprdiscussiondatasetno,yesinsufficient_samplesoutput_formatdeepreportannotated_brief,executive_brief,full_briefing,steunsupported_multi_variantsummary_detaildependabotcampaignbrief,detailedinsufficient_samplesprompt_styledependabotgocheckerconcise,detailed,step_by_stepunsupported_multi_varianttool_verbositygpcleanfull_bash,minimal_toolsetinsufficient_observationsprompt_styleissuearboristconcise,detailedinsufficient_observationsreasoning_depthplanbaseline,deep,shallowunsupported_multi_variantsub_agent_strategysmokeantigravitysingle_agent,sub_agentsinsufficient_observationscavemansmokecopilotno,yesinsufficient_observationssubagent_modelsmokecopilotlarge,smallinsufficient_observationscavemansmokecopilotaoaiapikeyno,yesinsufficient_observationssubagent_modelsmokecopilotaoaiapikeylarge,smallinsufficient_observationscavemansmokecopilotaoaientrano,yesinsufficient_observationssubagent_modelsmokecopilotaoaientralarge,smallinsufficient_observationssub_agent_strategysmokecopilotsubagentsdelegated_sequential,inline_strict,single_agent_controlunsupported_multi_variantsub_agent_strategysmokegeminisingle_agent,sub_agentsinsufficient_observationssub_agent_decompositionsmokepiparallel_sub_agents,single_agentinsufficient_observationsmodel_sizetestqualitysentinelclaude-haiku-4.5,claude-sonnet-4.6insufficient_observationstone_styletypistconversational,formalinsufficient_observationsprefetch_strategyweeklyblogpostwritereager,lazyinsufficient_samples🔧 Self-Tuning Continuation Plan
Top experiment actions:
run_success_rate/ other guardrail metrics into eval graders forcicoach,daily-safe-output-optimizer, anddaily-security-red-team— currently blocked onguardrail_unsupported, so no PROMOTE/REJECT decision is possible even at READY sample sizes.token_count_per_run,output_length_chars, and duration proxies soinsufficient_observationsexperiments (e.g.agent-performance-analyzer,testqualitysentinel,smoke-*) can progress past READY once metric data lands.unsupported_multi_variantblocks 15 experiments:awfailureinvestigator,dailycachestrategyanalyzer,dailycavemanoptimizer,dailycodemetrics,dailycompilerquality,dailydochealer,dailydocupdater,dailyissuesreport,dailysemgrepscan,dailysubagentoptimizer,deepreport,dependabotgochecker,plan,smokecopilotsubagents,copilotagentanalysis).Top eval actions:
run_success_rateandproposal_qualitymetrics referenced incicoach/daily-security-red-teamhypotheses so guardrails stop reportingguardrail_unsupported.token_count_per_run,output_length_chars) foroutput_format-style experiments (daily-code-metrics,daily-compiler-quality,deep-report,daily-issues-report) to unblock multi-variant comparisons once K≥3 support lands.min_samplesvisibility for low-volume experiments stillCOLLECTING(blog-auditor,daily-architecture-diagram,dataflow-pr-discussion-dataset,weekly-blog-post-writer,dependabot-campaign) — several are under 10 runs/variant.Decision-pipeline gaps for next PR:
guardrail_unsupportedappears forcicoach,daily-safe-output-optimizer, anddaily-security-red-team— these mandatory guardrails have no backing metric observation; core analysis correctly holds at EXTEND rather than guessing.unsupported_multi_variant(15 experiments) is the single largest blocker to converting COLLECTING/READY signals into deterministic PROMOTE/REJECT decisions; this should be the top priority for the core decision engine.All reactions