You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
51 active A/B experiments span 48 workflows in github/gh-aw. No experiment has reached a decision-quality evidence state (PROMOTE/REJECT) yet — every resolved analysis today is EXTEND or INCONCLUSIVE. 24 experiments are READY (samples collected) but all 24 are held back by data-pipeline gaps (missing per-run metric observations, guardrail metrics with no supported source, or >2-variant designs the core analyzer does not yet compare pairwise) rather than by an actual null result. 27 experiments are still COLLECTING samples.
⚡ Quick Stats
Metric
Value
Active experiments
51 (48 workflows)
Ready for analysis
24
Decisions with sufficient evidence (PROMOTE/REJECT)
All 24 have reached min_samples per variant, but the core analyzer returned a non-terminal decision for every one. Grouped by reason code:
`insufficient_observations` — primary metric has no per-run observation data to compare variants (19)
Workflow
Experiment
Variants (progress)
agentperformanceanalyzer
prompt_compression
caveman:██████████ 66/20 (100%)
agentpersonaexplorer
sub_agent_strategy
batch:██████████ 64/20 (100%)
breakingchangechecker
tone_variant
neutral:██████████ 24/20 (100%)
dailyagentrxtraceoptimizer
sub_agent_strategy
single_agent:██████████ 41/20 (100%)
dailyastrostylelitemarkdownspellcheck
prompt_style
concise:██████████ 71/20 (100%)
dailycommunityattribution
prompt_style
concise:██████████ 69/20 (100%)
dailynews
prompt_style
concise:██████████ 58/20 (100%)
gpclean
tool_verbosity
full_bash:██████████ 57/20 (100%)
issuearborist
prompt_style
concise:██████████ 43/20 (100%)
smokeantigravity
sub_agent_strategy
single_agent:██████████ 99/20 (100%)
smokecopilot
caveman
no:██████████ 184/20 (100%)
smokecopilot
subagent_model
large:██████████ 174/20 (100%)
smokecopilotaoaiapikey
caveman
no:██████████ 93/20 (100%)
smokecopilotaoaiapikey
subagent_model
large:██████████ 93/20 (100%)
smokecopilotaoaientra
caveman
no:██████████ 79/20 (100%)
smokecopilotaoaientra
subagent_model
large:██████████ 79/20 (100%)
smokegemini
sub_agent_strategy
single_agent:██████████ 127/20 (100%)
smokepi
sub_agent_decomposition
parallel_sub_agents:██████████ 42/20 (100%)
typist
tone_style
conversational:██████████ 45/20 (100%)
`guardrail_unsupported` — a mandatory guardrail metric has no supported native source (3)
Workflow
Experiment
Variants (progress)
cicoach
prompt_style
concise:██████████ 37/20 (100%)
dailyfact
reasoning_depth
multi_candidate:██████████ 38/30 (100%)
dailysecurityredteam
reasoning_depth
iterative:██████████ 79/30 (100%)
`unsupported_multi_variant` — experiment has >2 variants; core analyzer only supports pairwise control/candidate comparison (2)
Workflow
Experiment
Variants (progress)
awfailureinvestigator
tone_variant
assertive:██████████ 151/20 (100%)
deepreport
output_format
annotated_brief:██████████ 52/20 (100%)
🟡 Still Collecting (27 experiments)
View all COLLECTING experiments and sample progress
Workflow
Experiment
Decision
Reason code
Progress toward min_samples
architectureguardian
sub_agent_strategy
EXTEND
insufficient_samples
single_agent:████████░░ 17/20 (85%)
auditworkflows
audit_decomposition
EXTEND
insufficient_samples
phased_sub_agents:██░░░░░░░░ 43/264 (16%)
blogauditor
prompt_style
EXTEND
insufficient_samples
concise:██░░░░░░░░ 3/20 (15%)
copilotagentanalysis
output_format
INCONCLUSIVE
unsupported_multi_variant
prose:██████████ 34/30 (100%)
dailyarchitecturediagram
detail_level
EXTEND
insufficient_samples
brief:██████░░░░ 11/20 (55%)
dailycachestrategyanalyzer
model_size
EXTEND
insufficient_samples
gpt-5.3-codex:████░░░░░░ 7/20 (35%)
dailycavemanoptimizer
model_size
INCONCLUSIVE
unsupported_multi_variant
agent:░░░░░░░░░░ 1/20 (5%)
dailycodemetrics
output_format
INCONCLUSIVE
unsupported_multi_variant
executive_summary:██████████ 53/20 (100%)
dailycompilerquality
output_format
INCONCLUSIVE
unsupported_multi_variant
concise:██████████ 53/20 (100%)
dailydochealer
model_size
INCONCLUSIVE
unsupported_multi_variant
agent:█░░░░░░░░░ 2/20 (10%)
dailydocupdater
model_size
INCONCLUSIVE
unsupported_multi_variant
agent:░░░░░░░░░░ 1/20 (5%)
dailyfunctionnamer
model_size
EXTEND
insufficient_observations
dailyissuesreport
output_format
INCONCLUSIVE
unsupported_multi_variant
collapsible:██████████ 59/20 (100%)
dailyrenderingscriptsverifier
remove_redundant_context_v1
EXTEND
insufficient_observations
dailysafeoutputoptimizer
log_fetch_strategy
EXTEND
insufficient_observations
dailysemgrepscan
semgrep_output_format
INCONCLUSIVE
unsupported_multi_variant
bullet_list:█████████░ 28/30 (93%)
dailysubagentoptimizer
timeout_setting
INCONCLUSIVE
unsupported_multi_variant
default:█░░░░░░░░░ 2/20 (10%)
dataflowprdiscussiondataset
caveman_mode
EXTEND
insufficient_samples
no:███░░░░░░░ 6/20 (30%)
dependabotcampaign
summary_detail
EXTEND
insufficient_samples
brief:███░░░░░░░ 6/20 (30%)
dependabotgochecker
prompt_style
INCONCLUSIVE
unsupported_multi_variant
concise:█████████░ 18/20 (90%)
plan
reasoning_depth
INCONCLUSIVE
unsupported_multi_variant
baseline:██░░░░░░░░ 4/20 (20%)
prsouschef
remove_redundant_context_v1
EXTEND
insufficient_observations
smokecopilotsubagents
sub_agent_strategy
INCONCLUSIVE
unsupported_multi_variant
delegated_sequential:█████░░░░░ 15/30 (50%)
smokeproject
prompt_style_test
EXTEND
insufficient_observations
smoketemporaryid
sub_agent_strategy
EXTEND
insufficient_observations
testqualitysentinel
model_size
EXTEND
insufficient_samples
claude-haiku-4.5:████░░░░░░ 8/20 (40%)
weeklyblogpostwriter
prefetch_strategy
EXTEND
insufficient_samples
eager:██░░░░░░░░ 5/20 (25%)
🔗 Cross-Experiment Interaction Check
Three workflows (smokecopilot, smokecopilotaoaiapikey, smokecopilotaoaientra) run two simultaneous experiments each (caveman × subagent_model). Interaction cells from the last 10 recorded runs per workflow:
smokecopilot (caveman × subagent_model, n=10 of last 10 runs):
caveman \ subagent_model
large
small
yes
3
2
no
3
2
smokecopilotaoaiapikey (caveman × subagent_model, n=10 of last 10 runs):
caveman \ subagent_model
large
small
yes
4
1
no
1
4
smokecopilotaoaientra (caveman × subagent_model, n=10 of last 10 runs):
caveman \ subagent_model
large
small
yes
2
4
no
3
1
All cells have n < min_samples (20) — SPARSE_CELL_RISK. Since every underlying decision today is EXTEND (never PROMOTE), the reporting safety hold does not trigger for any of these three workflows; it will apply automatically before any future PROMOTE recommendation on these branches.
Fix unsupported_multi_variant (14 experiments) — output_format/model_size/reasoning_depth designs with 3+ variants (e.g. deepreport, dailycodemetrics, dailycavemanoptimizer, plan) cannot reach a decision until the core analyzer supports either a designated control + N pairwise candidate comparisons or an omnibus test with post-hoc contrasts.
Wire up guardrail_unsupported metrics (3 experiments) — dailysecurityredteam, dailyfact, and cicoach all fail on run_success_rate/empty_output_rate guardrails that have no supported native metric source; add these as first-class native metrics so mandatory guardrails can evaluate.
Close the insufficient_observations gap (25 experiments) — the majority of READY experiments (e.g. smokecopilot, typist, dailynews, agentperformanceanalyzer) have balanced sample counts but no per-run primary-metric observation recorded; prioritize wiring output_length_chars / aic / token_count-style metrics into the observation pipeline so these can convert to real decisions.
Top 3 eval actions
Tighten eval questions for workflows stuck at unsupported_multi_variant (e.g. daily-code-metrics.md, plan.md) to explicitly grade whether the agent recognized the >2-variant limitation instead of silently reporting a false decision.
Add an eval question to daily-security-red-team.md, daily-fact.md, and ci-coach.md prompts that checks whether guardrail metrics are actually emitted per run, surfacing the guardrail_unsupported gap earlier than the weekly experiment report.
For high-volume READY experiments with insufficient_observations (smoke-copilot.md family, typist.md, daily-news.md), add eval coverage asserting the primary metric: field frontmatter value is actually populated in run logs.
Decision-pipeline gaps for the next PR
guardrail_unsupported: run_success_rate and empty_output_rate guardrails declared in frontmatter have no backing native metric extractor yet (dailysecurityredteam, dailyfact, cicoach).
unsupported_multi_variant: 14 experiments with 3+ variants need a normalized core comparison strategy (pairwise vs. control, or omnibus + correction) instead of remaining permanently INCONCLUSIVE.
Interaction diagnostics (smokecopilot/smokecopilotaoaiapikey/smokecopilotaoaientra factorial cells) are currently a reporting-only overlay; they should become a normalized core analysis signal so SPARSE_CELL_RISK can gate PROMOTE decisions automatically rather than via this report's safety hold.
Descriptive window: last 30 runs per workflow (where available) · Decision thresholds: from each experiment's decision_policy · Run: 35701752501
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-22
51 active A/B experiments span 48 workflows in
github/gh-aw. No experiment has reached a decision-quality evidence state (PROMOTE/REJECT) yet — every resolved analysis today isEXTENDorINCONCLUSIVE. 24 experiments areREADY(samples collected) but all 24 are held back by data-pipeline gaps (missing per-run metric observations, guardrail metrics with no supported source, or >2-variant designs the core analyzer does not yet compare pairwise) rather than by an actual null result. 27 experiments are stillCOLLECTINGsamples.⚡ Quick Stats
🟢 Ready for Analysis (24 experiments)
All 24 have reached
min_samplesper variant, but the core analyzer returned a non-terminal decision for every one. Grouped by reason code:`insufficient_observations` — primary metric has no per-run observation data to compare variants (19)
agentperformanceanalyzerprompt_compressionagentpersonaexplorersub_agent_strategybreakingchangecheckertone_variantdailyagentrxtraceoptimizersub_agent_strategydailyastrostylelitemarkdownspellcheckprompt_styledailycommunityattributionprompt_styledailynewsprompt_stylegpcleantool_verbosityissuearboristprompt_stylesmokeantigravitysub_agent_strategysmokecopilotcavemansmokecopilotsubagent_modelsmokecopilotaoaiapikeycavemansmokecopilotaoaiapikeysubagent_modelsmokecopilotaoaientracavemansmokecopilotaoaientrasubagent_modelsmokegeminisub_agent_strategysmokepisub_agent_decompositiontypisttone_style`guardrail_unsupported` — a mandatory guardrail metric has no supported native source (3)
cicoachprompt_styledailyfactreasoning_depthdailysecurityredteamreasoning_depth`unsupported_multi_variant` — experiment has >2 variants; core analyzer only supports pairwise control/candidate comparison (2)
awfailureinvestigatortone_variantdeepreportoutput_format🟡 Still Collecting (27 experiments)
View all COLLECTING experiments and sample progress
min_samplesarchitectureguardiansub_agent_strategyinsufficient_samplesauditworkflowsaudit_decompositioninsufficient_samplesblogauditorprompt_styleinsufficient_samplescopilotagentanalysisoutput_formatunsupported_multi_variantdailyarchitecturediagramdetail_levelinsufficient_samplesdailycachestrategyanalyzermodel_sizeinsufficient_samplesdailycavemanoptimizermodel_sizeunsupported_multi_variantdailycodemetricsoutput_formatunsupported_multi_variantdailycompilerqualityoutput_formatunsupported_multi_variantdailydochealermodel_sizeunsupported_multi_variantdailydocupdatermodel_sizeunsupported_multi_variantdailyfunctionnamermodel_sizeinsufficient_observationsdailyissuesreportoutput_formatunsupported_multi_variantdailyrenderingscriptsverifierremove_redundant_context_v1insufficient_observationsdailysafeoutputoptimizerlog_fetch_strategyinsufficient_observationsdailysemgrepscansemgrep_output_formatunsupported_multi_variantdailysubagentoptimizertimeout_settingunsupported_multi_variantdataflowprdiscussiondatasetcaveman_modeinsufficient_samplesdependabotcampaignsummary_detailinsufficient_samplesdependabotgocheckerprompt_styleunsupported_multi_variantplanreasoning_depthunsupported_multi_variantprsouschefremove_redundant_context_v1insufficient_observationssmokecopilotsubagentssub_agent_strategyunsupported_multi_variantsmokeprojectprompt_style_testinsufficient_observationssmoketemporaryidsub_agent_strategyinsufficient_observationstestqualitysentinelmodel_sizeinsufficient_samplesweeklyblogpostwriterprefetch_strategyinsufficient_samples🔗 Cross-Experiment Interaction Check
Three workflows (
smokecopilot,smokecopilotaoaiapikey,smokecopilotaoaientra) run two simultaneous experiments each (caveman×subagent_model). Interaction cells from the last 10 recorded runs per workflow:smokecopilot(caveman × subagent_model, n=10 of last 10 runs):yesnosmokecopilotaoaiapikey(caveman × subagent_model, n=10 of last 10 runs):yesnosmokecopilotaoaientra(caveman × subagent_model, n=10 of last 10 runs):yesnoAll cells have
n < min_samples(20) — SPARSE_CELL_RISK. Since every underlying decision today isEXTEND(neverPROMOTE), the reporting safety hold does not trigger for any of these three workflows; it will apply automatically before any futurePROMOTErecommendation on these branches.📊 Summary
View Full Experiments Table
prompt_compressionagentperformanceanalyzercaveman,verboseinsufficient_observationssub_agent_strategyagentpersonaexplorerbatch,per_scenarioinsufficient_observationssub_agent_strategyarchitectureguardiansingle_agent,sub_agentsinsufficient_samplesaudit_decompositionauditworkflowsphased_sub_agents,single_agentinsufficient_samplestone_variantawfailureinvestigatorassertive,clinical,narrativeunsupported_multi_variantprompt_styleblogauditorconcise,detailedinsufficient_samplestone_variantbreakingchangecheckerneutral,urgentinsufficient_observationsprompt_stylecicoachconcise,detailedguardrail_unsupportedoutput_formatcopilotagentanalysisprose,ste,structuredunsupported_multi_variantsub_agent_strategydailyagentrxtraceoptimizersingle_agent,sub_agentsinsufficient_observationsdetail_leveldailyarchitecturediagrambrief,comprehensiveinsufficient_samplesprompt_styledailyastrostylelitemarkdownspellcheckconcise,detailedinsufficient_observationsmodel_sizedailycachestrategyanalyzergpt-5.3-codex,gpt-5.3-codex-sparkinsufficient_samplesmodel_sizedailycavemanoptimizeragent,claude-haiku-4.5,claude-sonnet-4.6,claude-sonnet-5,small-agentunsupported_multi_variantoutput_formatdailycodemetricsexecutive_summary,full_detail,steunsupported_multi_variantprompt_styledailycommunityattributionconcise,verboseinsufficient_observationsoutput_formatdailycompilerqualityconcise,detailed,steunsupported_multi_variantmodel_sizedailydochealeragent,claude-haiku-4.5,claude-sonnet-4.6,claude-sonnet-5,small-agentunsupported_multi_variantmodel_sizedailydocupdateragent,claude-haiku-4.5,claude-sonnet-4.6,claude-sonnet-5,small-agentunsupported_multi_variantreasoning_depthdailyfactmulti_candidate,single_passguardrail_unsupportedmodel_sizedailyfunctionnamerinsufficient_observationsoutput_formatdailyissuesreportcollapsible,inline,steunsupported_multi_variantprompt_styledailynewsconcise,detailedinsufficient_observationsremove_redundant_context_v1dailyrenderingscriptsverifierinsufficient_observationslog_fetch_strategydailysafeoutputoptimizerinsufficient_observationsreasoning_depthdailysecurityredteamiterative,single_passguardrail_unsupportedsemgrep_output_formatdailysemgrepscanbullet_list,prose,structured_sectionsunsupported_multi_varianttimeout_settingdailysubagentoptimizerdefault,relaxed,tightunsupported_multi_variantcaveman_modedataflowprdiscussiondatasetno,yesinsufficient_samplesoutput_formatdeepreportannotated_brief,executive_brief,full_briefing,steunsupported_multi_variantsummary_detaildependabotcampaignbrief,detailedinsufficient_samplesprompt_styledependabotgocheckerconcise,detailed,step_by_stepunsupported_multi_varianttool_verbositygpcleanfull_bash,minimal_toolsetinsufficient_observationsprompt_styleissuearboristconcise,detailedinsufficient_observationsreasoning_depthplanbaseline,deep,shallowunsupported_multi_variantremove_redundant_context_v1prsouschefinsufficient_observationssub_agent_strategysmokeantigravitysingle_agent,sub_agentsinsufficient_observationscavemansmokecopilotno,yesinsufficient_observationssubagent_modelsmokecopilotlarge,smallinsufficient_observationscavemansmokecopilotaoaiapikeyno,yesinsufficient_observationssubagent_modelsmokecopilotaoaiapikeylarge,smallinsufficient_observationscavemansmokecopilotaoaientrano,yesinsufficient_observationssubagent_modelsmokecopilotaoaientralarge,smallinsufficient_observationssub_agent_strategysmokecopilotsubagentsdelegated_sequential,inline_strict,single_agent_controlunsupported_multi_variantsub_agent_strategysmokegeminisingle_agent,sub_agentsinsufficient_observationssub_agent_decompositionsmokepiparallel_sub_agents,single_agentinsufficient_observationsprompt_style_testsmokeprojectinsufficient_observationssub_agent_strategysmoketemporaryidinsufficient_observationsmodel_sizetestqualitysentinelclaude-haiku-4.5,claude-sonnet-5insufficient_samplestone_styletypistconversational,formalinsufficient_observationsprefetch_strategyweeklyblogpostwritereager,lazyinsufficient_samples🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
unsupported_multi_variant(14 experiments) —output_format/model_size/reasoning_depthdesigns with 3+ variants (e.g.deepreport,dailycodemetrics,dailycavemanoptimizer,plan) cannot reach a decision until the core analyzer supports either a designated control + N pairwise candidate comparisons or an omnibus test with post-hoc contrasts.guardrail_unsupportedmetrics (3 experiments) —dailysecurityredteam,dailyfact, andcicoachall fail onrun_success_rate/empty_output_rateguardrails that have no supported native metric source; add these as first-class native metrics so mandatory guardrails can evaluate.insufficient_observationsgap (25 experiments) — the majority ofREADYexperiments (e.g.smokecopilot,typist,dailynews,agentperformanceanalyzer) have balanced sample counts but no per-run primary-metric observation recorded; prioritize wiringoutput_length_chars/aic/token_count-style metrics into the observation pipeline so these can convert to real decisions.Top 3 eval actions
unsupported_multi_variant(e.g.daily-code-metrics.md,plan.md) to explicitly grade whether the agent recognized the >2-variant limitation instead of silently reporting a false decision.daily-security-red-team.md,daily-fact.md, andci-coach.mdprompts that checks whether guardrail metrics are actually emitted per run, surfacing theguardrail_unsupportedgap earlier than the weekly experiment report.insufficient_observations(smoke-copilot.mdfamily,typist.md,daily-news.md), add eval coverage asserting the primarymetric:field frontmatter value is actually populated in run logs.Decision-pipeline gaps for the next PR
guardrail_unsupported:run_success_rateandempty_output_rateguardrails declared in frontmatter have no backing native metric extractor yet (dailysecurityredteam,dailyfact,cicoach).unsupported_multi_variant: 14 experiments with 3+ variants need a normalized core comparison strategy (pairwise vs. control, or omnibus + correction) instead of remaining permanentlyINCONCLUSIVE.smokecopilot/smokecopilotaoaiapikey/smokecopilotaoaientrafactorial cells) are currently a reporting-only overlay; they should become a normalized core analysis signal soSPARSE_CELL_RISKcan gatePROMOTEdecisions automatically rather than via this report's safety hold.All reactions