You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
51 experiments analysed across 48 workflows. 22 have reached minimum sample size (READY). No experiment reached a PROMOTE or REJECT decision today — the core decision layer currently reports EXTEND or INCONCLUSIVE for every experiment, driven almost entirely by a data-pipeline gap: per-run native primary-metric observations (e.g. token counts, success rates, quality scores) are not yet wired into the decision layer for most workflows, and core analysis does not yet support statistical comparison across 3+ variants.
Experiment : prompt_compression
Workflow : agent-performance-analyzer.md
Hypothesis : H0: no change in effective_tokens. H1: caveman reduces tokens by ≥20% while maintaining quality ≥90%
min_samples: 20 per variant | total_runs: 113
caveman n=62
verbose n=51
Balance test : chi_square p=0.301 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : sub_agent_strategy
Workflow : agent-persona-explorer.md
Hypothesis : H0: no change in effective_tokens or duration. H1: batch reduces tokens by ≥20% and duration by ≥15% without quality loss
min_samples: 20 per variant | total_runs: 113
batch n=59
per_scenario n=54
Balance test : chi_square p=0.643 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : tone_variant
Workflow : aw-failure-investigator.md
Hypothesis : H0: no change in output_length_chars across tone variants. H1: assertive tone produces shorter, more actionable outputs than clinical or narrative, with equivalent or better sub-issue quality.
min_samples: 20 per variant | total_runs: 406
assertive n=140
clinical n=135
narrative n=131
Balance test : chi_square p=0.858 (balanced)
Core decision : INCONCLUSIVE
Reason code : unsupported_multi_variant
Rationale : automatic decisions currently require exactly two variants
----------------------------------------------------------------------
Experiment : prompt_style
Workflow : ci-coach.md
Hypothesis : H0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
min_samples: 20 per variant | total_runs: 85
concise n=37
detailed n=48
Balance test : chi_square p=0.000 (balanced)
Core decision : EXTEND
Reason code : guardrail_unsupported
Rationale : mandatory guardrail "run_success_rate" is not backed by a supported metric observation
----------------------------------------------------------------------
Experiment : sub_agent_strategy
Workflow : daily-agentrx-trace-optimizer.md
Hypothesis : H0: no change in issue quality or run success rate. H1: sub_agents variant yields higher evidence completeness score with equal or lower token cost
min_samples: 20 per variant | total_runs: 74
single_agent n=37
sub_agents n=37
Balance test : chi_square p=1.000 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : prompt_style
Workflow : daily-astrostylelite-markdown-spellcheck.md
Hypothesis : Concise prompt reduces token consumption ≥20% without degrading fix precision. H0: no difference in fix rate.
min_samples: 20 per variant | total_runs: 131
concise n=67
detailed n=64
Balance test : chi_square p=0.783 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : prompt_style
Workflow : daily-community-attribution.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 129
concise n=64
verbose n=65
Balance test : chi_square p=0.891 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : reasoning_depth
Workflow : daily-fact.md
Hypothesis : H0: no change in discussion engagement rate. H1: multi_candidate produces more novel verses with higher reaction counts (expected +20% reactions).
min_samples: 30 per variant | total_runs: 86
multi_candidate n=34
single_pass n=52
Balance test : chi_square p=0.049 (imbalanced)
Core decision : EXTEND
Reason code : guardrail_unsupported
Rationale : mandatory guardrail "empty_output_rate" is not backed by a supported metric observation
----------------------------------------------------------------------
Experiment : prompt_style
Workflow : daily-news.md
Hypothesis : H0: no change in output quality. H1: concise prompt reduces token usage by ≥20% with no significant drop in output completeness score
min_samples: 20 per variant | total_runs: 101
concise n=54
detailed n=47
Balance test : chi_square p=0.493 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : reasoning_depth
Workflow : daily-security-red-team.md
Hypothesis : H0: no change in finding quality. H1: iterative reduces false-positive rate by >=20% at <=30% token overhead
min_samples: 30 per variant | total_runs: 122
iterative n=74
single_pass n=48
Balance test : chi_square p=0.018 (imbalanced)
Core decision : EXTEND
Reason code : guardrail_unsupported
Rationale : mandatory guardrail "run_success_rate" is not backed by a supported metric observation
----------------------------------------------------------------------
Experiment : tool_verbosity
Workflow : gpclean.md
Hypothesis : H0: no change in token consumption. H1: minimal toolset reduces tokens by 10-15% while maintaining issue quality (detection accuracy + alternative research depth)
min_samples: 20 per variant | total_runs: 110
full_bash n=52
minimal_toolset n=58
Balance test : chi_square p=0.575 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : prompt_style
Workflow : issue-arborist.md
Hypothesis : H0: no change in links_created. H1: detailed instructions produce ≥15% more correct links per run
min_samples: 20 per variant | total_runs: 69
concise n=37
detailed n=32
Balance test : chi_square p=0.555 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : sub_agent_strategy
Workflow : smokeantigravity (file removed)
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 208
single_agent n=99
sub_agents n=109
Balance test : chi_square p=0.495 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : caveman
Workflow : smoke-copilot.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 356
no n=178
yes n=178
Balance test : chi_square p=1.000 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : subagent_model
Workflow : smoke-copilot.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 336
large n=168
small n=168
Balance test : chi_square p=1.000 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : caveman
Workflow : smoke-copilot-aoai-apikey.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 181
no n=91
yes n=90
Balance test : chi_square p=0.899 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : subagent_model
Workflow : smoke-copilot-aoai-apikey.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 181
large n=90
small n=91
Balance test : chi_square p=0.899 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : caveman
Workflow : smoke-copilot-aoai-entra.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 154
no n=77
yes n=77
Balance test : chi_square p=1.000 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : subagent_model
Workflow : smoke-copilot-aoai-entra.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 154
large n=77
small n=77
Balance test : chi_square p=1.000 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : sub_agent_strategy
Workflow : smoke-gemini.md
Hypothesis : H0: no change in effective_tokens. H1: sub_agents reduces tokens by >=20%
min_samples: 20 per variant | total_runs: 264
single_agent n=124
sub_agents n=140
Balance test : chi_square p=0.326 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : sub_agent_decomposition
Workflow : smoke-pi.md
Hypothesis : (not specified)
min_samples: 20 per variant | total_runs: 76
parallel_sub_agents n=42
single_agent n=34
Balance test : chi_square p=0.362 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
Experiment : tone_style
Workflow : typist.md
Hypothesis : H0: no change in engagement. H1: conversational tone increases discussion views+reactions+comments by 20%+ while maintaining analysis quality
min_samples: 20 per variant | total_runs: 78
conversational n=43
formal n=35
Balance test : chi_square p=0.368 (balanced)
Core decision : EXTEND
Reason code : insufficient_observations
Rationale : primary metric observations are not available for statistical comparison
----------------------------------------------------------------------
📊 Summary
View Full Experiments Table
Experiment
Workflow
Control
Candidate
Evidence
Guardrails
Core decision
Report action
prompt_compression
agent-performance-analyzer.md
caveman
verbose
p=0.301
—
EXTEND
EXTEND
sub_agent_strategy
agent-persona-explorer.md
batch
per_scenario
p=0.643
—
EXTEND
EXTEND
sub_agent_strategy
architecture-guardian.md
single_agent
sub_agents
p=0.524
—
EXTEND
EXTEND
audit_decomposition
audit-workflows.md
single_agent
phased_sub_agents
p=0.334
—
EXTEND
EXTEND
tone_variant
aw-failure-investigator.md
—
—
p=0.858
—
INCONCLUSIVE
INCONCLUSIVE
prompt_style
blog-auditor.md
concise
detailed
p=0.319
—
EXTEND
EXTEND
tone_variant
breaking-change-checker.md
neutral
urgent
p=0.535
—
EXTEND
EXTEND
prompt_style
ci-coach.md
detailed
concise
p=0.000
n/a (unsupported)
EXTEND
EXTEND
output_format
copilot-agent-analysis.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
sub_agent_strategy
daily-agentrx-trace-optimizer.md
single_agent
sub_agents
p=1.000
—
EXTEND
EXTEND
detail_level
daily-architecture-diagram.md
brief
comprehensive
p=0.319
—
EXTEND
EXTEND
prompt_style
daily-astrostylelite-markdown-spellcheck.md
concise
detailed
p=0.783
—
EXTEND
EXTEND
model_size
daily-cache-strategy-analyzer.md
gpt-5.3-codex
gpt-5.3-codex-spark
p=0.659
—
EXTEND
EXTEND
model_size
daily-caveman-optimizer.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
output_format
daily-code-metrics.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
prompt_style
daily-community-attribution.md
concise
verbose
p=0.891
—
EXTEND
EXTEND
output_format
daily-compiler-quality.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
model_size
daily-doc-healer.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
model_size
daily-doc-updater.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
reasoning_depth
daily-fact.md
single_pass
multi_candidate
p=0.049
n/a (unsupported)
EXTEND
EXTEND
model_size
daily-function-namer.md
—
—
p=0.000
—
EXTEND
EXTEND
output_format
daily-issues-report.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
prompt_style
daily-news.md
concise
detailed
p=0.493
—
EXTEND
EXTEND
remove_redundant_context_v1
daily-rendering-scripts-verifier.md
—
—
p=0.000
—
EXTEND
EXTEND
log_fetch_strategy
daily-safe-output-optimizer.md
eager
lazy
p=0.000
—
EXTEND
EXTEND
reasoning_depth
daily-security-red-team.md
single_pass
iterative
p=0.018
n/a (unsupported)
EXTEND
EXTEND
semgrep_output_format
daily-semgrep-scan.md
—
—
p=0.023
—
INCONCLUSIVE
INCONCLUSIVE
timeout_setting
dailysubagentoptimizer (removed)
—
—
p=0.781
—
INCONCLUSIVE
INCONCLUSIVE
caveman_mode
dataflow-pr-discussion-dataset.md
no
yes
p=0.285
—
EXTEND
EXTEND
output_format
deep-report.md
—
—
p=0.000
—
INCONCLUSIVE
INCONCLUSIVE
summary_detail
dependabotcampaign (removed)
brief
detailed
p=1.000
—
EXTEND
EXTEND
prompt_style
dependabot-go-checker.md
—
—
p=0.819
—
INCONCLUSIVE
INCONCLUSIVE
tool_verbosity
gpclean.md
full_bash
minimal_toolset
p=0.575
—
EXTEND
EXTEND
prompt_style
issue-arborist.md
concise
detailed
p=0.555
—
EXTEND
EXTEND
reasoning_depth
plan.md
—
—
p=0.369
—
INCONCLUSIVE
INCONCLUSIVE
remove_redundant_context_v1
pr-sous-chef.md
—
—
p=0.000
—
EXTEND
EXTEND
sub_agent_strategy
smokeantigravity (removed)
single_agent
sub_agents
p=0.495
—
EXTEND
EXTEND
caveman
smoke-copilot.md
no
yes
p=1.000
—
EXTEND
EXTEND
subagent_model
smoke-copilot.md
large
small
p=1.000
—
EXTEND
EXTEND
caveman
smoke-copilot-aoai-apikey.md
no
yes
p=0.899
—
EXTEND
EXTEND
subagent_model
smoke-copilot-aoai-apikey.md
large
small
p=0.899
—
EXTEND
EXTEND
caveman
smoke-copilot-aoai-entra.md
no
yes
p=1.000
—
EXTEND
EXTEND
subagent_model
smoke-copilot-aoai-entra.md
large
small
p=1.000
—
EXTEND
EXTEND
sub_agent_strategy
smoke-copilot-sub-agents.md
—
—
p=0.802
—
INCONCLUSIVE
INCONCLUSIVE
sub_agent_strategy
smoke-gemini.md
single_agent
sub_agents
p=0.326
—
EXTEND
EXTEND
sub_agent_decomposition
smoke-pi.md
parallel_sub_agents
single_agent
p=0.362
—
EXTEND
EXTEND
prompt_style_test
smoke-project.md
—
—
p=0.000
—
EXTEND
EXTEND
sub_agent_strategy
smoke-temporary-id.md
single_agent
sub_agents
p=0.000
—
EXTEND
EXTEND
model_size
test-quality-sentinel.md
—
—
p=0.000
—
EXTEND
EXTEND
tone_style
typist.md
conversational
formal
p=0.368
—
EXTEND
EXTEND
prefetch_strategy
weekly-blog-post-writer.md
eager
lazy
p=0.410
—
EXTEND
EXTEND
Descriptive window: last 30 runs per workflow (where available) · Decision thresholds: from decision_policy (95% confidence, min_samples per experiment config)
Run: 34576316456
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
awfailureinvestigator / tone_variant and copilotagentanalysis, dailycodemetrics, dailycompilerquality, dailyissuesreport, dailysemgrepscan, dailydochealer, dailydocupdater, dailycavemanoptimizer, dependabotgochecker, plan, smokecopilotsubagents, dailysubagentoptimizer, deepreport are all INCONCLUSIVE due to unsupported_multi_variant — prioritize adding 3+ variant statistical support (e.g. ANOVA/Kruskal-Wallis) to the core analyzer so these 14 experiments (largest single blocker) can progress.
cicoach, dailyfact, dailysecurityredteam are READY but blocked by guardrail_unsupported — wire native observation collection for their guardrail metrics (run_success_rate, quality scores) so guardrail pass/fail can be evaluated and these experiments can reach a real decision.
The largest group (25 experiments, insufficient_observations) is READY or actively accumulating samples but has no primary-metric per-run observation pipeline at all — this is the single highest-leverage fix; prioritize effective_tokens, run_duration_ms, and run_success_rate observation collection since they are the most common primary metrics declared in frontmatter.
Top 3 eval actions
Workflows testing output_format variants (copilotagentanalysis, dailycodemetrics, dailycompilerquality, dailyissuesreport, deepreport, dailysemgrepscan) declare eval:output_format_adherence-style secondary metrics but lack graders — add an eval question per workflow asking "Does the report match the declared output_format (prose/ste/structured)?" to produce the observation the core analyzer needs.
dailycavemanoptimizer, dailydochealer, dailydocupdater, testqualitysentinel, dailycachestrategyanalyzer test model_size — add an eval question scoring output quality (e.g. "Did the smaller model produce an equally actionable recommendation?") to give the guardrail pipeline a quality signal alongside cost.
agentperformanceanalyzer, agentpersonaexplorer, dailyagentrxtraceoptimizer, smokeantigravity, smokegemini, smokepi test sub_agent_strategy — tighten the existing output_quality_score / scenarios_analyzed eval questions to be binary pass/fail so they map cleanly onto guardrail_metrics thresholds instead of free-text scores.
Decision-pipeline gaps for next PR
insufficient_observations (25 experiments) and guardrail_unsupported (3 experiments) both stem from the same root cause: no per-run primary/guardrail metric write-back into experiments/* branch state. This is the top priority gap.
unsupported_multi_variant (14 experiments) blocks every 3+ variant test regardless of sample size — core analysis needs an omnibus test (ANOVA / Kruskal-Wallis / chi-square-based) with pairwise post-hoc comparison and Bonferroni correction (the bonferroni_alpha field already exists in the schema but is unused while this gap persists).
No workflow currently runs 2+ concurrent experiments, so the factorial-interaction diagnostic (buildFactorialInteractionCells) was not exercised this run. Once dailycavemanoptimizer-style multi-experiment workflows appear, promote the interaction check from a reporting-only safety hold into a normalized core analysis signal.
Descriptive window: last 30 runs per workflow · Decision thresholds: from decision_policy
Run: 34576316456
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-11
51 experiments analysed across 48 workflows. 22 have reached minimum sample size (
READY). No experiment reached aPROMOTEorREJECTdecision today — the core decision layer currently reportsEXTENDorINCONCLUSIVEfor every experiment, driven almost entirely by a data-pipeline gap: per-run native primary-metric observations (e.g. token counts, success rates, quality scores) are not yet wired into the decision layer for most workflows, and core analysis does not yet support statistical comparison across 3+ variants.⚡ Quick Stats
READY)COLLECTING)📋 Experiments by Decision Reason
🟡 EXTEND — missing native metric observations — 25 experiments
primary-metric per-run observations are not available for statistical comparison
agent-performance-analyzer.mdprompt_compressioncaveman=62,verbose=51#33280agent-persona-explorer.mdsub_agent_strategybatch=59,per_scenario=54daily-agentrx-trace-optimizer.mdsub_agent_strategysingle_agent=37,sub_agents=37daily-astrostylelite-markdown-spellcheck.mdprompt_styleconcise=67,detailed=64daily-community-attribution.mdprompt_styleconcise=64,verbose=65daily-function-namer.mdmodel_sizedaily-news.mdprompt_styleconcise=54,detailed=47#31190daily-rendering-scripts-verifier.mdremove_redundant_context_v1daily-safe-output-optimizer.mdlog_fetch_strategy#38094gpclean.mdtool_verbosityfull_bash=52,minimal_toolset=58issue-arborist.mdprompt_styleconcise=37,detailed=32#30015pr-sous-chef.mdremove_redundant_context_v1smokeantigravity (file removed)sub_agent_strategysingle_agent=99,sub_agents=109smoke-copilot.mdcavemanno=178,yes=178smoke-copilot.mdsubagent_modellarge=168,small=168smoke-copilot-aoai-apikey.mdcavemanno=91,yes=90smoke-copilot-aoai-apikey.mdsubagent_modellarge=90,small=91smoke-copilot-aoai-entra.mdcavemanno=77,yes=77smoke-copilot-aoai-entra.mdsubagent_modellarge=77,small=77smoke-gemini.mdsub_agent_strategysingle_agent=124,sub_agents=140smoke-pi.mdsub_agent_decompositionparallel_sub_agents=42,single_agent=34smoke-project.mdprompt_style_test#37302smoke-temporary-id.mdsub_agent_strategytest-quality-sentinel.mdmodel_size#43530typist.mdtone_styleconversational=43,formal=35#34032⚪ INCONCLUSIVE — 3+ variant test unsupported — 14 experiments
core analysis does not yet support statistical comparison across 3+ variants
aw-failure-investigator.mdtone_variantassertive=140,clinical=135,narrative=131#36105copilot-agent-analysis.mdoutput_formatprose=34,ste=3,structured=43daily-caveman-optimizer.mdmodel_sizeagent=1,claude-haiku-4.5=57,claude-sonnet-4.6=34,claude-sonnet-5=5,small-agent=2daily-code-metrics.mdoutput_formatexecutive_summary=50,full_detail=55,ste=12#1daily-compiler-quality.mdoutput_formatconcise=50,detailed=56,ste=12#32390daily-doc-healer.mdmodel_sizeagent=2,claude-haiku-4.5=50,claude-sonnet-4.6=43,claude-sonnet-5=3,small-agent=1daily-doc-updater.mdmodel_sizeagent=1,claude-haiku-4.5=42,claude-sonnet-4.6=47,claude-sonnet-5=6,small-agent=2daily-issues-report.mdoutput_formatcollapsible=55,inline=58,ste=16#30573daily-semgrep-scan.mdsemgrep_output_formatbullet_list=28,prose=27,structured_sections=22#32795dailysubagentoptimizer (file removed)timeout_settingdefault=2,relaxed=1,tight=1deep-report.mdoutput_formatannotated_brief=43,executive_brief=64,full_briefing=51,ste=19dependabot-go-checker.mdprompt_styleconcise=16,detailed=16,step_by_step=13plan.mdreasoning_depthbaseline=4,deep=2,shallow=1#42941smoke-copilot-sub-agents.mdsub_agent_strategydelegated_sequential=14,inline_strict=12,single_agent_control=15#47551🟡 EXTEND — below min_samples — 9 experiments
one or more variants have not yet reached the configured min_samples threshold
architecture-guardian.mdsub_agent_strategysingle_agent=17,sub_agents=21audit-workflows.mdaudit_decompositionphased_sub_agents=38,single_agent=30#43177blog-auditor.mdprompt_styleconcise=3,detailed=6#32603breaking-change-checker.mdtone_variantneutral=22,urgent=18#42467daily-architecture-diagram.mddetail_levelbrief=10,comprehensive=6#31926daily-cache-strategy-analyzer.mdmodel_sizegpt-5.3-codex=2,gpt-5.3-codex-spark=3dataflow-pr-discussion-dataset.mdcaveman_modeno=5,yes=9#37102dependabotcampaign (file removed)summary_detailbrief=6,detailed=6weekly-blog-post-writer.mdprefetch_strategyeager=5,lazy=8#38590🟡 EXTEND — guardrail metric unsupported — 3 experiments
a required guardrail metric has no native observation pipeline yet
ci-coach.mdprompt_styleconcise=37,detailed=48#32335daily-fact.mdreasoning_depthmulti_candidate=34,single_pass=52#31324daily-security-red-team.mdreasoning_depthiterative=74,single_pass=48#31673📈 Full Sample-Size Progress Table (all 51 experiments)
agent-performance-analyzer.mdprompt_compressionagent-persona-explorer.mdsub_agent_strategyarchitecture-guardian.mdsub_agent_strategyaudit-workflows.mdaudit_decompositionaw-failure-investigator.mdtone_variantblog-auditor.mdprompt_stylebreaking-change-checker.mdtone_variantci-coach.mdprompt_stylecopilot-agent-analysis.mdoutput_formatdaily-agentrx-trace-optimizer.mdsub_agent_strategydaily-architecture-diagram.mddetail_leveldaily-astrostylelite-markdown-spellcheck.mdprompt_styledaily-cache-strategy-analyzer.mdmodel_sizedaily-caveman-optimizer.mdmodel_sizedaily-code-metrics.mdoutput_formatdaily-community-attribution.mdprompt_styledaily-compiler-quality.mdoutput_formatdaily-doc-healer.mdmodel_sizedaily-doc-updater.mdmodel_sizedaily-fact.mdreasoning_depthdaily-function-namer.mdmodel_sizedaily-issues-report.mdoutput_formatdaily-news.mdprompt_styledaily-rendering-scripts-verifier.mdremove_redundant_context_v1daily-safe-output-optimizer.mdlog_fetch_strategydaily-security-red-team.mdreasoning_depthdaily-semgrep-scan.mdsemgrep_output_formatdailysubagentoptimizer (file removed)timeout_settingdataflow-pr-discussion-dataset.mdcaveman_modedeep-report.mdoutput_formatdependabotcampaign (file removed)summary_detaildependabot-go-checker.mdprompt_stylegpclean.mdtool_verbosityissue-arborist.mdprompt_styleplan.mdreasoning_depthpr-sous-chef.mdremove_redundant_context_v1smokeantigravity (file removed)sub_agent_strategysmoke-copilot.mdcavemansmoke-copilot.mdsubagent_modelsmoke-copilot-aoai-apikey.mdcavemansmoke-copilot-aoai-apikey.mdsubagent_modelsmoke-copilot-aoai-entra.mdcavemansmoke-copilot-aoai-entra.mdsubagent_modelsmoke-copilot-sub-agents.mdsub_agent_strategysmoke-gemini.mdsub_agent_strategysmoke-pi.mdsub_agent_decompositionsmoke-project.mdprompt_style_testsmoke-temporary-id.mdsub_agent_strategytest-quality-sentinel.mdmodel_sizetypist.mdtone_styleweekly-blog-post-writer.mdprefetch_strategy🔬 Detailed ASCII Comparison — READY experiments
📊 Summary
View Full Experiments Table
prompt_compressionagent-performance-analyzer.mdcavemanverbosesub_agent_strategyagent-persona-explorer.mdbatchper_scenariosub_agent_strategyarchitecture-guardian.mdsingle_agentsub_agentsaudit_decompositionaudit-workflows.mdsingle_agentphased_sub_agentstone_variantaw-failure-investigator.md——prompt_styleblog-auditor.mdconcisedetailedtone_variantbreaking-change-checker.mdneutralurgentprompt_styleci-coach.mddetailedconciseoutput_formatcopilot-agent-analysis.md——sub_agent_strategydaily-agentrx-trace-optimizer.mdsingle_agentsub_agentsdetail_leveldaily-architecture-diagram.mdbriefcomprehensiveprompt_styledaily-astrostylelite-markdown-spellcheck.mdconcisedetailedmodel_sizedaily-cache-strategy-analyzer.mdgpt-5.3-codexgpt-5.3-codex-sparkmodel_sizedaily-caveman-optimizer.md——output_formatdaily-code-metrics.md——prompt_styledaily-community-attribution.mdconciseverboseoutput_formatdaily-compiler-quality.md——model_sizedaily-doc-healer.md——model_sizedaily-doc-updater.md——reasoning_depthdaily-fact.mdsingle_passmulti_candidatemodel_sizedaily-function-namer.md——output_formatdaily-issues-report.md——prompt_styledaily-news.mdconcisedetailedremove_redundant_context_v1daily-rendering-scripts-verifier.md——log_fetch_strategydaily-safe-output-optimizer.mdeagerlazyreasoning_depthdaily-security-red-team.mdsingle_passiterativesemgrep_output_formatdaily-semgrep-scan.md——timeout_settingdailysubagentoptimizer (removed)——caveman_modedataflow-pr-discussion-dataset.mdnoyesoutput_formatdeep-report.md——summary_detaildependabotcampaign (removed)briefdetailedprompt_styledependabot-go-checker.md——tool_verbositygpclean.mdfull_bashminimal_toolsetprompt_styleissue-arborist.mdconcisedetailedreasoning_depthplan.md——remove_redundant_context_v1pr-sous-chef.md——sub_agent_strategysmokeantigravity (removed)single_agentsub_agentscavemansmoke-copilot.mdnoyessubagent_modelsmoke-copilot.mdlargesmallcavemansmoke-copilot-aoai-apikey.mdnoyessubagent_modelsmoke-copilot-aoai-apikey.mdlargesmallcavemansmoke-copilot-aoai-entra.mdnoyessubagent_modelsmoke-copilot-aoai-entra.mdlargesmallsub_agent_strategysmoke-copilot-sub-agents.md——sub_agent_strategysmoke-gemini.mdsingle_agentsub_agentssub_agent_decompositionsmoke-pi.mdparallel_sub_agentssingle_agentprompt_style_testsmoke-project.md——sub_agent_strategysmoke-temporary-id.mdsingle_agentsub_agentsmodel_sizetest-quality-sentinel.md——tone_styletypist.mdconversationalformalprefetch_strategyweekly-blog-post-writer.mdeagerlazy🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
awfailureinvestigator/tone_variantandcopilotagentanalysis,dailycodemetrics,dailycompilerquality,dailyissuesreport,dailysemgrepscan,dailydochealer,dailydocupdater,dailycavemanoptimizer,dependabotgochecker,plan,smokecopilotsubagents,dailysubagentoptimizer,deepreportare allINCONCLUSIVEdue tounsupported_multi_variant— prioritize adding 3+ variant statistical support (e.g. ANOVA/Kruskal-Wallis) to the core analyzer so these 14 experiments (largest single blocker) can progress.cicoach,dailyfact,dailysecurityredteamareREADYbut blocked byguardrail_unsupported— wire native observation collection for their guardrail metrics (run_success_rate, quality scores) so guardrail pass/fail can be evaluated and these experiments can reach a real decision.insufficient_observations) isREADYor actively accumulating samples but has no primary-metric per-run observation pipeline at all — this is the single highest-leverage fix; prioritizeeffective_tokens,run_duration_ms, andrun_success_rateobservation collection since they are the most common primary metrics declared in frontmatter.Top 3 eval actions
output_formatvariants (copilotagentanalysis,dailycodemetrics,dailycompilerquality,dailyissuesreport,deepreport,dailysemgrepscan) declareeval:output_format_adherence-style secondary metrics but lack graders — add an eval question per workflow asking "Does the report match the declared output_format (prose/ste/structured)?" to produce the observation the core analyzer needs.dailycavemanoptimizer,dailydochealer,dailydocupdater,testqualitysentinel,dailycachestrategyanalyzertestmodel_size— add an eval question scoring output quality (e.g. "Did the smaller model produce an equally actionable recommendation?") to give the guardrail pipeline a quality signal alongside cost.agentperformanceanalyzer,agentpersonaexplorer,dailyagentrxtraceoptimizer,smokeantigravity,smokegemini,smokepitestsub_agent_strategy— tighten the existingoutput_quality_score/scenarios_analyzedeval questions to be binary pass/fail so they map cleanly ontoguardrail_metricsthresholds instead of free-text scores.Decision-pipeline gaps for next PR
insufficient_observations(25 experiments) andguardrail_unsupported(3 experiments) both stem from the same root cause: no per-run primary/guardrail metric write-back intoexperiments/*branch state. This is the top priority gap.unsupported_multi_variant(14 experiments) blocks every 3+ variant test regardless of sample size — core analysis needs an omnibus test (ANOVA / Kruskal-Wallis / chi-square-based) with pairwise post-hoc comparison and Bonferroni correction (thebonferroni_alphafield already exists in the schema but is unused while this gap persists).buildFactorialInteractionCells) was not exercised this run. Oncedailycavemanoptimizer-style multi-experiment workflows appear, promote the interaction check from a reporting-only safety hold into a normalized core analysis signal.All reactions