[experiments] Daily Experiment Report — 2026-09-27 #63805
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #63963. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Executive Summary
This report covers 51 experiments across 48 workflows with active
experiments:frontmatter. 20 experiments have a linked tracking issue; 31 do not (legacy shorthand config or noissue:field, and 3 branches with no matching source file).Decision distribution: EXTEND=37, INCONCLUSIVE=14
Readiness distribution: COLLECTING=25, READY=26
Warning
dailyfactis failing 100% of recent runs (30/30 failures) across bothreasoning_depthvariants (single_passn=15,multi_candidaten=14) in the last 30 runs — confirmed against raw run conclusions, not a data-parsing artifact. This is a workflow health incident independent of the A/B test and should be triaged first.Quick Stats
Tracked Experiments (linked to a GitHub issue)
Per-experiment detail (20 experiments)
agentperformanceanalyzer— prompt_compression (issue#33280)Display name: Agent Performance Analyzer - Meta-Orchestrator | Readiness: READY | Decision: EXTEND
auditworkflows— audit_decomposition (issue#43177)Display name: Agentic Workflow Audit Agent | Readiness: COLLECTING | Decision: EXTEND
awfailureinvestigator— tone_variant (issue#36105)Display name: [aw] Failure Investigator (6h) | Readiness: READY | Decision: INCONCLUSIVE
breakingchangechecker— tone_variant (issue#42467)Display name: Breaking Change Checker | Readiness: READY | Decision: EXTEND
cicoach— prompt_style (issue#32335)Display name: CI Optimization Coach | Readiness: READY | Decision: EXTEND
dailyarchitecturediagram— detail_level (issue#31926)Display name: 1. List all Go packages with their doc comments | Readiness: COLLECTING | Decision: EXTEND
dailycompilerquality— output_format (issue#32390)Display name: Daily Compiler Quality Check | Readiness: COLLECTING | Decision: INCONCLUSIVE
dailyfact— reasoning_depth (issue#31324)Display name: Daily Fact | Readiness: READY | Decision: EXTEND
dailyissuesreport— output_format (issue#30573)Display name: Daily Issues Report Generator | Readiness: READY | Decision: INCONCLUSIVE
dailynews— prompt_style (issue#31190)Display name: Daily News | Readiness: READY | Decision: EXTEND
dailysafeoutputoptimizer— log_fetch_strategy (issue#38094)Display name: Daily Safe Output Tool Optimizer | Readiness: COLLECTING | Decision: EXTEND
dailysecurityredteam— reasoning_depth (issue#31673)Display name: Cache directory setup | Readiness: READY | Decision: EXTEND
dailysemgrepscan— semgrep_output_format (issue#32795)Display name: Daily Semgrep Scan | Readiness: COLLECTING | Decision: INCONCLUSIVE
dataflowprdiscussiondataset— caveman_mode (issue#37102)Display name: DataFlow PR & Discussion Dataset Builder | Readiness: COLLECTING | Decision: EXTEND
issuearborist— prompt_style (issue#30015)Display name: Issue Arborist | Readiness: READY | Decision: EXTEND
plan— reasoning_depth (issue#42941)Display name: Plan Command | Readiness: COLLECTING | Decision: INCONCLUSIVE
smokecopilotsubagents— sub_agent_strategy (issue#47551)Display name: Smoke Copilot Sub Agents | Readiness: COLLECTING | Decision: INCONCLUSIVE
testqualitysentinel— model_size (issue#43530)Display name: Test Quality Sentinel | Readiness: READY | Decision: EXTEND
typist— tone_style (issue#34032)Display name: Typist - Go Type Analysis | Readiness: READY | Decision: EXTEND
weeklyblogpostwriter— prefetch_strategy (issue#38590)Display name: Weekly Blog Post Writer | Readiness: COLLECTING | Decision: EXTEND
Untracked Experiments (no linked issue)
Summary (31 experiments — legacy shorthand config or missing source file)
agentpersonaexplorerarchitectureguardianblogauditorcopilotagentanalysisdailyagentrxtraceoptimizerdailyastrostylelitemarkdownspellcheckdailycachestrategyanalyzerdailycavemanoptimizerdailycodemetricsdailycommunityattributiondailydochealerdailydocupdaterdailyfunctionnamerdailyrenderingscriptsverifierdailysubagentoptimizerdeepreportdependabotcampaigndependabotgocheckergpcleanprsouschefsmokeantigravitysmokecopilotsmokecopilotsmokecopilotaoaiapikeysmokecopilotaoaiapikeysmokecopilotaoaientrasmokecopilotaoaientrasmokegeminismokepismokeprojectsmoketemporaryidFactorial Interaction Diagnostics (multi-experiment workflows)
Three workflows run two simultaneous experiments (
caveman×subagent_model). The table below tests whether variant assignment across the two experiments is statistically independent (orthogonal randomization) — this is a randomization-health check, not an outcome-interaction test, and it only gates report recommendations when a core decision isPROMOTE(none exist in this run).smokecopilotinteraction_p_value= 0.7086 | min cell n = 85 |interaction_risk_status= OKsmokecopilotaoaiapikeyinteraction_p_value= 0.0006 | min cell n = 35 |interaction_risk_status= OKsmokecopilotaoaiapikey. No core decision isPROMOTEhere so no safety-hold applies, but this is worth flagging as a randomization/data-pipeline concern.smokecopilotaoaientrainteraction_p_value= 0.0593 | min cell n = 34 |interaction_risk_status= OKSelf-Tuning Continuation Plan
Top 3 experiment actions:
typist/tone_styleandtestqualitysentinel/model_sizeare READY with balanced assignment (chi-square not significant) — prioritize wiring a supported outcome metric (currentlyguardrail_unsupported/insufficient_observations) so these can progress past EXTEND.dailyfact/reasoning_depthshows 0% success rate across both variants (single_pass n=15, multi_candidate n=14) in the last 30 runs — this is a real signal worth investigating as a workflow health issue, independent of the A/B comparison itself.unsupported_multi_variant(3+ variant experiments the core analyzer cannot yet compare pairwise) — extend the analyzer to support >2-variant comparisons (e.g.,dailycompilerquality,awfailureinvestigator,dailyissuesreport).Top 3 eval / data-pipeline actions:
unsupportedacross all 51 analyzed experiments — zero experiments have a supported/passing guardrail metric observation. This is the single largest gap blocking anyPROMOTE/REJECTdecision; wire native guardrail metric collection into the experiment framework.smokecopilotaoaiapikey's two simultaneous experiments show a statistically significant assignment-interaction (p=0.0006) — audit the randomization logic for that workflow to confirm variant assignment is drawn independently per experiment.smokeantigravity,dependabotcampaign,dailysubagentoptimizer) have no matching.mdsource file — either restore/rename the source workflow or archive the stale experiment branch to avoid orphaned state.All reactions