You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
44 workflows in github/gh-aw currently run 47 active A/B experiments via experiments: frontmatter. 23 experiments have all variants at/above min_samples (ready for outcome analysis); 24 are still collecting data. No workflow declares a tracking issue: field, so this run posts findings only to this discussion (no issue comments/labels were applicable — see note in Step 8/9 below).
Important caveat: the p_value/chi_square figures below are the variant-assignment balance test (is traffic split as configured?), not a primary-metric significance test. None of these 47 experiments currently expose per-run outcome metrics (token_count, success_rate, etc.) in the git-branch state.json/state.jsonl used by gh aw experiments analyze — all declared guardrail_metrics report status: "unsupported". Full PROMOTE/ABANDON decisions require per-run outcome extraction (see Step 10 below) that is not yet wired up.
⚡ Quick Stats
Metric
Value
Active workflows with experiments
44
Active experiments
47
Ready for outcome analysis (min_samples reached, all variants)
23
Still collecting (EXTEND)
24
Assignment-balance test flagged imbalanced (p < 0.05)
14 (mostly multi-variant model/format experiments with skewed rollout, e.g. legacy variant labels retained from earlier experiment iterations)
Tracking issues configured
0
🟢 Ready for Analysis (23 experiments, all variants ≥ min_samples)
Significance: * p<0.05 (assignment balance test, not outcome significance)
Recommendation for this group: EXTEND (instrument outcome metrics before PROMOTE/ABANDON). Sample sizes are healthy, but without per-run outcome data (token cost, success rate, reviewer/reaction signals) none of these can be safely promoted or abandoned yet — recommend adding outcome capture to pick_experiment/finalize steps as the next iteration (see Step 10).
🟡 Still Collecting (24 experiments, at least one variant below min_samples)
Recommendation: EXTEND for all 24. The 8 output_format/model_size experiments flagged imbalanced above mostly carry legacy variant labels (agent, small-agent, or early ste runs) from earlier iterations of the same experiment key — these should be pruned from the variant set or the branch state reset so future balance tests reflect only the current variant lineup.
📊 Summary Table (all 47 experiments)
View Full Experiments Table (sorted by sample size)
See the two tables above — 23 rows in "Ready for Analysis", 24 rows in "Still Collecting".
Analysis window: full experiments/* branch state (git for-each-ref) · Significance threshold: p < 0.05 (two-tailed, assignment-balance test only — see caveat above)
Run: 32626952068
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
Instrument outcome-metric capture (token_count, success_rate, run_duration_ms) into each experiment's finalize step so gh aw experiments analyze guardrails stop reporting status: "unsupported" — currently blocking all 47 experiments from PROMOTE/ABANDON decisions.
Prune legacy/renamed variant labels from model_size and output_format experiments (dailycavemanoptimizer, dailydochealer, dailydocupdater, dailycachestrategyanalyzer, copilotagentanalysis, dailycodemetrics, dailycompilerquality, dailyissuesreport, deepreport) — stale agent/small-agent/early ste counts are skewing the balance test and diluting min_samples progress for the live variants.
Add issue: tracking fields to the highest-value, near-ready experiments (dailysecurityredteam, cicoach, dailysafeoutputoptimizer, testqualitysentinel) so future runs get automatic ready/significance/guardrail notifications per Step 8.
Top 3 eval actions
For output_format experiments comparing ste variants, tighten eval questions to explicitly score "clarity vs. information loss" — ste is consistently the lowest-n variant across 5 workflows, suggesting either low rollout weight or eval friction discouraging its selection.
For model_size/sub_agent_strategy experiments, add an eval question scoring "task completion equivalence" between small and large models to give a quality floor before promoting cost-saving variants.
Map low-scoring eval questions for prompt_style (concise vs. detailed) workflows to the token_count_per_run metric already declared in ci-coach.md, so eval quality and cost savings can be jointly reported.
Grader-ready hooks
Decision-path hook: extend the per-run outcome record (Step 2 schema: run_id, experiment, variant, conclusion, duration_ms) with a grader_score field once graders are available, so recommendation gating can weight quality alongside sample size.
Missing fields to capture now: per-run token_count/ai_credits_used, output_char_length, and a boolean guardrail_pass per declared guardrail — all currently absent from state.json/state.jsonl, which is why every guardrail in this run shows unsupported.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-08-23
44 workflows in
github/gh-awcurrently run 47 active A/B experiments viaexperiments:frontmatter. 23 experiments have all variants at/abovemin_samples(ready for outcome analysis); 24 are still collecting data. No workflow declares a trackingissue:field, so this run posts findings only to this discussion (no issue comments/labels were applicable — see note in Step 8/9 below).Important caveat: the
p_value/chi_squarefigures below are the variant-assignment balance test (is traffic split as configured?), not a primary-metric significance test. None of these 47 experiments currently expose per-run outcome metrics (token_count, success_rate, etc.) in the git-branchstate.json/state.jsonlused bygh aw experiments analyze— all declaredguardrail_metricsreportstatus: "unsupported". Full PROMOTE/ABANDON decisions require per-run outcome extraction (see Step 10 below) that is not yet wired up.⚡ Quick Stats
min_samplesreached, all variants)EXTEND)🟢 Ready for Analysis (23 experiments, all variants ≥ min_samples)
📈 View Full Table
Significance: * p<0.05 (assignment balance test, not outcome significance)
Recommendation for this group: EXTEND (instrument outcome metrics before PROMOTE/ABANDON). Sample sizes are healthy, but without per-run outcome data (token cost, success rate, reviewer/reaction signals) none of these can be safely promoted or abandoned yet — recommend adding outcome capture to
pick_experiment/finalize steps as the next iteration (see Step 10).🟡 Still Collecting (24 experiments, at least one variant below min_samples)
📈 View Full Table
Recommendation: EXTEND for all 24. The 8
output_format/model_sizeexperiments flagged imbalanced above mostly carry legacy variant labels (agent,small-agent, or earlysteruns) from earlier iterations of the same experiment key — these should be pruned from the variant set or the branch state reset so future balance tests reflect only the current variant lineup.📊 Summary Table (all 47 experiments)
View Full Experiments Table (sorted by sample size)
See the two tables above — 23 rows in "Ready for Analysis", 24 rows in "Still Collecting".
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
gh aw experiments analyzeguardrails stop reportingstatus: "unsupported"— currently blocking all 47 experiments from PROMOTE/ABANDON decisions.model_sizeandoutput_formatexperiments (dailycavemanoptimizer,dailydochealer,dailydocupdater,dailycachestrategyanalyzer,copilotagentanalysis,dailycodemetrics,dailycompilerquality,dailyissuesreport,deepreport) — staleagent/small-agent/earlystecounts are skewing the balance test and dilutingmin_samplesprogress for the live variants.issue:tracking fields to the highest-value, near-ready experiments (dailysecurityredteam,cicoach,dailysafeoutputoptimizer,testqualitysentinel) so future runs get automatic ready/significance/guardrail notifications per Step 8.Top 3 eval actions
output_formatexperiments comparingstevariants, tighten eval questions to explicitly score "clarity vs. information loss" —steis consistently the lowest-n variant across 5 workflows, suggesting either low rollout weight or eval friction discouraging its selection.model_size/sub_agent_strategyexperiments, add an eval question scoring "task completion equivalence" between small and large models to give a quality floor before promoting cost-saving variants.prompt_style(concise vs. detailed) workflows to thetoken_count_per_runmetric already declared inci-coach.md, so eval quality and cost savings can be jointly reported.Grader-ready hooks
run_id,experiment,variant,conclusion,duration_ms) with agrader_scorefield once graders are available, sorecommendationgating can weight quality alongside sample size.token_count/ai_credits_used,output_char_length, and a booleanguardrail_passper declared guardrail — all currently absent fromstate.json/state.jsonl, which is why every guardrail in this run showsunsupported.All reactions