[experiments] Daily Experiment Report — 2026-09-28 #63963
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #64219. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-28
48 workflows in
github/gh-awcurrently declareexperiments:sections, running 51 active A/B experiments. 26 experiments have reachedREADYsample thresholds, but none have aPROMOTEorREJECTcore decision yet — 37 are held atEXTEND(mostlyinsufficient_observations/guardrail_unsupported, since native metric/guardrail plumbing isn't wired up for most experiments) and 14 multi-variant experiments (K≥3) returnINCONCLUSIVE(unsupported_multi_variant) until core analysis supports more than two variants. Thesmoke-copilot*family (3 workflows) runs two crossed experiments (caveman×subagent_model) with all interaction cells sparse (n<20 in every cell), so a cross-experiment safety hold is documented below even though no individual core decision is currentlyPROMOTE.⚡ Quick Stats
🔎 All 26 READY experiments with a tracking issue (detailed statistics)
prompt_compression·agent-performance-analyzer.md📈 View Detailed Statistics
Sample Sizes & Progress
cavemanverboseCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_variant·aw-failure-investigator.md📈 View Detailed Statistics
Sample Sizes & Progress
assertiveclinicalnarrativeCore decision: INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstone_variant·breaking-change-checker.md📈 View Detailed Statistics
Sample Sizes & Progress
neutralurgentCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·ci-coach.md📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationreasoning_depth·daily-fact.md📈 View Detailed Statistics
Sample Sizes & Progress
multi_candidatesingle_passCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "empty_output_rate" is not backed by a supported metric observationoutput_format·daily-issues-report.md📈 View Detailed Statistics
Sample Sizes & Progress
collapsibleinlinesteCore decision: INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantsprompt_style·daily-news.md📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·daily-security-red-team.md📈 View Detailed Statistics
Sample Sizes & Progress
iterativesingle_passCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationprompt_style·issue-arborist.md📈 View Detailed Statistics
Sample Sizes & Progress
concisedetailedCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonmodel_size·test-quality-sentinel.md📈 View Detailed Statistics
Sample Sizes & Progress
claude-haiku-4.5claude-sonnet-5Core decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_style·typist.md📈 View Detailed Statistics
Sample Sizes & Progress
conversationalformalCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman×subagent_modelinteraction ·smoke-copilotfamily (3 workflows)Each workflow crosses two active experiments (
caveman: yes/no,subagent_model: small/large). Individual per-experiment core decisions are allEXTEND(insufficient_observations), but interaction cell diagnostics are computed below as a reporting safety check.📈 View Interaction Cell Diagnostics
smoke-copilot
interaction_p_value= 0.270 ·interaction_risk_status= SPARSE_CELL_RISK (all cells n<20)smoke-copilot-aoai-apikey
interaction_p_value= N/A (degenerate table) ·interaction_risk_status= SPARSE_CELL_RISKsmoke-copilot-aoai-entra
interaction_p_value= N/A (degenerate table) ·interaction_risk_status= SPARSE_CELL_RISKSince no individual core decision is currently
PROMOTE, this hold is informational only for now — but the pattern is recorded so that if any experiment in this family later reachesPROMOTE, the sparse interaction cells will forcereport_action: EXTEND(interaction_underpowered) instead of an outright promotion recommendation.📊 Summary
View Full Experiments Table (all 51 experiments)
cavemanverbosebatchper_scenariosingle_agentsub_agentssingle_agentphased_sub_agentsassertiveclinicalconcisedetailedneutralurgentdetailedconciseprosestesingle_agentsub_agentsbriefcomprehensiveconcisedetailedgpt-5.3-codexgpt-5.3-codex-sparkagentclaude-haiku-4.5executive_summaryfull_detailconciseverboseconcisedetailedagentclaude-haiku-4.5agentclaude-haiku-4.5single_passmulti_candidaten/an/acollapsibleinlineconcisedetailedn/an/aeagerlazysingle_passiterativebullet_listprosedefaultrelaxednoyesannotated_briefexecutive_briefbriefdetailedconcisedetailedfull_bashminimal_toolsetconcisedetailedbaselinedeepn/an/asingle_agentsub_agentsnoyeslargesmallnoyeslargesmallnoyeslargesmalldelegated_sequentialinline_strictsingle_agentsub_agentsparallel_sub_agentssingle_agentn/an/asingle_agentsub_agentsclaude-haiku-4.5claude-sonnet-5conversationalformaleagerlazy🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
dailyissuesreport/deepreport/dailycodemetrics/dailycompilerquality/dailydochealer/dailydocupdater/dailycavemanoptimizer/dailysemgrepscan/dependabotgochecker/awfailureinvestigator/plan/smokecopilotsubagents/dailysubagentoptimizer/copilotagentanalysis(14 total,unsupported_multi_variant) — these experiments declare 3+ variants (e.g.output_format: [collapsible, inline, ste],model_size: [agent, claude-haiku-4.5, ...]). Core analysis currently only supports pairwise (K=2) decisions. Action: extend the core decision engine to support K≥3 comparisons (omnibus test + pairwise post-hoc with Bonferroni correction, already partially scaffolded viabonferroni_alpha), or split these into paired sub-experiments.cicoach,dailyfact,dailysecurityredteam(guardrail_unsupported) — mandatory guardrail metrics (run_success_rate,empty_output_rate) are declared in frontmatter but have no supported native observation source. Action: wire these guardrail names to the existingrun_success_rate/output-emptiness detectors sodecision_guardrails.passedcan be computed instead of blocking onEXTEND.insufficient_observationsexperiments (majority of the READY set) — primarymetric:fields (e.g.token_count_per_run,aic,discussion_engagement_score) aren't backed by a supported observation pipeline yet, even though sample counts are well pastmin_samples. Action: prioritize native metric collection for the highest-volume workflows first (smoke-copilot,smoke-gemini,smoke-copilot-aoai-*,dailyissuesreport) since they already have the largest run histories and would unblock aPROMOTE/REJECTdecision fastest.Top 3 eval actions
output_formatvariants (dailyissuesreport,deepreport,dailycodemetrics,dailycompilerquality,copilotagentanalysis,dailysemgrepscan) declare"eval:output_format_goal_met"/"eval:output_format_adherence"secondary metrics — tighten these grader questions to score format compliance (e.g. STE readability, collapsible-section usage) numerically rather than pass/fail, so they can feed the core decision engine as a supported observation.copilot-pr-nlp-analysisdeclareseval:insights_report_producedandeval:pr_conversations_analyzed— these should be validated against the workflow's actual PR-conversation-analysis prompt section to ensure the grader checks the produced artifact, not just its presence.aic(ai_credits) as the sole primary metric (agent-persona-explorer,blog-auditor,daily-news,gpclean,smoke-temporary-id,smoke-gemini) should add a secondary quality-eval question so cost reductions aren't rewarded at the expense of output quality — currently these can only ever answer "is it cheaper," never "is it still good."Decision-pipeline gaps for the next PR
insufficient_observations(26 experiments) andguardrail_unsupported(3 experiments) together account for 100% of theEXTENDdecisions among READY experiments — this is the single largest lever to unlockPROMOTE/REJECToutcomes. The native-metric-to-core-analysis bridge (fortoken_count_per_run,aic,ai_credits_total,discussion_engagement_score,run_success_rate,empty_output_rate) is the top priority gap.unsupported_multi_variant(14 experiments, all K≥3) needs a normalized core analysis signal (omnibus + Bonferroni-corrected pairwise) before these experiments can ever resolve pastINCONCLUSIVE.caveman×subagent_modelin thesmoke-copilotfamily) are currently computed only in this report as a descriptive safety check. These should become a normalized core analysis signal (interaction_p_value,SPARSE_CELL_RISK) so that a futurePROMOTEdecision in either experiment is automatically held pending sufficient interaction cell samples, rather than relying on manual reporting-layer logic.All reactions