[experiments] Daily Experiment Report — 2026-09-25 #63394
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #63590. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Daily Experiment Report — 2026-09-25
Automated statistical summary of all active agentic workflow A/B experiments in
github/gh-aw. Decisions below are the canonical output ofgh aw experiments analyze(deterministic core logic) — descriptive success-rate/duration stats are supplementary context only and never override the core decision.Quick Stats
Decision-Pipeline Gaps (surfaced this run)
unsupported_multi_variant(14 experiments): every K≥3-variant experiment returnsINCONCLUSIVEbecause the core analysis engine does not yet support statistical comparison beyond pairwise (2-variant) tests. These cannot progress toPROMOTE/REJECTuntil multi-variant support (e.g. ANOVA / Kruskal-Wallis with Bonferroni-corrected pairwise follow-up) is added.guardrail_unsupported(3 experiments): guardrail metrics are configured but no per-run observation pipeline currently feeds them, so guardrail pass/fail cannot be evaluated end-to-end. All guardrails repo-wide currently reportstatus: unsupported.insufficient_observations(26 experiments): primary metric values are not yet being recorded per run for statistical comparison (assignment counts exist, but the metric itself is missing), blocking readiness advancement.gh aw experiments list/analyze --repo <owner/repo>fails becausepkg/cli/experiments_fetch.gocallsgh api ... --repo <repo>, andgh apihas no--repoflag. Today's report used a local git-branch fallback (git fetch origin 'refs/heads/experiments/*:...') instead of the documented remote path. Filing as a follow-up dev task.Per-Workflow Experiment Detail
Each workflow's active experiment(s) are shown below (compact ASCII comparison table, evidence, guardrails, and core decision).
agentperformanceanalyzer— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#33280``prompt_compression
agentpersonaexplorer— 1 experiment(s) — decision(s): EXTEND — readiness: READYsub_agent_strategy
architectureguardian— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGsub_agent_strategy
auditworkflows— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTING · tracking: ``#43177``audit_decomposition
awfailureinvestigator— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: READY · tracking: ``#36105``tone_variant
blogauditor— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGprompt_style
breakingchangechecker— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#42467``tone_variant
cicoach— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#32335``prompt_style
copilotagentanalysis— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGoutput_format
dailyagentrxtraceoptimizer— 1 experiment(s) — decision(s): EXTEND — readiness: READYsub_agent_strategy
dailyarchitecturediagram— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTING · tracking: ``#31926``detail_level
dailyastrostylelitemarkdownspellcheck— 1 experiment(s) — decision(s): EXTEND — readiness: READYprompt_style
dailycachestrategyanalyzer— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGmodel_size
dailycavemanoptimizer— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGmodel_size
dailycodemetrics— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTING · tracking: ``#1``output_format
dailycommunityattribution— 1 experiment(s) — decision(s): EXTEND — readiness: READYprompt_style
dailycompilerquality— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTING · tracking: ``#32390``output_format
dailydochealer— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGmodel_size
dailydocupdater— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGmodel_size
dailyfact— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#31324``reasoning_depth
dailyfunctionnamer— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGmodel_size
dailyissuesreport— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: READY · tracking: ``#30573``output_format
dailynews— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#31190``prompt_style
dailyrenderingscriptsverifier— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGremove_redundant_context_v1
dailysafeoutputoptimizer— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTING · tracking: ``#38094``log_fetch_strategy
dailysecurityredteam— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#31673``reasoning_depth
dailysemgrepscan— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTING · tracking: ``#32795``semgrep_output_format
dailysubagentoptimizer— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGtimeout_setting
dataflowprdiscussiondataset— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTING · tracking: ``#37102``caveman_mode
deepreport— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: READYoutput_format
dependabotcampaign— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGsummary_detail
dependabotgochecker— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTINGprompt_style
gpclean— 1 experiment(s) — decision(s): EXTEND — readiness: READYtool_verbosity
issuearborist— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#30015``prompt_style
plan— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTING · tracking: ``#42941``reasoning_depth
prsouschef— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGremove_redundant_context_v1
smokeantigravity— 1 experiment(s) — decision(s): EXTEND — readiness: READYsub_agent_strategy
smokecopilot— 2 experiment(s) — decision(s): EXTEND — readiness: READYcaveman
subagent_model
smokecopilotaoaiapikey— 2 experiment(s) — decision(s): EXTEND — readiness: READYcaveman
subagent_model
smokecopilotaoaientra— 2 experiment(s) — decision(s): EXTEND — readiness: READYcaveman
subagent_model
smokecopilotsubagents— 1 experiment(s) — decision(s): INCONCLUSIVE — readiness: COLLECTING · tracking: ``#47551``sub_agent_strategy
smokegemini— 1 experiment(s) — decision(s): EXTEND — readiness: READYsub_agent_strategy
smokepi— 1 experiment(s) — decision(s): EXTEND — readiness: READYsub_agent_decomposition
smokeproject— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGprompt_style_test
smoketemporaryid— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTINGsub_agent_strategy
testqualitysentinel— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#43530``model_size
typist— 1 experiment(s) — decision(s): EXTEND — readiness: READY · tracking: ``#34032``tone_style
weeklyblogpostwriter— 1 experiment(s) — decision(s): EXTEND — readiness: COLLECTING · tracking: ``#38590``prefetch_strategy
Factorial Interaction Diagnostics
Three workflows (
smokecopilot,smokecopilotaoaiapikey,smokecopilotaoaientra) run two concurrent experiments (caveman×subagent_model). Pairwise interaction cells were computed for each; all cells are sparse (n < min_samples=20 per cell), yieldinginteraction_risk_status: SPARSE_CELL_RISK. However, none of these experiments' core decisions arePROMOTE, so the interaction safety hold (which only applies to aPROMOTEcore decision) does not trigger for any of them today —report_actionequalscore_decision(EXTEND) in all three cases.Self-Tuning Continuation Plan
Top experiment actions:
unsupported_multi_variant(K≥3) experiments until multi-variant statistical comparison lands in core; consider splitting long-running 3+-variant experiments into pairwise A/B pairs to unblock decisions sooner.insufficient_observationsexperiments — assignment counts are recorded but the metric values needed for comparison are not, which is the single largest blocker to readiness across the fleet.PROMOTE/REJECTand none newly reached READY today — continue monitoring.Top eval/pipeline actions:
guardrail_unsupportedexperiments (and all guardrail-configured experiments generally) can produce real pass/fail guardrail results instead ofstatus: unsupported.gh aw experiments list/analyze --repoCLI bug (gh api ... --repo <repo>is an invalid flag combination inpkg/cli/experiments_fetch.go) so future runs of this report don't need the local git-branch fallback.hypothesis/metric/secondary_metricsfrontmatter for the ~27 experiments currently missing an explicit hypothesis, to make the ASCII tables above more informative.Decision-pipeline gaps:
unsupported_multi_variant(14),guardrail_unsupported(3, and universally unsupported across all guardrail-configured experiments),insufficient_observations(26), and the--repo/gh apiCLI flag bug noted above.Generated 2026-09-25T08:12:56.389082Z by the daily-experiment-report workflow.
All reactions