You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
48 experiment workflows declare experiments: in this repository; 47 have written state to their experiments/<id> branch, spanning 51 individual A/B experiments across model, prompt, output-format, and sub-agent-strategy variants. 25 experiments have reached min_samples for every variant (🟢 READY); 22 are still accumulating runs (🟡 COLLECTING); 4 declared-experiment workflows (dailyrenderingscriptsverifier, prsouschef, smokeproject, smoketemporaryid) have no state file yet — likely brand-new or not yet executed on their experiment branch.
⚠️Data-pipeline gap: gh aw experiments analyze could not run in this environment (outbound git access to github.com is proxy-blocked), so this report was built directly from each experiments/<id> branch's state.json (assignment counts only) fetched via the GitHub MCP server, plus each workflow's frontmatter (min_samples, issue, metric, hypothesis). No core decision/reason_code/readiness field, no per-run outcome (success rate, duration), and no guardrail pass/fail could be computed — those require the gh aw experiments analyze binary's decision engine and 30-run outcome correlation, which needs authenticated gh/git access not available here. All statuses below are descriptive readiness (n ≥ min_samples) and chi-square balance only, not statistical decisions. Treat this as a partial/interim report.
⚡ Quick Stats
Metric
Value
Workflows declaring experiments:
48
Workflows with state data
47
Total individual experiments
51
Reached min_samples (🟢 READY)
25
Still collecting (🟡 COLLECTING)
22
Missing state file
4
Core decisions computed
0 (blocked — see gap above)
📈 View Full Experiment Table (n, min_samples, balance p-value, status)
Balance p-value is a chi-square goodness-of-fit test against equal assignment across variants (higher = more balanced random assignment). It is not an outcome-significance test.
Note: daily-cache-strategy-analyzer, daily-caveman-optimizer, daily-doc-healer, daily-doc-updater show model_size variant sets that include stray legacy labels (small-agent, agent, single-digit counts) alongside the current model names — these look like leftover assignments from before a variant-set change and should be reviewed/pruned so balance tests aren't skewed.
🟢 Ready for Analysis (min_samples reached on all variants)
View 25 ready experiments with tracking issues
Experiment
Workflow
Tracking issue
prompt_compression
agent-performance-analyzer
#33280
tone_variant
aw-failure-investigator
#36105
prompt_style
ci-coach
#32335
output_format
daily-code-metrics
#1
output_format
daily-compiler-quality
#32390
output_format
daily-issues-report
#30573
prompt_style
daily-news
#31190
reasoning_depth
daily-security-red-team
#31673
model_size
test-quality-sentinel
#43530
tone_style
typist
#34032
The remaining 15 ready experiments (agent-persona-explorer, daily-agentrx-trace-optimizer, daily-astrostylelite-markdown-spellcheck, daily-community-attribution, deep-report, gpclean, smoke-antigravity, smoke-copilot ×2, smoke-copilot-aoai-apikey ×2, smoke-copilot-aoai-entra ×2, smoke-gemini, smoke-pi) have no issue: field configured, so no tracking-issue notification applies per Step 8 rules.
No core decisions (PROMOTE/EXTEND/REJECT/INCONCLUSIVE) are reported this run — the decision engine (gh aw experiments analyze) requires outbound git/API access that this environment's network policy blocks. Readiness above reflects sample-size and balance only.
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
Restore gh aw experiments analyze execution in the reporting environment (grant outbound git/gh access to github.com, or add an MCP-based equivalent) so core decision/reason_code/readiness fields can be computed instead of this descriptive-only fallback.
Investigate the 4 workflows with experiments: declared but no state.json on their branch (daily-rendering-scripts-verifier, pr-sous-chef, smoke-project, smoke-temporary-id) — confirm push_experiments_state job is running and succeeding for these.
Clean up stray legacy variant labels in model_size experiments (daily-cache-strategy-analyzer, daily-caveman-optimizer, daily-doc-healer, daily-doc-updater) so future balance/decision tests aren't computed over obsolete single-digit variant buckets.
Top 3 eval actions
Once decision-engine access is restored, prioritize evals for the 10 issue:-tracked ready experiments above (ci-coach, daily-security-red-team, test-quality-sentinel, etc.) since they already have a stakeholder-facing tracking issue awaiting a result.
Add native outcome-metric collection (success rate, duration, guardrail pass/fail) directly into state.json at the workflow level, rather than requiring a secondary 30-run GitHub Actions API correlation — this is the same "insufficient_observations" gap seen across nearly all experiments here.
Tighten grading for model_size experiments with >2 variants and imbalanced legacy buckets (daily-cache-strategy-analyzer chi-square p≈0.000) so pruning of stale variants happens automatically rather than via manual audit.
Decision-pipeline gaps for next PR
No experiment in this run produced evidence/decision_guardrails — this indicates guardrail and outcome-metric collection is not yet wired to a queryable source usable outside a full gh-authenticated environment; consider exposing state.json outcome fields (success/duration) directly so descriptive-only sandboxes can still compute meaningful guardrail pass/fail.
No interaction diagnostics were computed since no workflow currently runs 2+ simultaneous experiments observed together in this dataset; smoke-copilot and its AOAI variants run two experiments (caveman, subagent_model) per workflow and would be the first good candidates for factorial interaction-cell analysis once outcome correlation is available.
Descriptive window: full experiments/<id> branch state (assignment counts only) · Decision thresholds: not evaluated (see data-pipeline gap)
Run: 33850615177
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
github.com
To allow these domains, add them to the network.allowed list in your workflow frontmatter:
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-04
48 experiment workflows declare
experiments:in this repository; 47 have written state to theirexperiments/<id>branch, spanning 51 individual A/B experiments across model, prompt, output-format, and sub-agent-strategy variants. 25 experiments have reachedmin_samplesfor every variant (🟢 READY); 22 are still accumulating runs (🟡 COLLECTING); 4 declared-experiment workflows (dailyrenderingscriptsverifier,prsouschef,smokeproject,smoketemporaryid) have no state file yet — likely brand-new or not yet executed on their experiment branch.⚡ Quick Stats
experiments:min_samples(🟢 READY)📈 View Full Experiment Table (n, min_samples, balance p-value, status)
Balance p-value is a chi-square goodness-of-fit test against equal assignment across variants (higher = more balanced random assignment). It is not an outcome-significance test.
Note:
daily-cache-strategy-analyzer,daily-caveman-optimizer,daily-doc-healer,daily-doc-updatershowmodel_sizevariant sets that include stray legacy labels (small-agent,agent, single-digit counts) alongside the current model names — these look like leftover assignments from before a variant-set change and should be reviewed/pruned so balance tests aren't skewed.🟢 Ready for Analysis (min_samples reached on all variants)
View 25 ready experiments with tracking issues
prompt_compression#33280tone_variant#36105prompt_style#32335output_format#1output_format#32390output_format#30573prompt_style#31190reasoning_depth#31673model_size#43530tone_style#34032The remaining 15 ready experiments (
agent-persona-explorer,daily-agentrx-trace-optimizer,daily-astrostylelite-markdown-spellcheck,daily-community-attribution,deep-report,gpclean,smoke-antigravity,smoke-copilot×2,smoke-copilot-aoai-apikey×2,smoke-copilot-aoai-entra×2,smoke-gemini,smoke-pi) have noissue:field configured, so no tracking-issue notification applies per Step 8 rules.No core decisions (PROMOTE/EXTEND/REJECT/INCONCLUSIVE) are reported this run — the decision engine (
gh aw experiments analyze) requires outbound git/API access that this environment's network policy blocks. Readiness above reflects sample-size and balance only.🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
gh aw experiments analyzeexecution in the reporting environment (grant outboundgit/ghaccess togithub.com, or add an MCP-based equivalent) so coredecision/reason_code/readinessfields can be computed instead of this descriptive-only fallback.experiments:declared but no state.json on their branch (daily-rendering-scripts-verifier,pr-sous-chef,smoke-project,smoke-temporary-id) — confirmpush_experiments_statejob is running and succeeding for these.model_sizeexperiments (daily-cache-strategy-analyzer,daily-caveman-optimizer,daily-doc-healer,daily-doc-updater) so future balance/decision tests aren't computed over obsolete single-digit variant buckets.Top 3 eval actions
issue:-tracked ready experiments above (ci-coach,daily-security-red-team,test-quality-sentinel, etc.) since they already have a stakeholder-facing tracking issue awaiting a result.state.jsonat the workflow level, rather than requiring a secondary 30-run GitHub Actions API correlation — this is the same "insufficient_observations" gap seen across nearly all experiments here.model_sizeexperiments with >2 variants and imbalanced legacy buckets (daily-cache-strategy-analyzerchi-square p≈0.000) so pruning of stale variants happens automatically rather than via manual audit.Decision-pipeline gaps for next PR
evidence/decision_guardrails— this indicates guardrail and outcome-metric collection is not yet wired to a queryable source usable outside a fullgh-authenticated environment; consider exposingstate.jsonoutcome fields (success/duration) directly so descriptive-only sandboxes can still compute meaningful guardrail pass/fail.smoke-copilotand its AOAI variants run two experiments (caveman,subagent_model) per workflow and would be the first good candidates for factorial interaction-cell analysis once outcome correlation is available.Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
github.comTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions