[experiments] Daily Experiment Report — 2026-09-24 #63130
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #63394. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-24
51 experiments analysed across 48 workflow branches. 26 have reached
min_samples(READY); the remaining 25 are still collecting. No experiment has a deterministic PROMOTE or REJECT decision yet — every core decision isEXTEND(37) orINCONCLUSIVE(14), mostly because outcome-metric plumbing (guardrail_unsupported,insufficient_observations) or ≥3-variant support (unsupported_multi_variant) is not yet wired into the decision layer.⚡ Quick Stats
experiments:)prompt_style·ci-coach.mdH0: no change in PR merge rate or proposal quality. H1: concise prompt reduces token usage ≥25% without degrading proposal quality
📈 View Detailed Statistics
Sample Sizes & Progress
detailedconciseDescriptive outcomes (last 10 recent runs)
detailedconciseCore decision: EXTEND (
guardrail_unsupported) — mandatory guardrailrun_success_ratehas no supported metric observation yet; sample size is sufficient (94 runs) but the guardrail pipeline is the blocker.reasoning_depth·daily-fact.mdH0: no change in discussion engagement rate. H1: multi_candidate produces more novel/verified facts
conclusion=failurein the Actions API — 0% success in the descriptive sample despiteinsufficient_observations/guardrail_unsupportedmasking it in the core decision. This looks like an infrastructure or workflow regression independent of the A/B assignment and should be triaged before drawing any A/B conclusion.📈 View Detailed Statistics
Sample Sizes & Progress
single_passmulti_candidateDescriptive outcomes (last 10 recent runs)
single_passmulti_candidateCore decision: EXTEND (
guardrail_unsupported) — sample size sufficient (94 runs) but flagged: recent-run success rate is 0% for both variants, which likely reflects an infrastructure/workflow break, not an experiment signal.reasoning_depth·daily-security-red-team.mdH0: no change in finding quality. H1: iterative reduces false-positive rate
conclusion=cancelled. This is very likely a timeout/cancellation pattern (mean duration ≈3800s ≈63 min in both variants) unrelated to the reasoning-depth assignment, and should be investigated as an infra issue before any A/B read.📈 View Detailed Statistics
Sample Sizes & Progress
single_passiterativeDescriptive outcomes (last 10 recent runs — all cancelled)
single_passiterativeCore decision: EXTEND (
guardrail_unsupported) — flagged: all 10 recent runs cancelled at ~63 minutes in both variants, consistent with atimeout-minutesceiling being hit regardless of assignment. Recommend checking the workflow's timeout budget before continuing this experiment.🟢 Other READY experiments (2-variant, `insufficient_observations`)
These experiments have reached
min_samplesbut the decision layer still requires more matched outcome observations (native metric pipeline gaps), not more runs.EXTENDhere means "keep collecting outcome telemetry," not "keep running more assignments."agent-performance-analyzer.md#33280agent-persona-explorer.mdbreaking-change-checker.md#42467daily-agentrx-trace-optimizer.mddaily-astrostylelite-markdown-spellcheck.mddaily-community-attribution.mddaily-news.md#31190gpclean.mdissue-arborist.md#30015smoke-antigravity.md*smoke-copilot.mdsmoke-copilot.mdsmoke-copilot-aoai-apikey.mdsmoke-copilot-aoai-apikey.mdsmoke-copilot-aoai-entra.mdsmoke-copilot-aoai-entra.mdsmoke-gemini.mdsmoke-pi.md*test-quality-sentinel.md#43530typist.md#34032*= workflow no longer declaresexperiments:in frontmatter but the branch has residual history; treat as concluded/stale.🔀 Factorial interaction:
caveman×subagent_model(smoke-copilot family)Three workflows (
smoke-copilot.md,smoke-copilot-aoai-apikey.md,smoke-copilot-aoai-entra.md) run two simultaneous 2-variant experiments. Both individual experiments are coreEXTEND(insufficient_observations), but a 2×2 interaction diagnostic was computed from the last 10 matched runs per workflow:All 12 cells have
n < 20(min_samples).core_decision: EXTENDfor both underlying experiments in all three workflows, so the sparse-cell safety hold does not change anything today (it only downgrades a would-bePROMOTE), but is logged here since these are the only multi-experiment factorial workflows in the fleet.⚪ INCONCLUSIVE experiments (3–4 variants, `unsupported_multi_variant`)
The core decision engine currently only computes automatic PROMOTE/REJECT for exactly 2 variants. These 6 experiments have ≥3 variants and enough samples, but return
INCONCLUSIVE/unsupported_multi_variantby design — they need either a K-variant decision extension or manual analyst review.aw-failure-investigator.md#36105daily-issues-report.md#30573deep-report.mdcopilot-agent-analysis.mdtotal_runs≥mindaily-caveman-optimizer.mdagent/small-agentlabelsdaily-code-metrics.md#1(likely placeholder/typo — not a valid tracking issue)daily-compiler-quality.md#32390daily-doc-healer.md/daily-doc-updater.mddaily-semgrep-scan.md#32795dependabot-go-checker.mdplan.md#42941smoke-copilot-sub-agents.md#47551🟡 COLLECTING experiments (below `min_samples`)
audit-workflows.md#43177blog-auditor.mddaily-architecture-diagram.md#31926daily-cache-strategy-analyzer.mddaily-safe-output-optimizer.md#38094dataflow-pr-discussion-dataset.md#37102dependabot-campaign(removed workflow)daily-rendering-scripts-verifier.md/pr-sous-chef.mdweekly-blog-post-writer.md#38590daily-function-namer/architecture-guardian(removed workflows)📊 Summary
View Full Experiments Table (READY only)
🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
daily-fact.md(0/10 recent runs succeeded) anddaily-security-red-team.md(10/10 recent runs cancelled, ~63 min each) as likely infra/timeout regressions unrelated to their A/B assignments — file/triage before trusting either experiment's future decision.issue: 1ondaily-code-metrics.md'soutput_formatexperiment frontmatter (points to an unrelated closed issue) so lifecycle notifications route correctly.unsupported_multi_variantfor 6 workflows includingdeep-report.md,copilot-agent-analysis.md,daily-issues-report.md) — these have enough samples and, indaily-issues-report.md's case, a statistically significant assignment imbalance (p<0.001) worth a native K-variant test.Top 3 eval actions
copilot-agent-analysis.md,daily-code-metrics.md,daily-compiler-quality.md,deep-report.md, anddaily-issues-report.mdall use anste(Simplified Technical English)output_formatvariant that is consistently under-assigned relative to peers (e.g. 3/34/43 for copilot-agent-analysis) — tighten weighting or an eval question asking "was thestevariant actually selected and rendered distinctly fromstructured/prose?" to confirm the variant isn't silently degrading to a default template.daily-fact.mdanddaily-security-red-team.md, add an eval question mapped to their prompts asking "did the run complete without failure/cancellation?" so guardrail-blocking infra issues surface as a graded observation rather than only via manual Action-run inspection.daily-caveman-optimizer.md/daily-doc-healer.md/daily-doc-updater.mdcarry stale variant labels (agent,small-agent) alongside current ones (claude-sonnet-5,claude-haiku-4.5) — add an eval/grading check that flags variant-label drift whenexperiments:frontmatter is edited without a corresponding branch-state migration.Decision-pipeline gaps for the next PR
guardrail_unsupportedblocks 3 READY, high-sample experiments (ci-coach,daily-fact,daily-security-red-team) purely becauserun_success_rate/empty_output_ratehave no native metric observation wired — this is the single highest-leverage fix to unblock automatic decisions today.unsupported_multi_variantblocks 14 experiments (6 workflows) from ever reaching PROMOTE/REJECT; the 2-cell factorial interaction helper built for this report (§ smoke-copilot family) should become a normalized core analysis signal rather than a reporting-only diagnostic, especially once K-variant support lands.smoke-antigravity,dependabot-campaign,daily-subagent-optimizer,architecture-guardian,daily-function-namer,smoke-pi) have historical state onexperiments/*branches but no matchingexperiments:frontmatter in any current workflow file — these should be explicitly archived/pruned or their frontmatter restored, otherwisegh aw experiments listwill keep surfacing stale data indefinitely.Warning
Firewall blocked 7 domains
The following domains were blocked by the firewall during workflow execution:
files.pythonhosted.orggithub.como205451.ingest.us.sentry.ioproxy.golang.orgpypi.orgstorage.googleapis.comsum.golang.orgTo allow these domains, add them to the
network.allowedlist in your workflow frontmatter:See Network Configuration for more information.
All reactions