[experiments] Daily Experiment Report — 2026-09-18 #61760
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #61982. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-18
51 experiments analysed across 48 workflows in
github/gh-aw. 24 have reached min_samples for at least descriptive comparison. Core decisions: ✅ PROMOTE 0 · 🟡 EXTEND 37 · ❌ REJECT 0 · ⚪ INCONCLUSIVE 14. No PROMOTE or REJECT decisions today — the majority remain in evidence-collection or unsupported-multi-variant status; no tracking issues are configured so no lifecycle labels or issue comments apply.⚡ Quick Stats
🟢 Experiments Ready for Analysis (24)
prompt_compression·agentperformanceanalyzercavemanverboseEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·agentpersonaexplorerbatchper_scenarioEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_variant·awfailureinvestigatorassertiveclinicalnarrativeEvidence: analysis_type=
n/a, metric=n/aCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstone_variant·breakingchangecheckerneutralurgentEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·cicoachconcisedetailedGuardrails:
run_success_rate>=0.85: UNSUPPORTED;empty_output_rate<=0.05: UNSUPPORTEDEvidence: analysis_type=
mann_whitney, metric=token_count_per_runCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationsub_agent_strategy·dailyagentrxtraceoptimizersingle_agentsub_agentsEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·dailyastrostylelitemarkdownspellcheckconcisedetailedEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·dailycommunityattributionconciseverboseEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·dailyfactmulti_candidatesingle_passGuardrails:
empty_output_rate<0.05: UNSUPPORTED;run_success_rate>=0.95: UNSUPPORTEDEvidence: analysis_type=
n/a, metric=discussion_reaction_countCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "empty_output_rate" is not backed by a supported metric observationprompt_style·dailynewsconcisedetailedEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonreasoning_depth·dailysecurityredteamiterativesingle_passGuardrails:
run_success_rate>=0.9: UNSUPPORTED;empty_output_rate<=0.05: UNSUPPORTEDEvidence: analysis_type=
proportion_test, metric=false_positive_rateCore decision: 🟡 EXTEND (
guardrail_unsupported) — mandatory guardrail "run_success_rate" is not backed by a supported metric observationoutput_format·deepreportannotated_briefexecutive_brieffull_briefingsteEvidence: analysis_type=
n/a, metric=n/aCore decision: ⚪ INCONCLUSIVE (
unsupported_multi_variant) — automatic decisions currently require exactly two variantstool_verbosity·gpcleanfull_bashminimal_toolsetEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonprompt_style·issuearboristconcisedetailedEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·smokeantigravitysingle_agentsub_agentsEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilotnoyesEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilotlargesmallEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilotaoaiapikeynoyesEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilotaoaiapikeylargesmallEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisoncaveman·smokecopilotaoaientranoyesEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsubagent_model·smokecopilotaoaientralargesmallEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_strategy·smokegeminisingle_agentsub_agentsEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonsub_agent_decomposition·smokepiparallel_sub_agentssingle_agentEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisontone_style·typistconversationalformalEvidence: analysis_type=
n/a, metric=n/aCore decision: 🟡 EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparison🟡 Experiments Still Collecting Samples (27)
View all COLLECTING experiments
sub_agent_strategyarchitectureguardiansingle_agent████████░░ 17/20 (85%);sub_agents██████████ 21/20 (100%)insufficient_samples)audit_decompositionauditworkflowsphased_sub_agents██░░░░░░░░ 41/264 (16%);single_agent█░░░░░░░░░ 34/264 (13%)insufficient_samples)prompt_styleblogauditorconcise██░░░░░░░░ 3/20 (15%);detailed████░░░░░░ 7/20 (35%)insufficient_samples)output_formatcopilotagentanalysisprose██████████ 34/30 (100%);ste█░░░░░░░░░ 3/30 (10%);structured██████████ 43/30 (100%)unsupported_multi_variant)detail_leveldailyarchitecturediagrambrief██████░░░░ 11/20 (55%);comprehensive███░░░░░░░ 6/20 (30%)insufficient_samples)model_sizedailycachestrategyanalyzergpt-5.3-codex██░░░░░░░░ 5/20 (25%);gpt-5.3-codex-spark████░░░░░░ 7/20 (35%)insufficient_samples)model_sizedailycavemanoptimizeragent░░░░░░░░░░ 1/20 (5%);claude-haiku-4.5██████████ 63/20 (100%);claude-sonnet-4.6██████████ 34/20 (100%);claude-sonnet-5███░░░░░░░ 6/20 (30%);small-agent█░░░░░░░░░ 2/20 (10%)unsupported_multi_variant)output_formatdailycodemetricsexecutive_summary██████████ 51/20 (100%);full_detail██████████ 57/20 (100%);ste███████░░░ 14/20 (70%)unsupported_multi_variant)output_formatdailycompilerqualityconcise██████████ 52/20 (100%);detailed██████████ 59/20 (100%);ste██████░░░░ 13/20 (65%)unsupported_multi_variant)model_sizedailydochealeragent█░░░░░░░░░ 2/20 (10%);claude-haiku-4.5██████████ 55/20 (100%);claude-sonnet-4.6██████████ 43/20 (100%);claude-sonnet-5██░░░░░░░░ 5/20 (25%);small-agent░░░░░░░░░░ 1/20 (5%)unsupported_multi_variant)model_sizedailydocupdateragent░░░░░░░░░░ 1/20 (5%);claude-haiku-4.5██████████ 44/20 (100%);claude-sonnet-4.6██████████ 47/20 (100%);claude-sonnet-5████░░░░░░ 9/20 (45%);small-agent█░░░░░░░░░ 2/20 (10%)unsupported_multi_variant)model_sizedailyfunctionnamerinsufficient_observations)output_formatdailyissuesreportcollapsible██████████ 58/20 (100%);inline██████████ 60/20 (100%);ste████████░░ 17/20 (85%)unsupported_multi_variant)remove_redundant_context_v1dailyrenderingscriptsverifierinsufficient_observations)log_fetch_strategydailysafeoutputoptimizerinsufficient_observations)semgrep_output_formatdailysemgrepscanbullet_list█████████░ 28/30 (93%);prose█████████░ 27/30 (90%);structured_sections███████░░░ 22/30 (73%)unsupported_multi_variant)timeout_settingdailysubagentoptimizerdefault█░░░░░░░░░ 2/20 (10%);relaxed░░░░░░░░░░ 1/20 (5%);tight░░░░░░░░░░ 1/20 (5%)unsupported_multi_variant)caveman_modedataflowprdiscussiondatasetno███░░░░░░░ 6/20 (30%);yes████░░░░░░ 9/20 (45%)insufficient_samples)summary_detaildependabotcampaignbrief███░░░░░░░ 6/20 (30%);detailed███░░░░░░░ 6/20 (30%)insufficient_samples)prompt_styledependabotgocheckerconcise█████████░ 18/20 (90%);detailed████████░░ 17/20 (85%);step_by_step██████░░░░ 13/20 (65%)unsupported_multi_variant)reasoning_depthplanbaseline██░░░░░░░░ 4/20 (20%);deep█░░░░░░░░░ 2/20 (10%);shallow░░░░░░░░░░ 1/20 (5%)unsupported_multi_variant)remove_redundant_context_v1prsouschefinsufficient_observations)sub_agent_strategysmokecopilotsubagentsdelegated_sequential█████░░░░░ 15/30 (50%);inline_strict████░░░░░░ 13/30 (43%);single_agent_control█████░░░░░ 16/30 (53%)unsupported_multi_variant)prompt_style_testsmokeprojectinsufficient_observations)sub_agent_strategysmoketemporaryidinsufficient_observations)model_sizetestqualitysentinelinsufficient_observations)prefetch_strategyweeklyblogpostwritereager██░░░░░░░░ 5/20 (25%);lazy████░░░░░░ 9/20 (45%)insufficient_samples)🔀 Interaction Diagnostics (multi-experiment workflows)
Three workflows (
smoke-copilot,smoke-copilot-aoai-apikey,smoke-copilot-aoai-entra) runcaveman×subagent_modelfactorial experiments simultaneously. All observed cells have n < min_samples (20), sointeraction_risk_status = SPARSE_CELL_RISKin every case. NoPROMOTEdecisions are affected today since all component decisions areEXTEND.🔀 Interaction diagnostics: `caveman` × `subagent_model` (smokecopilot)
nolargenosmallyeslargeinteraction_p_value: 0.1472
interaction_risk_status: SPARSE_CELL_RISK (all cells n < min_samples=20)
Since both
cavemanandsubagent_modeldecisions areEXTEND(notPROMOTE), no report-level promotion safety hold is triggered here — but interaction risk remainsSPARSE_CELL_RISKand must be re-checked before any futurePROMOTE.🔀 Interaction diagnostics: `caveman` × `subagent_model` (smokecopilotaoaiapikey)
nolargenosmallyeslargeyessmallinteraction_p_value: 1.0000
interaction_risk_status: SPARSE_CELL_RISK (all cells n < min_samples=20)
Since both
cavemanandsubagent_modeldecisions areEXTEND(notPROMOTE), no report-level promotion safety hold is triggered here — but interaction risk remainsSPARSE_CELL_RISKand must be re-checked before any futurePROMOTE.🔀 Interaction diagnostics: `caveman` × `subagent_model` (smokecopilotaoaientra)
nolargenosmallyeslargeyessmallinteraction_p_value: 1.0000
interaction_risk_status: SPARSE_CELL_RISK (all cells n < min_samples=20)
Since both
cavemanandsubagent_modeldecisions areEXTEND(notPROMOTE), no report-level promotion safety hold is triggered here — but interaction risk remainsSPARSE_CELL_RISKand must be re-checked before any futurePROMOTE.📊 Summary
View Full Experiments Table (51 experiments)
🧭 Self-Tuning Continuation Plan
🔧 Top 3 experiment actions
output_formatindeep-report.md—INCONCLUSIVE(unsupported_multi_variant, 4 variants). Reduce to a 2-variant test (e.g.executive_briefvsste) so the core decision layer can produce a deterministic PROMOTE/REJECT.prompt_styleinci-coach.md—EXTEND(guardrail_unsupported:run_success_rate,empty_output_rate). Wire these to a supported native observation so the guardrail gate can pass/fail instead of blocking every cycle.caveman/subagent_modelinsmoke-copilot*.md(3 workflows) — allEXTENDwithSPARSE_CELL_RISKinteraction diagnostics. Increase weekly run volume or consolidate the near-duplicate smoke workflows so factorial cells reach ≥20 samples faster.📋 Top 3 eval actions
labels-applied(deep-report.md) — 0% pass rate (0/115). Tighten the prompt's explicit instruction to attachcode-quality,automation,task-mininglabels when creating issues.output_format_goal_met(deep-report.md) — 25% pass rate (3/12). Reports don't reliably match their assignedoutput_formatvariant style (especiallyste). Add per-variant style examples/checklists.tool_verbosity_goal_met(gpclean.md) — 0% pass rate (0/3),issue-or-noopat 58% (18/31). Clarifyminimal_toolsetvsfull_bashinstructions and issue-vs-noop decision criteria.🕳️ Decision-pipeline gaps for next PR
guardrail_unsupportedblocks 3 experiments (ci-coach,daily-fact,daily-security-red-team) —run_success_rate/empty_output_rate/false_positive_rateneed a native-metric resolver.unsupported_multi_variantblocks 14 experiments (deep-report, daily-code-metrics, daily-compiler-quality, daily-doc-healer, daily-doc-updater, daily-issues-report, daily-caveman-optimizer, daily-semgrep-scan, daily-subagent-optimizer, dependabot-go-checker, plan, smoke-copilot-sub-agents) — core needs a normalized K≥3 variant decision path.SPARSE_CELL_RISK) in 3 workflows is a reporting-only diagnostic today; it should become a normalized core analysis signal to natively gate promotion.All reactions