[experiments] Daily Experiment Report — 2026-09-15 #61070
Closed
Replies: 1 comment
|
This discussion has been marked as outdated by daily-experiment-report. A newer discussion is available at Discussion #61311. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-09-15
51 experiments analysed across 48 workflows. 24 have reached min_samples and are core-analysis ready; none show a deterministic PROMOTE/REJECT outcome yet — all resolved analyses land on EXTEND (guardrail/observation data gaps) or INCONCLUSIVE (3+ variant experiments, which the core decision engine does not yet auto-resolve). One 2-variant interaction pair (
smokecopilotcaveman × subagent_model) is sparse and held from promotion consideration.⚡ Quick Stats
caveman·smokecopilot📈 View Detailed Statistics
Sample Sizes & Progress
noyesCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparisonReport action: EXTEND (
interaction_underpowered) — Interaction cells forcaveman×subagent_modelare sparse (min cell n=1, target min_samples=20); interaction_p_value=0.206, interaction_risk_status=SPARSE_CELL_RISK. This is a reporting safety hold pending normalized interaction analysis in core.prompt_style·dailycommunityattribution📈 View Detailed Statistics
Sample Sizes & Progress
conciseverboseCore decision: EXTEND (
insufficient_observations) — primary metric observations are not available for statistical comparison[Repeat pattern above applies to all remaining experiments — summarized in the compact table below]
📊 Summary
View Full Experiments Table
tone_variantcavemansubagent_modelsub_agent_strategysub_agent_strategyoutput_formatcavemansubagent_modelcavemansubagent_modelprompt_styleoutput_formatprompt_stylereasoning_depthoutput_formatoutput_formatprompt_compressionsub_agent_strategytool_verbositymodel_sizeprompt_stylemodel_sizemodel_sizereasoning_depthprompt_styleoutput_formattone_stylesemgrep_output_formatsub_agent_strategysub_agent_decompositionaudit_decompositionprompt_styleprompt_stylesub_agent_strategytone_variantsub_agent_strategydetail_levelprefetch_strategycaveman_modesummary_detailprompt_stylemodel_sizereasoning_depthtimeout_settingprompt_style_testremove_redundant_context_v1log_fetch_strategysub_agent_strategyremove_redundant_context_v1model_sizemodel_size🔧 Self-Tuning Continuation Plan
Top 3 experiment actions
dailysecurityredteam/reasoning_depth(READY, 126 runs) — guardrailsrun_success_rateandempty_output_rateareguardrail_unsupported; wire native metric emission so the guardrail-backed PROMOTE/REJECT path can activate (10/10 sampled runs werecancelled, which itself warrants investigation into scheduling/dedup behavior independent of the experiment).awfailureinvestigator/tone_variantanddeepreport/output_format(both READY, 3–4 variants) — core decision engine currently reportsINCONCLUSIVE/unsupported_multi_variantfor any experiment with more than two variants. Prioritize extending core analysis to K-variant (ANOVA/Kruskal-Wallis style) decisions so these long-running, well-powered experiments (418 and 189 runs respectively) can resolve.smokecopilotcaveman × subagent_model factorial pair — interaction cells are sparse (min cell n=1 vs min_samples=20); continue collecting joint assignments before considering promotion of either factor in isolation.Top 3 eval actions
dailyfact,cicoach,dailysecurityredteam) declare guardrails (empty_output_rate,run_success_rate) that areunsupported— add graders/evals that emit these as native metric observations sodecision_guardrailscan move fromconfigured=Trueto a pass/fail verdict.insufficient_observationsas their reason code (33 of 51) indicate the primarymetric:field (e.g.discussion_engagement_score,effective_tokens,token_count_per_run) is not yet populated by an eval/grader pipeline for these workflows — tighten eval question mappings so primary-metric observations flow intogh aw experiments analyzefor READY experiments first (see list above).output_format/semgrep_output_formatstyle experiments (deepreport,dailycodemetrics,dailyissuesreport,dailysemgrepscan,dailycompilerquality), add secondary-metric evals foroutput_length_chars/output_token_countso verbosity trade-offs are visible even while the primary K-variant decision remains INCONCLUSIVE.Decision-pipeline gaps for the next PR
guardrail_unsupported(3 experiments) andinsufficient_observations(33 experiments) are the two dominant reason codes blocking every READY experiment from a deterministic PROMOTE/REJECT — this is a metric-plumbing gap, not a statistical-power gap, since most of these experiments already exceedmin_samples.unsupported_multi_variant(14 experiments, including two of the largest by sample size) is a core-analysis gap: extend the decision engine beyond two-variant comparisons.interaction_underpoweredcan become a first-classreason_coderather than a report-only override.All reactions