Overview
v1.1 re-trains only DPO, on v1.0's preference pairs, changing only the learning rate and the number of epochs, to find what caused v1.0's DPO results. Everything else (β, LoRA, NF4, effective batch 16, seed, evaluation protocol) is v1.0's.
Evaluation results (HumanEval pass@1, EvalPlus chat protocol, greedy, 164 problems)
| DPO run | lr × epochs | pass@1 (95% CI) | Δ vs. SFT (McNemar p) | Mean CC |
|---|---|---|---|---|
| SFT (v1.0), reference | — | 70.7% [63.4, 77.2] | — | 3.54 |
| Execution-only (v1.0) | 2e-4 × 3 | 53.0% [45.4, 60.5] | −17.7 (<0.001) | 3.28 |
| Execution-only | 2e-4 × 1 | 74.4% [67.2, 80.5] | +3.7 (0.44) | 4.09 |
| Execution-only | 5e-6 × 3 | 69.5% [62.1, 76.0] | −1.2 (0.85) | 3.16 |
| Execution-only | 5e-6 × 1 | 68.9% [61.5, 75.5] | −1.8 (0.69) | 3.15 |
| Composite, size-matched | 5e-6 × 1 | 69.5% [62.1, 76.0] | −1.2 (0.84) | 3.14 |
| Composite (30,628 pairs) | 5e-6 × 1 | 68.3% [60.8, 74.9] | −2.4 (0.60) | 3.12 |
Key findings
- v1.0's DPO collapse came from training three epochs at lr 2e-4. One epoch at the same rate scores 74.4%, 21.3 points above three (p < 0.001). At lr 5e-6 the epoch count makes no difference.
- No v1.1 run is distinguishable from SFT (McNemar p 0.44–0.85). Trained gently, DPO on these "runs without errors" pairs neither costs nor measurably adds correctness.
- With a gentle recipe the reward makes no detectable difference, in pass@1 or in complexity (composite vs. execution-only CC −0.02 [−0.20, +0.21]).
- The pairs favour no-op code, rarely: 3.0% of chosen vs. 1.3% of rejected Case A candidates are mostly comments, define nothing, or gut the file they edit (p < 10⁻⁷). Over-optimization amplifies it: no-op HumanEval solutions go 0 → 5 → 34 as training gets more aggressive.
What's new in the code
dpo_trainer.py:--learning_rate,--num_train_epochs,--dataset,--run_name(defaults reproduce v1.0); each v1.1 run saves arun_config.json.scripts/audit_preference_pairs.py,scripts/filter_noop_pairs.py,src/preference_generation/noop.py: no-op audit and filter for preference pairs.scripts/eval_checkpoint.sh,src/evaluation/per_problem.py: single-adapter evaluation with v1.0's flags, and per-problem results for paired tests.src/notebooks/06_v1.1_dpo_recipe.ipynb: the full v1.1 analysis.ROADMAP.md: the plan for v2.x–v3.0.
Full results and caveats: README, Results (v1.1). Curated dataset (unchanged): Sergasgr/codealign-commitpackft.