Skip to content

CodeAlign v1.1

Latest

Choose a tag to compare

@Sergasgr Sergasgr released this 04 Oct 15:50

Overview

v1.1 re-trains only DPO, on v1.0's preference pairs, changing only the learning rate and the number of epochs, to find what caused v1.0's DPO results. Everything else (β, LoRA, NF4, effective batch 16, seed, evaluation protocol) is v1.0's.

Evaluation results (HumanEval pass@1, EvalPlus chat protocol, greedy, 164 problems)

DPO run lr × epochs pass@1 (95% CI) Δ vs. SFT (McNemar p) Mean CC
SFT (v1.0), reference — 70.7% [63.4, 77.2] — 3.54
Execution-only (v1.0) 2e-4 × 3 53.0% [45.4, 60.5] −17.7 (<0.001) 3.28
Execution-only 2e-4 × 1 74.4% [67.2, 80.5] +3.7 (0.44) 4.09
Execution-only 5e-6 × 3 69.5% [62.1, 76.0] −1.2 (0.85) 3.16
Execution-only 5e-6 × 1 68.9% [61.5, 75.5] −1.8 (0.69) 3.15
Composite, size-matched 5e-6 × 1 69.5% [62.1, 76.0] −1.2 (0.84) 3.14
Composite (30,628 pairs) 5e-6 × 1 68.3% [60.8, 74.9] −2.4 (0.60) 3.12

Key findings

  • v1.0's DPO collapse came from training three epochs at lr 2e-4. One epoch at the same rate scores 74.4%, 21.3 points above three (p < 0.001). At lr 5e-6 the epoch count makes no difference.
  • No v1.1 run is distinguishable from SFT (McNemar p 0.44–0.85). Trained gently, DPO on these "runs without errors" pairs neither costs nor measurably adds correctness.
  • With a gentle recipe the reward makes no detectable difference, in pass@1 or in complexity (composite vs. execution-only CC −0.02 [−0.20, +0.21]).
  • The pairs favour no-op code, rarely: 3.0% of chosen vs. 1.3% of rejected Case A candidates are mostly comments, define nothing, or gut the file they edit (p < 10⁻⁷). Over-optimization amplifies it: no-op HumanEval solutions go 0 → 5 → 34 as training gets more aggressive.

What's new in the code

  • dpo_trainer.py: --learning_rate, --num_train_epochs, --dataset, --run_name (defaults reproduce v1.0); each v1.1 run saves a run_config.json.
  • scripts/audit_preference_pairs.py, scripts/filter_noop_pairs.py, src/preference_generation/noop.py: no-op audit and filter for preference pairs.
  • scripts/eval_checkpoint.sh, src/evaluation/per_problem.py: single-adapter evaluation with v1.0's flags, and per-problem results for paired tests.
  • src/notebooks/06_v1.1_dpo_recipe.ipynb: the full v1.1 analysis.
  • ROADMAP.md: the plan for v2.x–v3.0.

Full results and caveats: README, Results (v1.1). Curated dataset (unchanged): Sergasgr/codealign-commitpackft.