Repository navigation
Releases: Sergasgr/CodeAlign
Release list
CodeAlign v1.1
Overview
v1.1 re-trains only DPO, on v1.0's preference pairs, changing only the learning rate and the number of epochs, to find what caused v1.0's DPO results. Everything else (β, LoRA, NF4, effective batch 16, seed, evaluation protocol) is v1.0's.
Evaluation results (HumanEval pass@1, EvalPlus chat protocol, greedy, 164 problems)
| DPO run | lr × epochs | pass@1 (95% CI) | Δ vs. SFT (McNemar p) | Mean CC |
|---|---|---|---|---|
| SFT (v1.0), reference | — | 70.7% [63.4, 77.2] | — | 3.54 |
| Execution-only (v1.0) | 2e-4 × 3 | 53.0% [45.4, 60.5] | −17.7 (<0.001) | 3.28 |
| Execution-only | 2e-4 × 1 | 74.4% [67.2, 80.5] | +3.7 (0.44) | 4.09 |
| Execution-only | 5e-6 × 3 | 69.5% [62.1, 76.0] | −1.2 (0.85) | 3.16 |
| Execution-only | 5e-6 × 1 | 68.9% [61.5, 75.5] | −1.8 (0.69) | 3.15 |
| Composite, size-matched | 5e-6 × 1 | 69.5% [62.1, 76.0] | −1.2 (0.84) | 3.14 |
| Composite (30,628 pairs) | 5e-6 × 1 | 68.3% [60.8, 74.9] | −2.4 (0.60) | 3.12 |
Key findings
- v1.0's DPO collapse came from training three epochs at lr 2e-4. One epoch at the same rate scores 74.4%, 21.3 points above three (p < 0.001). At lr 5e-6 the epoch count makes no difference.
- No v1.1 run is distinguishable from SFT (McNemar p 0.44–0.85). Trained gently, DPO on these "runs without errors" pairs neither costs nor measurably adds correctness.
- With a gentle recipe the reward makes no detectable difference, in pass@1 or in complexity (composite vs. execution-only CC −0.02 [−0.20, +0.21]).
- The pairs favour no-op code, rarely: 3.0% of chosen vs. 1.3% of rejected Case A candidates are mostly comments, define nothing, or gut the file they edit (p < 10⁻⁷). Over-optimization amplifies it: no-op HumanEval solutions go 0 → 5 → 34 as training gets more aggressive.
What's new in the code
dpo_trainer.py:--learning_rate,--num_train_epochs,--dataset,--run_name(defaults reproduce v1.0); each v1.1 run saves arun_config.json.scripts/audit_preference_pairs.py,scripts/filter_noop_pairs.py,src/preference_generation/noop.py: no-op audit and filter for preference pairs.scripts/eval_checkpoint.sh,src/evaluation/per_problem.py: single-adapter evaluation with v1.0's flags, and per-problem results for paired tests.src/notebooks/06_v1.1_dpo_recipe.ipynb: the full v1.1 analysis.ROADMAP.md: the plan for v2.x–v3.0.
Full results and caveats: README, Results (v1.1). Curated dataset (unchanged): Sergasgr/codealign-commitpackft.
CodeAlign v1.0
First complete run of the pipeline: CommitPackFT curation (122,066 samples, 8 languages) → QLoRA SFT → preference pairs labelled by a sandbox and static analysis → DPO with a composite reward, an execution-only reward and a size-matched control → HumanEval with EvalPlus's chat protocol.
| Model | HumanEval pass@1 (95% CI) | Mean CC |
|---|---|---|
Base (Qwen2.5-Coder-7B-Instruct) |
90.2% [84.7, 93.9] | 3.70 |
| SFT | 70.7% [63.4, 77.2] | 3.54 |
| DPO, execution-only reward (3,267 pairs) | 53.0% [45.4, 60.5] | 3.28 |
| DPO, composite reward, size-matched (3,267 pairs) | 45.1% [37.7, 52.8] | 2.03 |
| DPO, composite reward (30,628 pairs) | 24.4% [18.5, 31.5] | 1.42 |
- The evaluation reproduces the published baseline: 90.2% against the 88.4% in the Qwen2.5-Coder technical report.
- SFT and every DPO run lowered pass@1. DPO reused SFT's learning rate (2e-4) and reached a reward accuracy of 1.00.
- At a matched pair count, the composite reward lowered cyclomatic complexity (−1.25) with no detectable pass@1 difference.
- Both rewards were gamed: "runs without errors" is satisfied by commented-out code.
Protocol, controls and limitations: README → Results (v1.0). Curated dataset: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft