CodeAlign v1.0
First complete run of the pipeline: CommitPackFT curation (122,066 samples, 8 languages) → QLoRA SFT → preference pairs labelled by a sandbox and static analysis → DPO with a composite reward, an execution-only reward and a size-matched control → HumanEval with EvalPlus's chat protocol.
| Model | HumanEval pass@1 (95% CI) | Mean CC |
|---|---|---|
Base (Qwen2.5-Coder-7B-Instruct) |
90.2% [84.7, 93.9] | 3.70 |
| SFT | 70.7% [63.4, 77.2] | 3.54 |
| DPO, execution-only reward (3,267 pairs) | 53.0% [45.4, 60.5] | 3.28 |
| DPO, composite reward, size-matched (3,267 pairs) | 45.1% [37.7, 52.8] | 2.03 |
| DPO, composite reward (30,628 pairs) | 24.4% [18.5, 31.5] | 1.42 |
- The evaluation reproduces the published baseline: 90.2% against the 88.4% in the Qwen2.5-Coder technical report.
- SFT and every DPO run lowered pass@1. DPO reused SFT's learning rate (2e-4) and reached a reward accuracy of 1.00.
- At a matched pair count, the composite reward lowered cyclomatic complexity (−1.25) with no detectable pass@1 difference.
- Both rewards were gamed: "runs without errors" is satisfied by commented-out code.
Protocol, controls and limitations: README → Results (v1.0). Curated dataset: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft