Skip to content

CodeAlign v1.0

Choose a tag to compare

@Sergasgr Sergasgr released this 27 Sep 11:38
· 2 commits to main since this release

First complete run of the pipeline: CommitPackFT curation (122,066 samples, 8 languages) → QLoRA SFT → preference pairs labelled by a sandbox and static analysis → DPO with a composite reward, an execution-only reward and a size-matched control → HumanEval with EvalPlus's chat protocol.

Model HumanEval pass@1 (95% CI) Mean CC
Base (Qwen2.5-Coder-7B-Instruct) 90.2% [84.7, 93.9] 3.70
SFT 70.7% [63.4, 77.2] 3.54
DPO, execution-only reward (3,267 pairs) 53.0% [45.4, 60.5] 3.28
DPO, composite reward, size-matched (3,267 pairs) 45.1% [37.7, 52.8] 2.03
DPO, composite reward (30,628 pairs) 24.4% [18.5, 31.5] 1.42
  • The evaluation reproduces the published baseline: 90.2% against the 88.4% in the Qwen2.5-Coder technical report.
  • SFT and every DPO run lowered pass@1. DPO reused SFT's learning rate (2e-4) and reached a reward accuracy of 1.00.
  • At a matched pair count, the composite reward lowered cyclomatic complexity (−1.25) with no detectable pass@1 difference.
  • Both rewards were gamed: "runs without errors" is satisfied by commented-out code.

Protocol, controls and limitations: README → Results (v1.0). Curated dataset: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft