Skip to content

Releases: Sergasgr/CodeAlign

CodeAlign v1.1

Choose a tag to compare

@Sergasgr Sergasgr released this 04 Oct 15:50

Overview

v1.1 re-trains only DPO, on v1.0's preference pairs, changing only the learning rate and the number of epochs, to find what caused v1.0's DPO results. Everything else (β, LoRA, NF4, effective batch 16, seed, evaluation protocol) is v1.0's.

Evaluation results (HumanEval pass@1, EvalPlus chat protocol, greedy, 164 problems)

DPO run lr × epochs pass@1 (95% CI) Δ vs. SFT (McNemar p) Mean CC
SFT (v1.0), reference — 70.7% [63.4, 77.2] — 3.54
Execution-only (v1.0) 2e-4 × 3 53.0% [45.4, 60.5] −17.7 (<0.001) 3.28
Execution-only 2e-4 × 1 74.4% [67.2, 80.5] +3.7 (0.44) 4.09
Execution-only 5e-6 × 3 69.5% [62.1, 76.0] −1.2 (0.85) 3.16
Execution-only 5e-6 × 1 68.9% [61.5, 75.5] −1.8 (0.69) 3.15
Composite, size-matched 5e-6 × 1 69.5% [62.1, 76.0] −1.2 (0.84) 3.14
Composite (30,628 pairs) 5e-6 × 1 68.3% [60.8, 74.9] −2.4 (0.60) 3.12

Key findings

  • v1.0's DPO collapse came from training three epochs at lr 2e-4. One epoch at the same rate scores 74.4%, 21.3 points above three (p < 0.001). At lr 5e-6 the epoch count makes no difference.
  • No v1.1 run is distinguishable from SFT (McNemar p 0.44–0.85). Trained gently, DPO on these "runs without errors" pairs neither costs nor measurably adds correctness.
  • With a gentle recipe the reward makes no detectable difference, in pass@1 or in complexity (composite vs. execution-only CC −0.02 [−0.20, +0.21]).
  • The pairs favour no-op code, rarely: 3.0% of chosen vs. 1.3% of rejected Case A candidates are mostly comments, define nothing, or gut the file they edit (p < 10⁻⁷). Over-optimization amplifies it: no-op HumanEval solutions go 0 → 5 → 34 as training gets more aggressive.

What's new in the code

  • dpo_trainer.py: --learning_rate, --num_train_epochs, --dataset, --run_name (defaults reproduce v1.0); each v1.1 run saves a run_config.json.
  • scripts/audit_preference_pairs.py, scripts/filter_noop_pairs.py, src/preference_generation/noop.py: no-op audit and filter for preference pairs.
  • scripts/eval_checkpoint.sh, src/evaluation/per_problem.py: single-adapter evaluation with v1.0's flags, and per-problem results for paired tests.
  • src/notebooks/06_v1.1_dpo_recipe.ipynb: the full v1.1 analysis.
  • ROADMAP.md: the plan for v2.x–v3.0.

Full results and caveats: README, Results (v1.1). Curated dataset (unchanged): Sergasgr/codealign-commitpackft.

CodeAlign v1.0

Choose a tag to compare

@Sergasgr Sergasgr released this 27 Sep 11:38

First complete run of the pipeline: CommitPackFT curation (122,066 samples, 8 languages) → QLoRA SFT → preference pairs labelled by a sandbox and static analysis → DPO with a composite reward, an execution-only reward and a size-matched control → HumanEval with EvalPlus's chat protocol.

Model HumanEval pass@1 (95% CI) Mean CC
Base (Qwen2.5-Coder-7B-Instruct) 90.2% [84.7, 93.9] 3.70
SFT 70.7% [63.4, 77.2] 3.54
DPO, execution-only reward (3,267 pairs) 53.0% [45.4, 60.5] 3.28
DPO, composite reward, size-matched (3,267 pairs) 45.1% [37.7, 52.8] 2.03
DPO, composite reward (30,628 pairs) 24.4% [18.5, 31.5] 1.42
  • The evaluation reproduces the published baseline: 90.2% against the 88.4% in the Qwen2.5-Coder technical report.
  • SFT and every DPO run lowered pass@1. DPO reused SFT's learning rate (2e-4) and reached a reward accuracy of 1.00.
  • At a matched pair count, the composite reward lowered cyclomatic complexity (−1.25) with no detectable pass@1 difference.
  • Both rewards were gamed: "runs without errors" is satisfied by commented-out code.

Protocol, controls and limitations: README → Results (v1.0). Curated dataset: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft