Skip to content

v1.14.2

Choose a tag to compare

@qgallouedec qgallouedec released this 06 Oct 20:30
· 143 commits to main since this release

Patch release fixing two cases of silently wrong training and three crashes.

Worth calling out: #7505 affects GRPO, RLOO and Distillation on the default generation path (use_vllm=False). Models that end a turn with an id only their generation config declares (Gemma 3/4, Phi-3.5, ...) did not stop at the end of the turn, and everything generated after it was trained on, up to max_completion_length. The only visible sign was completions/clipped_ratio close to 1. With vLLM, generation stopped correctly but the metrics were still wrong, and mask_truncated_completions=True dropped every finished completion. Runs on models whose tokenizer eos is their only end-of-turn id are unchanged.

What's Changed

  • Keep the DFT loss finite on a batch with no trainable tokens by @albertvillanova in #7471
  • Fix precompute_ref_log_probs under FSDP with no reference model by @qgallouedec in #7507
  • Fix: Start the completion where the tokenized prompt and prompt+completion diverge by @qgallouedec in #7463
  • End completions on every eos id the model declares in GRPO, RLOO and Distillation by @albertvillanova in #7505
  • Guard the requests and urllib3 imports in the vLLM client by @qgallouedec in #7545

Full Changelog: v1.14.1...v1.14.2