Repository navigation
v1.14.2
Patch release fixing two cases of silently wrong training and three crashes.
Worth calling out: #7505 affects GRPO, RLOO and Distillation on the default generation path (use_vllm=False). Models that end a turn with an id only their generation config declares (Gemma 3/4, Phi-3.5, ...) did not stop at the end of the turn, and everything generated after it was trained on, up to max_completion_length. The only visible sign was completions/clipped_ratio close to 1. With vLLM, generation stopped correctly but the metrics were still wrong, and mask_truncated_completions=True dropped every finished completion. Runs on models whose tokenizer eos is their only end-of-turn id are unchanged.
What's Changed
- Keep the DFT loss finite on a batch with no trainable tokens by @albertvillanova in #7471
- Fix
precompute_ref_log_probsunder FSDP with no reference model by @qgallouedec in #7507 - Fix: Start the completion where the tokenized prompt and prompt+completion diverge by @qgallouedec in #7463
- End completions on every eos id the model declares in GRPO, RLOO and Distillation by @albertvillanova in #7505
- Guard the requests and urllib3 imports in the vLLM client by @qgallouedec in #7545
Full Changelog: v1.14.1...v1.14.2