Repository navigation
Features
⚡ Fused LM head: up to 6.9× longer sequences on the same GPU
SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built. It is on by default, there is nothing to turn on.
Max trainable sequence length (gemma-3-1b, 262k vocabulary, 79 GiB of GPU memory, bf16, batch size 1, gradient checkpointing, sdpa):
| Trainer | v1.14.2 | v1.15.0 | |
|---|---|---|---|
| DPO | 10,240 | 59,392 | 5.80× |
| KTO (batch size 2) | 9,216 | 63,488 | 6.89× |
| GRPO (scoring only) | 28,672 | 114,688 | 4.00× |
| RLOO (scoring only) | 23,552 | 100,352 | 4.26× |
SFT (loss_type="nll") |
20,480 | 107,520 | 5.25× |
SFT (default chunked_nll) |
107,520 | 107,520 | 1.00× |
At 8,192 tokens, peak memory drops 52% to 82% and steps are 2.3% to 10.9% faster:
| Trainer | Peak GiB v1.14.2 → v1.15.0 | Tokens/s |
|---|---|---|
| DPO | 48.97 → 12.00 (-75.5%) | 1.035× |
| KTO | 64.91 → 11.98 (-81.5%) | 1.045× |
| GRPO | 24.15 → 10.30 (-57.4%) | 1.023× |
| RLOO | 29.28 → 14.02 (-52.1%) | 1.029× |
SFT (nll) |
30.33 → 8.87 (-70.7%) | 1.029× |
The gain is largest for small models with large vocabularies. SFT's previous default (chunked_nll) already avoided full logits, so it is unchanged. Scoring needs Triton on a GPU (Linux with CUDA, ROCm or XPU).
💥 What this changes for you. The full-logits scoring path, the use_liger_kernel chunked path and _forward_redirection are gone from GRPO, RLOO, DPO and KTO:
- A PEFT adapter on
lm_headnow raises. Usemodules_to_save=["lm_head"]instead. use_liger_kernel=Trueis deprecated in these four trainers and will be removed in v2.0.0. Liger's layer kernels still apply, butfused_linear_cross_entropyis forced toFalse(it would replace the fused head) and setting it explicitly raises. Usemodel_init_kwargs={"use_kernels": True}.- In DPO and KTO,
compute_metricsandcompute_loss(..., return_outputs=True)still receive the full logits, from an extra forward pass taken only when they are used. Both warn that this is deprecated and will be removed in v2.0.0.
Nothing else the trainers logged is lost: the kernel gained log_sum_sq_probs (WPO weighting), mean_logits (logits/chosen, logits/rejected) and is_top1 (DPO's mean_token_accuracy) as opt-in outputs, each checked against the full logits in tests.
Benchmark details
1× B300 with the PyTorch allocator capped at 79 GiB (about an H100 80GB's usable memory), google/gemma-3-1b-pt built from its config with random bf16 weights, lr 1e-6, AdamW, synthetic token ids at exact length. Max length: one fresh process per attempt, 2 training steps, doubling from 8k then bisection to 1024 tokens, arms interleaved round-robin in the same job. Throughput: 12 steps, first 2 discarded, median of 3 A/B/A/B repeats in the same job. Noise floor (A vs A) is 0.1% to 0.8%, below every gap reported. GRPO/RLOO replace generation with fixed random completions, so only scoring is measured.
Not measured: H100 hardware itself, flash-attention (sdpa only), real generation, multi-GPU.
- Compute the chunked log-probabilities with a Triton kernel by @qgallouedec in #7386
- Add an opt-in fused LM head to the patched forward by @qgallouedec in #7387
- 💥 Score GRPO and RLOO tokens with the fused LM head by @qgallouedec in #7389
- 💥 Score DPO and KTO tokens with the fused LM head by @qgallouedec in #7390
- Score SFT tokens with the fused LM head by @qgallouedec in #7466
- Compute the distillation divergence with a Triton kernel by @qgallouedec in #7465
- Compute only the requested fused LM head outputs by @qgallouedec in #7406
- Rename
patch_fused_lm_headtoadd_fused_lm_headby @qgallouedec in #7415 - Honor
logits_scalingandlm_head_multiplierin the fused LM head by @YaseenBashaT in #7439 - Fix the fused LM head for models split across devices by @albertvillanova in #7479
Selective activation checkpointing in SFT
With gradient checkpointing on, the attention output is now saved during the forward instead of being recomputed in the backward, recovering most of the checkpointing slowdown at long context for one extra hidden-state-sized tensor per layer. Same eager SAC approach torchtitan uses under FSDP2, no torch.compile needed.
SFTConfig(
gradient_checkpointing=True,
gradient_checkpointing_kwargs={"selective": True},
)PEFT + DeepSpeed ZeRO-3 is rejected: that combination needs reentrant checkpointing, which SAC cannot use.
Assistant-only loss on vision datasets in SFT
assistant_only_loss=True now works on vision datasets. SFTTrainer used to refuse it because processors did not return usable assistant masks once image placeholders were expanded; transformers#48793 fixed that, and it ships in transformers 5.18.
SFTConfig(assistant_only_loss=True) # now fine with an image datasetThe vision collator tokenizes conversational examples through the processor's apply_chat_template and requests the assistant masks to mask non-assistant tokens in the labels. Older transformers still raises, now with the required version in the message.
by @albertvillanova in #7377
Conversations in the completions table
Prompts and completions are logged as raw conversation lists instead of batch_decode'd flat strings, so the table renders turn by turn. Much more readable for multi-turn and tool-calling setups, and for VLMs.
The rendered system prompt (including tool definitions injected by the chat template) is no longer visible, since the original messages are shown rather than the fully-templated string.
by @qgallouedec in #5309
AsyncGRPO and AsyncDistillation
- AsyncGRPO: train OpenEnv harnesses from validated token captures by @adithya-s-k in #6947
- Add PEFT support to
AsyncDistillationTrainerby @kashif in #7476 - Fix duplicate completions in AsyncGRPO groups with data-parallel vLLM by @lewtun in #7549
Other
- Add SmolVLM original and training chat template with generation markers by @aazizyan in #5868
- Declare gradient accumulation loss scaling with
loss_is_scaled_for_gaby @qgallouedec in #7508 - Keep SFT's MoE aux loss independent of world size and gradient accumulation by @qgallouedec in #7570
- Discover environment tools on the class, not the instance by @Rome-1 in #7360
- Accept an IPv6
base_urlinVLLMClientby @qgallouedec in #7462 - Pass processor kwargs via
processor_kwargsin the VLM SFT collator by @qgallouedec in #7547 - Call
trainer.end()at the end of the scripts and examples by @qgallouedec in #7384 - Install vLLM in the TRL Docker images by @qgallouedec in #7464
- Train harnesses through Harbor: drop the standalone opencode example by @sergiopaniego in #7457
- Add support for vLLM 0.31.0 by @qgallouedec in #7567
Breaking changes and removals
- Drop Python 3.10 support by @qgallouedec in #7497
- Refuse
nn.DataParallel(and drop the multi-GPU slow test job) by @qgallouedec in #7407 - Remove the experimental MiniLLM trainer by @qgallouedec in #7500
use_liger_kernelis deprecated inDPOTrainer,KTOTrainer,GRPOTrainerandRLOOTrainerand will be removed in v2.0.0. These trainers now use the fused LM head, so only Liger's layer kernels apply; usemodel_init_kwargs={"use_kernels": True}instead. Docs updated by @qgallouedec in #7535- Drop vLLM 0.20.0 / 0.20.1 / 0.20.2 support by @qgallouedec in #7435, #7499 and #7568
- Remove the unreachable
prepare_multimodal_messages_vllmby @behroozazarkhalili in #6946
Fixes
- Fix chat template drift: recognize the current LFM2 (#7494) and LFM2.5 (#7493) chat templates, the other Gemma 4 revisions (#7495), parse the current LFM2 tool calls (#7522), and read templates shipped as named variants by their default variant (#7530), all by @albertvillanova
- Fix
SFTTrainernot appending EOS whendataset_text_fieldis not"text"by @JoeyTan21 in #7446 - End completions on every eos id the model declares in GRPO, RLOO and Distillation by @albertvillanova in #7505, and in the experimental trainers in #7506
- Start the completion where the tokenized prompt and prompt+completion diverge by @qgallouedec in #7463
- Fix the server-mode crash at the first weight sync on vLLM 0.20 to 0.25 by @albertvillanova in #7410
- Generate one completion per sample after tool calls in vLLM server mode by @qgallouedec in #7418
- Fix
precompute_ref_log_probsunder FSDP with no reference model by @qgallouedec in #7507 - Keep the DFT loss finite on a batch with no trainable tokens by @albertvillanova in #7471
- Guard the
requestsandurllib3imports in the vLLM client by @qgallouedec in #7545, and in the experimental async vLLM clients by @albertvillanova in #7561 fix(cpo, orpo): truncate responses independently to prevent empty completions by @RohanMali2003 in #6588- Fix FLOPs overcount for models with untied word embeddings by @jayzuccarelli in #7450
[MFU]Use device peak flops instead of hardcoded H100's by @dwarez in #7197- Fix segfault on Apple Silicon when loading bf16 checkpoints from a model ID by @albertvillanova in #7455
- Keep the DoRA magnitude vector in float32 under QLoRA by @qgallouedec in #7342
- Fix missing OpenReward outcome rewards by @adithya-s-k in #7467
- Keep the next turn's first token out of the Llava-Next assistant mask by @albertvillanova in #7375
- Tie weights after loading remote-code models with transformers 5.3.0 by @albertvillanova in #7559
- Make SDFT importable without peft and align SSD by @albertvillanova in #7562
- Preserve explicit
model_init_kwargsin CLI scripts by @DaoyuanLi2816 in #7394 - Honor
resume_from_checkpointin CLI scripts by @DaoyuanLi2816 in #7391 - Seed before creating the model so the PEFT adapter init is reproducible by @qgallouedec in #7385, and in experimental trainers by @albertvillanova in #7401
- Drop the redundant spawn start method from
trl vllm-serveby @albertvillanova in #7379 FIX:test_offloading_with_peft_modelson L40, register model params in activation offloading tests by @AmineDiro in #7380- Drop the dead
gate_projLoRA target from the Nemotron 3 SFT example by @behroozazarkhalili in #6866 - Target the Mamba
in_projin the Nemotron 3 LoRA example by @behroozazarkhalili in #6870 - Replace the introspection guard in
sft_qwen3_8b_1m_context.pywith a version check by @kumarrah2002 in #7537
Documentation
- Expand the Jobs training guide by @davanstrien in #7441
- Document
expandable_segmentsallocator config for memory tuning by @akshansh47 in #5794 - Update "What's New" for the fused LM head by @qgallouedec in #7593
- Document every stored chat template and check it in CI by @albertvillanova in #7502
- Document what goes into a patch release by @albertvillanova in #7440
- Document input validation guidelines by @albertvillanova in #7472
- Unwrap hard-wrapped paragraphs in the docs by @qgallouedec in #7405
- Discourage
SimpleNamespace,object.__new__and mocks in tests in the agent instructions by @qgallouedec in #7459 - Require first-time contributors to link an issue assigned to them by @qgallouedec in #7313
- Extend the AI usage policy to issues by @albertvillanova in #7473
- Require bug reports to state how the problem was hit by @albertvillanova in #7474
- Label CONTRIBUTING.md changes as documentation by @albertvillanova in #7475
CI
Chat template coverage, so drift stops recurring, all by @albertvillanova: document every stored chat template and check it in CI (#7502), add a script checking that reference Hub repos still ship a stored chat template (#7523), check it weekly (#7521), and keep the earlier LFM2 revision under test (#7527)
Tiny-model config alignment, continuing the sweep, all by @albertvillanova: DeepSeek-V3 (#7398) and V3-0528 (#7399), GPT-OSS (#7424), Qwen2-VL (#7425) and Qwen2.5-VL (#7426) mrope sections, Nemotron-3-Nano (#7443), Nemotron-3.5-Lightning (#7444), Nemotron-3-Super (#7447), Nemotron-3-Ultra (#7448), LFM2 (#7400), LFM2.5 (#7532)
Test quality, following the "real objects, not mocks" rule, all by @albertvillanova: build real trainers instead of object.__new__ in the GRPO and SDFT tests (#7531), build real configs instead of SimpleNamespace in the GOLD and GKD tests (#7564), build the async rollout loops through __init__ (#7555), run the async GRPO rollout loop tests on the real tokenizer and response parser (#7556), align the DFT loss tests on a real model output (#7481)
Suite hygiene, all by @albertvillanova unless noted: move the vLLM training (#7488), continuous batching (#7487) and SFT activation offloading (#7477) tests to the regular suite, unmark the vLLM client/server tests as slow (#7490), run them on a single xdist worker (#7489), replace the SFT slow tests with an fp16 test (#7550), remove the always-skipped Gemma 3n tests (#7485), drop the no-op slow and low_priority markers (#7486), stop piping the vLLM test server output (#7414), run the VLM test server on the last visible accelerator (#7412), test the vLLM client against the documented vllm serve command (#7378), inject MODEL_REVISIONS into the Auto loaders (#7528), mirror the pad token onto the model config in the callback tests (#7383), remove the DataParallel gather warning filter (#7554), tie the mamba-ssm autocast warning filter removal to transformers 5.6 (#7431), remove the nightly cache cleanup workflow (#7395), scope GITHUB_TOKEN permissions per job by @hf-security-analysis (#7381), bump the doc-builder workflow pin by @paulinebm (#7542), Dependabot group bump (#7411)
Hotfixes: xfail the GRPO continuous batching test under DataParallel (#7422), xfail Qwen3.5-MoE VLM tests against a broken transformers dev build (#7437) and revert once fixed (#7460)
New Contributors
- @dwarez made their first contribution in #7197
- @Rome-1 made their first contribution in #7360
- @JoeyTan21 made their first contribution in #7446
- @jayzuccarelli made their first contribution in #7450
- @RohanMali2003 made their first contribution in #6588
- @kumarrah2002 made their first contribution in #7537
What's Changed
- ⬆️ Bump dev version by @albertvillanova in #7393
- Scope GITHUB_TOKEN permissions per job by @hf-security-analysis[bot] in #7381
- Mirror the pad token onto the model config in the callback tests by @albertvillanova in #7383
- Drop the redundant spawn start method from
trl vllm-serveby @albertvillanova in #7379 - Seed before creating the model so the PEFT adapter init is reproducible by @qgallouedec in #7385
- Compute the chunked log-probabilities with a Triton kernel by @qgallouedec in #7386
- Add an opt-in fused LM head to the patched forward by @qgallouedec in #7387
- Unwrap hard-wrapped paragraphs in the docs by @qgallouedec in #7405
- Score GRPO and RLOO tokens with the fused LM head by @qgallouedec in #7389
- Fix the server-mode crash at the first weight sync on vLLM 0.20 to 0.25 by @albertvillanova in #7410
- Bump the "github-actions-and-pre-commit" group with 2 updates across multiple ecosystems by @dependabot[bot] in #7411
- Seed before creating the model in experimental trainers by @albertvillanova in #7401
- Align tiny DeepSeek-V3-0528 config with deepseek-ai/DeepSeek-R1-0528 by @albertvillanova in #7399
- Hotfix CI: xfail the GRPO continuous batching test under DataParallel by @albertvillanova in #7422
- Align tiny DeepSeek-V3 config with deepseek-ai/DeepSeek-R1 by @albertvillanova in #7398
- Remove the nightly cache cleanup workflow by @albertvillanova in #7395
- [Fix] Generate one completion per sample after tool calls in vLLM server mode by @qgallouedec in #7418
- Preserve explicit model_init_kwargs in CLI scripts by @DaoyuanLi2816 in #7394
- Honor resume_from_checkpoint in CLI scripts by @DaoyuanLi2816 in #7391
- Call trainer.end() at the end of the scripts and examples by @qgallouedec in #7384
- Rename
patch_fused_lm_headtoadd_fused_lm_headby @qgallouedec in #7415 - Hotfix CI: Xfail Qwen3.5-MoE VLM tests against a broken transformers dev build by @albertvillanova in #7437
- Align tiny GPT-OSS config with openai/gpt-oss-20b by @albertvillanova in #7424
- Align tiny Qwen2-VL mrope_section with Qwen/Qwen2-VL-2B-Instruct by @albertvillanova in #7425
- Align tiny Qwen2.5-VL mrope_section with Qwen/Qwen2.5-VL-3B-Instruct by @albertvillanova in #7426
- Expand the Jobs training guide by @davanstrien in #7441
- Keep the next turn's first token out of the Llava-Next assistant mask by @albertvillanova in #7375
- Tie the mamba-ssm autocast warning filter removal to transformers 5.6 by @albertvillanova in #7431
- Stop piping the vLLM test server output so it no longer hides errors or blocks the server by @albertvillanova in #7414
- Run the VLM test server on the last visible accelerator so it can start on 24 GB GPUs by @albertvillanova in #7412
- Align tiny Nemotron-3-Nano config with nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 by @albertvillanova in #7443
- Align tiny Nemotron-3.5-Lightning config with nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 by @albertvillanova in #7444
- Align tiny Nemotron-3-Super config with nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 by @albertvillanova in #7447
- Align tiny Nemotron-3-Ultra config with nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 by @albertvillanova in #7448
- Fix FLOPs overcount for models with untied word embeddings by @jayzuccarelli in #7450
- Add selective activation checkpointing (SAC) to SFT by @kashif in #7194
- Fix SFTTrainer not appending EOS when
dataset_text_fieldis not"text"by @JoeyTan21 in #7446 - Require first-time contributors to link an issue assigned to them by @qgallouedec in #7313
- Revert xfail for Qwen3.5-MoE VLM tests now that transformers fixed router_logits by @albertvillanova in #7460
- Drop vLLM 0.20.0 support by @qgallouedec in #7435
- [MFU] Use device peak flops instead of hardcoded H100's by @dwarez in #7197
- Fix segfault on Apple Silicon when loading bf16 checkpoints from a model ID by @albertvillanova in #7455
- Test the vLLM client against the documented
vllm servecommand by @albertvillanova in #7378 - Support assistant-only loss for vision datasets in SFT by @albertvillanova in #7377
- AsyncGRPO: train OpenEnv harnesses from validated token captures by @adithya-s-k in #6947
- Discourage SimpleNamespace, object.new and mocks in tests in agent instructions by @qgallouedec in #7459
- Accept an IPv6 base_url in VLLMClient by @qgallouedec in #7462
- FIX: test_offloading_with_peft_models on L40. Register model params in activation offloading tests by @AmineDiro in #7380
- Remove the unreachable prepare_multimodal_messages_vllm by @behroozazarkhalili in #6946
- docs: document expandable_segments allocator config for memory tuning by @akshansh47 in #5794
- Document input validation guidelines by @albertvillanova in #7472
- Extend the AI usage policy to issues by @albertvillanova in #7473
- Require bug reports to state how the problem was hit by @albertvillanova in #7474
- Label CONTRIBUTING.md changes as documentation by @albertvillanova in #7475
- Keep the DFT loss finite on a batch with no trainable tokens by @albertvillanova in #7471
- Move SFT activation offloading test to the regular suite by @albertvillanova in #7477
- Keep the DoRA magnitude vector in float32 under QLoRA by @qgallouedec in #7342
- Document what goes into a patch release by @albertvillanova in #7440
- Install vLLM in the TRL Docker images by @qgallouedec in #7464
- Drop the dead gate_proj LoRA target from the Nemotron 3 SFT example by @behroozazarkhalili in #6866
- Target the Mamba in_proj in the Nemotron 3 LoRA example by @behroozazarkhalili in #6870
- Add SmolVLM original and training chat template with generation markers by @aazizyan in #5868
- Honor logits_scaling and lm_head_multiplier in the fused LM head by @YaseenBashaT in #7439
- Discover environment tools on the class, not the instance by @Rome-1 in #7360
- fix(cpo, orpo): truncate responses independently to prevent empty completions by @RohanMali2003 in #6588
- Recognize the current LFM2.5 chat template by @albertvillanova in #7493
- Recognize the other Gemma 4 chat template revisions by @albertvillanova in #7495
- Run the vLLM server tests on a single xdist worker by @albertvillanova in #7489
- Unmark the vLLM client/server tests as slow by @albertvillanova in #7490
- Recognize the current LFM2 chat template for assistant-only loss by @albertvillanova in #7494
- Align tiny LFM2 config with LiquidAI/LFM2-1.2B by @albertvillanova in #7400
- Remove the always-skipped Gemma 3n tests by @albertvillanova in #7485
- Drop the no-op slow and low_priority markers from the experimental tests by @albertvillanova in #7486
- Move the vLLM training test to the regular suite and add it to RLOO by @albertvillanova in #7488
- Drop Python 3.10 support by @qgallouedec in #7497
- Drop vLLM 0.20.1 support by @qgallouedec in #7499
- Train harnesses through Harbor: drop the standalone opencode example by @sergiopaniego in #7457
- Fix
precompute_ref_log_probsunder FSDP with no reference model by @qgallouedec in #7507 - Fix missing OpenReward outcome rewards by @adithya-s-k in #7467
- Fix the fused LM head for models split across devices by @albertvillanova in #7479
- Inject MODEL_REVISIONS into the Auto tokenizer and processor loaders in tests by @albertvillanova in #7528
- Align tiny LFM2.5 config with LiquidAI/LFM2.5-230M by @albertvillanova in #7532
- Add a script checking that reference Hub repos still ship a stored chat template by @albertvillanova in #7523
- Check weekly that reference Hub repos still ship a stored chat template by @albertvillanova in #7521
- Align the DFT loss tests on a real model output by @albertvillanova in #7481
- Document every stored chat template and check it in CI by @albertvillanova in #7502
- Bump doc-builder workflow pin by @paulinebm in #7542
- Remove the experimental MiniLLM trainer by @qgallouedec in #7500
- Refuse nn.DataParallel and drop the multi-GPU slow test job by @qgallouedec in #7407
- Replace the introspection guard in sft_qwen3_8b_1m_context.py with a version check by @kumarrah2002 in #7537
- Fix: Start the completion where the tokenized prompt and prompt+completion diverge by @qgallouedec in #7463
- Show conversations instead of decoded text in the completions table by @qgallouedec in #5309
- End completions on every eos id the model declares in GRPO, RLOO and Distillation by @albertvillanova in #7505
- Guard the requests and urllib3 imports in the vLLM client by @qgallouedec in #7545
- Add PEFT support to AsyncDistillationTrainer by @kashif in #7476
- Fix duplicate completions in AsyncGRPO groups with data-parallel vLLM by @lewtun in #7549
- End completions on every eos id the model declares in the experimental trainers by @albertvillanova in #7506
- Build real trainers instead of object.new in the GRPO and SDFT tests by @albertvillanova in #7531
- Move the continuous batching test to the regular suite and add it to RLOO by @albertvillanova in #7487
- Pass processor kwargs via processor_kwargs in the VLM SFT collator by @qgallouedec in #7547
- Compute only the requested fused LM head outputs by @qgallouedec in #7406
- Score DPO and KTO tokens with the fused LM head by @qgallouedec in #7390
- Add support for vLLM 0.31.0 by @qgallouedec in #7567
- Drop vLLM 0.20.2 support by @qgallouedec in #7568
- Score SFT tokens with the fused LM head by @qgallouedec in #7466
- Parse the tool calls of the current LFM2 chat template by @albertvillanova in #7522
- Keep the earlier LFM2 chat template revision under test by @albertvillanova in #7527
- Read chat templates shipped as named variants by their default variant by @albertvillanova in #7530
- Replace the SFT slow tests with an fp16 test by @albertvillanova in #7550
- Remove the DataParallel gather warning filter by @albertvillanova in #7554
- Tie weights after loading remote-code models with transformers 5.3.0 by @albertvillanova in #7559
- Guard the requests import in the experimental async vLLM clients by @albertvillanova in #7561
- Make SDFT importable without peft and align SSD by @albertvillanova in #7562
- Build real configs instead of SimpleNamespace in the GOLD and GKD tests by @albertvillanova in #7564
- Keep SFT's MoE aux loss independent of world size and gradient accumulation by @qgallouedec in #7570
- Build the async rollout loops in tests through init by @albertvillanova in #7555
- Run the async GRPO rollout loop tests on the real tokenizer and response parser by @albertvillanova in #7556
- Compute the distillation divergence with a Triton kernel by @qgallouedec in #7465
- Drop the deprecated
use_liger_kernelfrom the docs by @qgallouedec in #7535 - Declare gradient accumulation loss scaling with
loss_is_scaled_for_gaby @qgallouedec in #7508 - Update "What's New" for the fused LM head by @qgallouedec in #7593
- Release: v1.15 by @qgallouedec in #7594
Full Changelog: v1.14.0...v1.15.0