Highlights
- Frontier-scale MoE and multimodal models. Train Moonshot AI's Kimi K3 (2.8T/104B active), Thinking Machines Lab's Inkling (975B/41B active), MiniMax AI's MiniMax-M3 VL (428B/22B active), Poolside's Laguna, and Zhipu AI's GLM-5.2.
- Long-context training. Shard packed multi-document sequences with block-diagonal variable-length context parallelism, and shard vision towers by frame to remove the redundant vision compute of a CP fold; recipes reach 128K tokens at EP8/CP32.
- New attention and kernel backends. Select FlashAttention 3 and 4, MagiAttention, FFPA for headdim=512 layers, QuACK linear/RMSNorm/RoPE kernels, and scoped partial CUDA graphs for MoE.
- Resilient checkpointing. Checkpoint on preemption signals, write to
msc://cloud directories, consolidate async saves in the background, bound checkpoint retention, and resume past an interrupted save. - Diffusion. Fine-tune and generate with LTX-2.3 joint video+audio and Qwen-Image-Edit-2511, and run context parallelism through the diffusers
ContextParallelConfigAPI.
New Hardware and Precision Support
- FP8 draft training. Extends the SFT recipes' top-level
fp8:block to every speculative recipe, swapping the draft'snn.Linearlayers to torchaoFloat8Linearon SM89 and newer, withemulate: truefor older GPUs. (#2963) - Quantized checkpoint ingestion. Dequantizes MXFP4 dequantization for Kimi K3 and MXFP8 dequantization for MiniMax-M3 for BF16 training, keeps Kimi-K2.5-VL LoRA keys out of INT4 quantization on save, and fixes a TileLang
boolx8backward codegen failure on the DeepSeek-V4 path. (#3259, #3295, #3467)
New and Expanded Model Support
LLM and MoE
- Moonshot AI: Kimi K3 (2.8T / 104B-active MoE) — text-only HellaSwag full SFT at EP32/PP8 on 256 GB200 GPUs. (docs, #3259, #3283)
- Zhipu AI: GLM-5.2 (MoE with IndexShare DSA) — HellaSwag and Tulu-3 4K at EP64/PP4, Tulu-3 32K at CP8/EP64/PP4, LoRA at EP128; selectable TileLang DSA kernels. (docs, #2633, #2691, #2695, #3222)
- Poolside: Laguna S 2.1 (118B / ~8B-active MoE) — HellaSwag full SFT at EP16. (docs, #3148)
- DeepSeek: DeepSeek-V4 Flash (MoE) — Tulu-3 SFT at CP8/PP4/EP32 and pretraining at PP4/EP32; packed THD under context parallelism. (docs, #2590, #2731, #3227, #3231)
VLM
- MiniMax AI: MiniMax-M3 VL (428B / 22B-active MoE) — MedPix full SFT at EP32/PP4 on 16 nodes, MedPix LoRA at PP4/EP8, a CP2 comparison recipe, and 16K packed Tulu-3 text SFT at CP8. (docs, #2538, #2551)
- Thinking Machines Lab: Inkling (975B / 41B-active MoE; text, image, video, audio) — MedPix SFT at PP8/EP32 on 32 nodes. (docs, #3095, #3281, #3358)
- Alibaba: Qwen3.5 and Qwen3.6 VLM (122B-A10B MoE, 27B dense) — 128K packed sequences at EP8/CP32 on 16 nodes with a trainable vision tower, dense context parallelism, and matched CP1/CP2 27B MedPix recipes. (docs, #2505, #3186, #3209)
- Google: Gemma 4 (31B and E2B/E4B dense, 26B-A4B MoE) — context parallelism with 16K Tulu-3 at CP8, 64K Tulu-3 at CP16, and 4K MedPix at EP8/CP2; plus fp32 SDPA to avoid NaNs on Hopper and an exact fp32 reference router for MoE parity. (docs, #1914, #2592, #2621, #3141)
- Moonshot AI: Kimi-K2.5 VL (MoE) — LoRA tensors kept out of INT4 expert quantization on save, so adapters reload on resume. (#3295, #3431)
- NVIDIA: Nemotron-Parse — corrected preprocessing stopping the RADIO encoder from normalizing already-normalized images a second time. (#3331)
- MTP on multimodal inputs. Pre-rolls pre-fused
inputs_embedsper depth so audio and vision positions keep their continuous embedding instead of being re-embedded from a padding token id. (#2510) - Pre-extracted video frames. Loads a
videofield holding a list of image paths directly as a frame sequence with no decoder, padded to the processor'stemporal_patch_size. (#3211)
Diffusion
- Lightricks: LTX-2.3 (video + audio) — full-finetuning, LoRA, and generation recipes that train video and audio jointly and mux the generated waveform into the output mp4;
--processor ltx2preprocessing requires an8n+1frame count. (docs, #3165, #3372) - Alibaba: Qwen-Image-Edit-2511 — cached, full-parameter instruction-based editing with an offline latent and prompt cache encoder, an eight-GPU BF16 recipe, and an
image-editpreprocessing subcommand. (docs, #3217) - Diffusion context parallelism. Drives the diffusers
ContextParallelConfigAPI fromfsdp.cp_size, reusing thecpaxis of the existing FSDP2 mesh; only pure Ulysses sharding is accepted. (#3157)
dLLM
- Google: DiffusionGemma 26B-A4B (block-diffusion MoE) — a native implementation with GSM8K SFT and LoRA recipes at
ep_size: 8. (docs, #2506) - Generation and LoRA. Exposes
--samplerpresetsllada,llada2,nemotron, andgemma, an--adaptermerge path, andllada_lora,llada2_lora, andnemotron_labs_diffusion_lorarecipes. (#3092, #3161, #3163) - Reproducible corruption. Seeds corruption noise per sample from that example's global index in the shuffled stream, so noise is independent of parallel topology and resume reproduces exactly. (#3162)
Speculative Decoding
- DSpark. Introduces a semi-autoregressive draft in which a parallel backbone proposes a whole block, a serial Markov head adds intra-block dependency, and a confidence head predicts acceptance, with configs for Qwen3, Gemma 4, GLM-5.2, MiniMax-M3, and DeepSeek-V4-Flash. (docs, #2810, #2866, #2885, #2909)
- Domino, JetSpec, and ViSpec. Adds a GRU correction head on the DFlash backbone (Domino), causal in-block attention distilled with a temperature-scaled forward KL (JetSpec), and two-stage VLM draft training (ViSpec). (#2819, #2867, #3173)
- New EAGLE-3 targets. Registers a DeepSeek-V3 MLA draft class reusing the target's low-rank q/kv projections, and registers Gemma 4 as a target with E2B, E4B, 31B, and 26B-A4B configs. (docs, #2849, #3071, #3077, #3079)
- Sequence packing, TP, and CP. Enables block-causal packed training for every draft family, target tensor parallelism for EAGLE-1/2/3 and DFlash, and context parallelism for both target and draft through a differentiable ring attention. (#2827, #2918, #3002, #3005)
- fp8, compile, LoRA, and new objectives. Exposes a
compile:block, EAGLE-3 adapter-only training throughpeft:, DFlashloss_type: variable_prefix, and EAGLE-3lk_loss_type: alpha/lambda. (#2963, #3032, #3081) - Feature-noise augmentation. Perturbs the target features fed to the EAGLE-1/2 draft with uniform noise on the target features fed to the EAGLE-1/2 draft through
recipe_args.feature_noise, defaulting to the paper value0.1. (#2470) - On-policy regeneration. Interleaves EAGLE-3 training with online target regeneration through a
recipe_args.regenblock that interleaves EAGLE-3 training with online target regeneration on a reserved GPU, hot-swapping the dataloader across all ranks in lockstep. (#3042) - Serving, benchmarking, and caches. Covers
sglangandvllmtarget backends for training supervision,serve_vllmfor trained EAGLE-3, P-EAGLE, and DFlash-family drafts,bench_vllm/bench_sweepacceptance and speedup benchmarks, and distributed DSpark offline precompute for targets too large for one node. (#2798, #2841, #2912, #2953) - Metrics and resume. Reports simulated accept length, per-position
accept_rate@k, and periodic real accept length throughdecode_eval, and rejects a mismatchedmask_token_id, a stale draft-vocab mapping, and mixed precompute shards on resume. (#2959, #3037, #3237, #3238)
Training, Parallelism, and Performance
- Block-diagonal variable-length context parallelism. Introduces a
blockdiag_cpimplementation that shards packed multi-document sequences while keeping masking block-causal per document, driving varlen kernels from precomputedcu_seqlenswith a fused K/V all-gather. (#2989, #3223) - Unified context-parallel input preparation. Returns a
ContextParallelSharderfrom a singleprepare_model_inputs_for_cphook and replacescp_utilswithcomponents.distributed.context_parallel. (#2937) - Frame-level vision-tower sharding. Partitions frame units across the CP group through
distributed.multimodal.vision.frame_shardingand reassembles embeddings with a differentiable variable-length all-gather. (docs, #2990, #3186) - Attention backends. Selects
flash_attention_3andflash_attention_4with afa4 → fa3 → fa2 → sdpa → eagerfallback ladder, MagiAttention as both a kernel and a context-parallel backend with CP1/CP2 Qwen3-MoE-30B packed recipes, and a CuTeDSL FFPA backend for Gemma 4's head_dim=512 layers (8K, packed CP8 16K). (#2384, #2929, #3070) - QuACK kernels and partial CUDA graphs. Exposes
quackoptions forbackend.linear,backend.rms_norm, andbackend.rope, andbackend.cuda_graph.modulescapture ofattn,moe_router, andmoe_preprocesswhile dispatch, expert compute, and combine stay eager. (#2917, #3115) - Activation checkpointing. Accepts
activation_checkpointing: selectivefor DDP and anactivation_checkpointing_scopeselecting which layer groups are wrapped, and saves the MoE router and its top-k output so recompute preserves expert assignments. (#2786, #2840, #3140, #3247) - Cross-entropy memory. Supports
FusedLinearCrossEntropyunder pipeline parallelism and boundsChunkedCrossEntropymemory with a kernel that saves only original-dtype logits and recomputes each softmax chunk in backward. (#2927, #2996) - FSDP2 Wraps HF transformer layers with PyTorch's non-reentrant
checkpoint_wrapperto remove duplicate prefetch all-gathers, and addsdistributed.multimodal.frozen_shardingfor fully frozen towers. (#3328, #3411, #3513) - Gradient correctness. Normalizes the MoE auxiliary loss by the gradient-accumulation and pipeline-microbatch count, all-reduces grad-norm accumulators on the mesh device, and preserves HSDP replica gradient synchronization for MoE. (#3135, #3359, #3461)
- Damaged-embedding repair. Replaces, through an optional
embedding_row_repair:section, input-embedding rows whose L2 norm is non-finite or belowmin_normwith the scaled output-embedding direction; it is rejected under pipeline parallelism. (docs, #3136) - Setup-time prewarms and mesh timeouts. Initializes lazily created resources through an opt-in
prewarm:section coveringcublas_backward,fla_gdn_autotune,mamba_ssd_autotune, andcomm_groups, and appliesdist_env.timeout_minutesto flattened and expert-parallel meshes. (#2846, #2992, #3296) - Rollout Routing Replay. Reuses the rollout's discrete top-k expert selection in the training forward through
MoEConfig.enable_routing_replay, which reuses the rollout's discrete top-k expert selection in the training forward to remove the rollout/training routing mismatch in on-policy RL. (#2797) - Optimizer parameter groups and throughput. Adds
optimizer.param_group_overrideswithlr_multandwd_multmultipliers, removes Python overhead from Transformer Engine attention and MoE expert hot paths, and shapesTPLinear/LinearLoRAgraphs so async-TP fusion fires. (#2987, #3046, #3374) - PEFT under pipeline and expert parallelism. Gathers adapter weights across pipeline stages before writing and saves optimizer state for every PEFT model part namespaced by stage. (#3096, #3250, #3316)
- PEFT v5 expert adapters. Exports MoE expert LoRA through an opt-in PEFT v0.18+ ParamWrapper export for MoE expert LoRA with fused
target_parameters, validated on Qwen3-MoE, MiniMax-M2, and Nemotron v3. (#3439, #3458) - LoRA correctness. Makes
use_memory_efficient_lora: falsealso suppress the fused SwiGLU/ReLU² LoRA MLP, restores adapters through DDP wrappers, and reaches the absorbed GLM MLA KV weight through a newmaterialize_effective_weight(). (#3126, #3150, #3176, #3470)
Checkpointing
- Preemption checkpointing. Introduces
step_scheduler.preemption_signal, which watches one or more signals (defaultSIGTERM), gathers the flag across ranks, and saves at the next step boundary before exiting cleanly. (docs, #3007) - Cloud checkpoint directories. Accepts
msc://paths for sharded DCP state; cloud roots requiresave_consolidated: false, rejectmax_recent_checkpoints, and keep RNG, dataloader, and pointer files local. (#1709) - Retention and interrupted saves. Prunes older directories through
max_recent_checkpoints, and marks in-progress directories until every component is published so resume falls back instead of hanging all ranks. (#2416, #3261) - Async consolidation and new knobs. Runs consolidated safetensors export through the same all-rank consolidation as synchronous mode on a background thread, and adds
cpu_offload,wait_for_staging, andconsolidation_timeout_minutes. (#3108, #3125, #3130, #3131) - torchsave export to Hugging Face. Adds
scripts/export_llm_dcp_to_hf.py, which rebuilds the original topology from the checkpoint's recordedconfig.yamland dispatches the LLM or VLM recipe named there, so a VLM checkpoint keeps its vision tower. (#2487) - Export fidelity and safety. Keeps fp32 parameters fp32 even under an explicit
cast_dtype, writes safetensors without append, retains later-stage keys in pipeline-parallel consolidation, and loads auxiliarytorch.savestate withweights_only=True. (#2627, #3193, #3240, #3248) - Per-global-rank RNG state. Saves and restores RNG state per global rank rather than per data-parallel rank, so tensor- and pipeline-parallel peers keep distinct streams across a resume. (#3437)
- Tied-embedding guards. Declares a
TieSupportpolicy per model class and rejects an unsupportedtie_word_embeddingsvalue. (#2805, #2896, #2998)
Data
- Typed dataset configs. Gives every shipped dataset a typed
*Configwith abuildmethod and resolvesdataset:anddataloader:into a singleDataloaderConfig; legacy_target_ values still resolve through a compatibility registry. (#2390) - Prefix-tree attention for rollouts. Folds, through
prefix_tree_collate_fn, one shared prompt and its completions into a single deduplicated sequence with a block-sparse mask so the prompt is encoded once. (#2564) - Long-context agent-SFT data pipeline. Tokenizes CoderForge trajectories once in a dedicated stage and drops over-length ones rather than truncating, with a paired Gemma-4-31B recipe at
seq_length: 65536andcp_size: 8. (#3151, #3446) - Chat-template masking and pre-tokenization. Validates prefix-built answer-only masks against the full-conversation render, keeps
systemturns in the ShareGPT converter, and addspacked_sequence.num_procfor parallel pre-tokenization before packing. (#3024, #3045, #3144)
Embedding and Re-ranker
- Normalized Arrow retrieval datasets. Reads, through
NormalizedRetrievalDatasetConfig, a portable Arrow bundle that stores each referenced document or image once, with a CPU-side preparation script and Slurm wrappers. (docs, #2596) - NVIDIA: Nemotron VL 1B. Adds a native
llama_nemotron_vlmodel and processor fornvidia/llama-nemotron-embed-vl-1b-v2, trained as a vision bi-encoder with a bi-encoder recipe and a ColPali conversion notebook. (docs, #2354) - Cross-encoder re-ranking.
use_text_in_documentapplies to the cross-encoder transform as well as the bi-encoder, so a re-ranker can score an image document together with its text. (#3189) - Embedding distillation. Distills a teacher embedding model into a student with
EmbeddingDistillRecipeand an 8-GPU example, mixing cosine, MSE, and listwise InfoNCE terms with support for intermediate-layer distillation and cross-tokenizer cached teachers. (#3058) - Ministral 3 recipes. Loads the stock bidirectional Ministral backbone instead of a bespoke extraction, repairs the bi-encoder recipe, and guards the optional Weights & Biases import so the retrieval recipes import without it. (#3103, #3380, #3441)
- Sentence Transformers metadata. Writes
modules.json,1_Pooling/config.json, andconfig_sentence_transformers.jsonalongside consolidated bi-encoder exports, which the hard-negative miner reads back for pooling and normalization. (#3464)
Distillation
- Distillation on separate meshes. Places the teacher on its own ranks through
separate_meshes: trueand ateacher_distributed:block, with Llama-3.2 1B/3B and Qwen3.5 4B/9B examples under TP2, CP2, PP2, or DP2 teacher meshes. (#2954, #3280) - Distillation correctness. Resolves fp32 master weights for the student, runs gradient clipping on the recipe's own mesh with the expert TP replication factor, and passes tensor-valued hidden states through unchanged. (#3019, #3302, #3324)
Packaging and Dependencies
- Framework pins. Moves
transformersfrom 5.8.1 to 5.12.1 andmegatron-fsdpfrom 0.2.3 to a pinned 0.5.0 resolved from a Megatron-LM fork commit; both force an environment rebuild. (#2873) - QuACK is a default dependency. Promotes
quack-kernels==0.6.1to the base dependencies on Linux, which pullsnvidia-cutlass-dsl,apache-tvm-ffi, andtorch-c-dlpack-extinto every Linux install. (#3115) - New optional dependency sets. Adds an
ffpaextra withffpa-attnandnvidia-cutlass-dsl, and ships MagiAttention as a uv dependency group installed withuv sync --group magi; neither is part of[all]. (#2436, #3070, #3225) - Container. Builds FlashAttention 3 from source on x86 by default and leaves FlashAttention 4 opt-in behind
INSTALL_FA4=true, moves the optional PyTorch base stage topytorch:26.06-py3, and moves the deploy image tovllm:26.04-py3. (#2929, #2983, #3061, #3329) - Security constraints. Advances
aiohttp,cryptography,gitpython,pillow,starlette,mlflow, andthrift, and addspyarrowandrayconstraints. (#3346, #3398, #3482, #3523)
Media dependencies
Media packages remain opt-in. The vlm-media extra adds albumentations for the Nemotron-Parse processor and diffusion-media adds PyAV for LTX-2 audio decoding; the diffusion extra now requires diffusers>=0.39.0:
pip install 'nemo-automodel[media]'Use nemo-automodel[vlm-media] or nemo-automodel[diffusion-media] when only one of those media stacks is required.
Breaking Changes
tie_word_embeddingsis validated per model family and raises instead of leaving a randomly initializedlm_head. (#2805, #2896, #2998)- Tokenizer BOS/EOS insertion is opt-in; set
add_bos_token/add_eos_tokento keep the previous tokenization. (#3337) - Multi-turn masking raises on chat templates that rewrite earlier turns; affected templates must carry
{% generation %}blocks. (#3024, #3051) - Unrecognized keys under
dataloader:or on a typed dataset config now raise instead of being silently forwarded. (#2390) - DFlash requires an explicit
mask_token_id, andloss_decay_gammadefaults to 7.0. (#2461, #2908) BaichuanForCausalLM,Qwen2ForCausalLM, andKimiVLForConditionalGenerationare deprecated and scheduled for removal in 26.10. (#2807, #2884)
External Contributors 🎉
This release includes work from the following contributors outside the NVIDIA-NeMo organization. Thank you.
- @khazic — most of the speculative-decoding stack in this release: the DSpark, Domino, and JetSpec drafts, ViSpec VLM draft training, sequence packing and target tensor/context parallelism across every draft family, the vLLM and SGLang target backends, serving and acceptance benchmarks, and Gemma 4 ring context parallelism (94 PRs). (#2449, #2810, #2819, #2867, #3173)
- @Achyuthan-S — the
TieSupportpolicy andtie_word_embeddingsguards across the registered model families. (#2732, #2805, #2896, #2998) - @kashif — the DSpark draft model and training objective, context parallelism for draft training, and DSpark acceptance and confidence metrics. (#2810, #2918, #2957)
- @edjson — MSC cloud-storage support for DCP checkpoints and preemption checkpointing. (#1709, #3007)
- @Butterfingrz — the FFPA headdim=512 attention backend for Gemma 4. (#2436)
- @beccohov —
FusedLinearCrossEntropysupport under pipeline parallelism. (#2927) - @fkuner — the DSpark offline target-cache path. (#2924)
- @GITsologun — pre-extracted video frame sequences in the VLM data path. (#3211)
- @hyfine — gathering PEFT adapters across pipeline stages before save. (#3096)
- @wangzhxg — preserving HSDP replica gradient synchronization for MoE models. (#3135)
- @shahafwa — applying LoRA to the TileLang MLA KV projection. (#3176)
- @amolkhanna — activation checkpointing for Qwen3-Next linear-attention layers. (#3192)
- @aminehd — running Gemma 4 SDPA in fp32 to avoid NaNs on Hopper. (#3141)
- @huahuajhu — writing consolidated safetensors without append. (#2627)
- @zhiqi-li — omitting the unset FSDP reshard argument. (#2926)
- @DOGEUNNKIM — casting Gemma 4 dense parameters without casting buffers. (#2359)
- @grgkovac — per-validation-dataset logging in Weights & Biases. (#2526)
Known Issues
- Transformer Engine fused RoPE is force-disabled globally and overrides an explicit
rope_fusion: true. (#3028) - Inter-node DeepEP dispatch still faults and
deepepremains the default whenever DeepEP is importable; use HybridEP for inter-node EP. - FlashAttention 4 and the TileLang backend cannot coexist in one image; an
INSTALL_FA4=truebuild loses the TileLang backend used by DeepSeek-V4 and GLM MoE DSA. (#2929) - P-EAGLE rejects a remote target backend, a cached target path, and packed sequences. (#2466)
- Qwen-Image-Edit-2511 training is limited to cached latents on a single node, and diffusion context parallelism accepts only pure Ulysses sharding. (#3157, #3217)
Changelog Details
- fix(speculative): serialize remote EAGLE-3 /generate end to end by @khazic :: PR: #2479
- feat(speculative): add EAGLE feature-noise augmentation to EAGLE-1/2 training by @khazic :: PR: #2470
- fix(speculative): gate unsupported P-EAGLE backend / cache / packing combos by @khazic :: PR: #2466
- feat(dllm): add DiffusionGemma 26B-A4B block-diffusion SFT (full + LoRA) by @zyzhou5 :: PR: #2506
- test(checkpoint): fix TestFormatLoad directory-based reader selection by @khazic :: PR: #2520
- feat(speculative): add tqdm training progress bar to EAGLE and DFlash recipes by @khazic :: PR: #2522
- feat(vlm): add Gemma 4 31B joint drafter example config by @khazic :: PR: #2620
- fix: skip fused LoRA MLP install for meta weights by @akoumpa :: PR: #2775
- fix(checkpoint): resolve tie_word_embeddings top-level-first to match HF tying by @Achyuthan-S :: PR: #2732
- fix(benchmark): skip unsupported MTP flops by @akoumpa :: PR: #2767
- fix: Qwen3.5 MedPix EP32 NCCL timeout by @akoumpa :: PR: #2777
- fix(ci): address go-git/go-billy and rustls-webpki CVEs by @thomasdhc :: PR: #2780
- fix(ci): stabilize diffusion finetune smoke tests by @pthombre :: PR: #2788
- build(deps): move ffmpeg/opencv deps to opt-in media extra by @thomasdhc :: PR: #2743
- fix(vlm): keep Qwen3.5 media tokens aligned by @yuhezhang-ai :: PR: #2772
- fix(ci): HybridEP for multi-node MoE benchmarks + LoRA OOM fixes by @hemildesai :: PR: #2789
- fix: qwen3.5 and 3.6 mtp expert checkpoint layout by @HuiyingLi :: PR: #2778
- feat: CP support for MiniMax M3 by @athitten :: PR: #2551
- docs(speculative): fix EAGLE drafter layer count and EAGLE-1 loss by @khazic :: PR: #2796
- fix(ci): drop base-image uv/wandb copies flagged for CVEs by @thomasdhc :: PR: #2800
- chore(skills): refresh automodel skill signatures by @akoumpa :: PR: #2804
- feat(distributed): enable selective checkpointing for DDP by @yuhezhang-ai :: PR: #2786
- docs: align nightly navigation routes by @akoumpa :: PR: #2801
- feat: vision biencoder finetuning + Nemotron VL 1B finetuning support by @gabrielspmoreira :: PR: #2354
- fix(docs): catch Fern MDX syntax errors by @akoumpa :: PR: #2806
- feat(moe): add Rollout Routing Replay (R3) for MoE RL training by @khazic :: PR: #2797
- fix(optim): align Dion mesh with FSDP sharding by @akoumpa :: PR: #2808
- fix(ci): stabilize failed benchmark recipes by @akoumpa :: PR: #2817
- docs: document opt-in media extras (vlm-media/diffusion-media) by @thomasdhc :: PR: #2799
- feat(speculative): add SGLang target backend for EAGLE-3 training by @khazic :: PR: #2449
- feat(speculative): add Domino online training path on top of DFlash by @khazic :: PR: #2819
- feat(speculative): add dspark draft model and training objective by @kashif :: PR: #2810
- docs: add GLM-5.2 and speculative decoding (DSpark) updates by @khazic :: PR: #2828
- feat(diffusion): support Hugging Face datasets by @pthombre :: PR: #2816
- feat(speculative): add vLLM target backend for EAGLE-3 training by @khazic :: PR: #2798
- fix(speculative): gather sharded target lm_head in EAGLE-1/2 token loss by @khazic :: PR: #2823
- fix(model): honor yarn/linear/dynamic RoPE in LlamaRotaryEmbedding by @khazic :: PR: #2825
- fix(speculative): keep DFlash padding blocks self-attending to avoid NaN by @khazic :: PR: #2826
- feat(speculative): support target tensor parallelism in EAGLE-3 colocated path by @khazic :: PR: #2827
- feat(speculative): support target tensor parallelism in EAGLE-1/2 by @khazic :: PR: #2829
- feat(speculative): support target tensor parallelism in DFlash by @khazic :: PR: #2830
- fix(speculative): validate DFlash mask_token_id on checkpoint resume by @khazic :: PR: #2824
- docs: enable Fern multi-source by @lbliii :: PR: #2845
- fix(distributed): propagate NCCL timeout to derived device meshes by @HuiyingLi :: PR: #2846
- feat(speculative): add serve_vllm for EAGLE-3 / P-EAGLE drafts by @khazic :: PR: #2841
- fix(speculative): keep DSpark draft RoPE inv_freq in fp32 under bf16 by @khazic :: PR: #2859
- feat(speculative): compress EAGLE-3 offline cache target_probs via top-k by @khazic :: PR: #2847
- perf: FFPA D=512 attention backend for Gemma4 (3× fwd / 6× bwd vs SDPA) by @Butterfingrz :: PR: #2436
- feat(speculative): add JetSpec causal parallel drafting training by @khazic :: PR: #2867
- feat(models): reject tie_word_embeddings=True on separate-head model families by @Achyuthan-S :: PR: #2805
- ci: Ensure diffusion-media extra installed for diffusion tests by @chtruong814 :: PR: #2858
- feat(speculative): add DeepSeek-V3 (MLA) EAGLE-3 draft model by @khazic :: PR: #2849
- feat(speculative): add DeepSeek V4 DSpark drafter and V4-Flash training by @khazic :: PR: #2866
- chore(models): add 26.10 deprecation warnings for custom model classes by @athitten :: PR: #2807
- feat(speculative): add DSpark draft for MiniMax M3 VL (text + multimodal) by @khazic :: PR: #2877
- ci: use NVIDIA inference for Claude review by @chtruong814 :: PR: #2882
- feat(speculative): add GLM-5.2 DSpark draft model and training by @khazic :: PR: #2885
- fix(checkpoint): write consolidated safetensors without append by @huahuajhu :: PR: #2627
- ci(automodel): set AM-576 release timeouts by @yuhezhang-ai :: PR: #2871
- fix(diffusion_gemma): make parity test work on transformers >= 5.11 by @zyzhou5 :: PR: #2883
- fix(bagel): torch2.12 DTensor sharding-prop leak by @zyzhou5 :: PR: #2888
- fix(speculative): plumb EAGLE-3.1 fc_norm/norm_output into the draft config by @khazic :: PR: #2897
- fix(speculative): make save_consolidated final export the last EAGLE checkpoint by @khazic :: PR: #2898
- fix(speculative): gather TP target outputs in the Domino trainer by @khazic :: PR: #2899
- fix(speculative): default DFlash/Domino loss_decay_gamma to the paper value by @khazic :: PR: #2908
- refactor(speculative): dispatch all DSpark drafts through the registry by @khazic :: PR: #2909
- feat(speculative): serve DFlash and JetSpec drafts on vLLM by @khazic :: PR: #2913
- docs(speculative): document DSpark, Domino, JetSpec, and the vLLM serve path by @khazic :: PR: #2914
- feat(speculative): add bench_vllm acceptance/speedup benchmark by @khazic :: PR: #2912
- fix(speculative): read DeepSeek EAGLE-3 draft rope_theta from rope_parameters by @khazic :: PR: #2920
- fix(speculative): send chat_template_kwargs as a top-level request field by @khazic :: PR: #2921
- ci(deps): add albumentations to vlm-media for nemotron-parse processor by @thomasdhc :: PR: #2933
- fix(glm): select HybridEP for GLM-5.2 recipes by @HuiyingLi :: PR: #2934
- fix(distributed): omit unset FSDP reshard argument by @zhiqi-li :: PR: #2926
- feat(speculative): add DSpark offline cache path by @fkuner :: PR: #2924
- chore: update readme for 26.06-26.08 by @akoumpa :: PR: #2886
- ci: validate ci section on newly added recipes by @thomasdhc :: PR: #2936
- fix(bagel): make non-backbone init topology independent by @zyzhou5 :: PR: #2939
- feat(speculative): add fp8 draft training and LoRA draft adaptation by @khazic :: PR: #2963
- feat(speculative): report simulated accept length during EAGLE-3 training by @khazic :: PR: #2959
- fix(distributed): fix and scope vlm activation checkpointing by @yuhezhang-ai :: PR: #2840
- feat(models): complete tie_word_embeddings guards for #2512 families by @Achyuthan-S :: PR: #2896
- feat(speculative): add distributed DSpark offline precompute by @khazic :: PR: #2953
- ci: route gb200 L2_HF_DCP to the IMEX runner for DeepEP NVLink P2P by @ko3n1g :: PR: #2968
- docs: add 0.5.0 release notes by @lbliii :: PR: #2881
- fix(vlm): gemma4 FFPA mock + e4b cp16 CI recipe fixes (AM-626, AM-627) by @athitten :: PR: #2949
- build(deps): upgrade transformers to 5.12.1 by @athitten :: PR: #2873
- ci: update claude review guidelines by @akoumpa :: PR: #2739
- fix(moe): preserve MLP dispatch through checkpoint wrappers by @akoumpa :: PR: #2955
- feat(speculative): add multi-dataset acceptance-length benchmark sweep by @khazic :: PR: #2966
- fix(transformers): gate _tie_weights_nemo on tie_word_embeddings flag by @yuhezhang-ai :: PR: #2942
- fix(loss): avoid pkg_resources in linear CE by @akoumpa :: PR: #2700
- fix(deepseek_v4): support packed THD with context parallel by @HuiyingLi :: PR: #2731
- feat(speculative): FSDP2-shard dense DSpark target + add Gemma4-31B config by @khazic :: PR: #2976
- fix(attention): initialize FlexAttention masks by @akoumpa :: PR: #2982
- fix(deepseek-v4): initialize random training state by @akoumpa :: PR: #2991
- ci(review): strengthen model review invariants by @akoumpa :: PR: #2994
- ci: Bump claude_review to v1.8.2 by @chtruong814 :: PR: #3012
- ci: AUT-779 repin _claude_review.yml to v1.8.3 by @svcnemo-autobot :: PR: #3013
- test(ci): accept SHA-pinned claude-review template refs in the policy test by @khazic :: PR: #3022
- feat(speculative): context parallelism for draft-model training by @kashif :: PR: #2918
- feat(pp): support FusedLinearCrossEntropy under pipeline parallelism by @beccohov :: PR: #2927
- feat(speculative): support sequence packing in the DeepSeek MLA EAGLE-3 draft by @khazic :: PR: #3002
- fix(datasets): reject non-prefix multiturn mask templates by @khazic :: PR: #3024
- feat(speculative): support sequence packing in the EAGLE-1/2 draft by @khazic :: PR: #3003
- feat(speculative): add variable-prefix DFlash and LK EAGLE-3 training losses by @khazic :: PR: #3032
- feat(dspark): log acceptance and confidence metrics during training by @kashif :: PR: #2957
- feat(speculative): support sequence packing in the DFlash draft by @khazic :: PR: #3004
- feat(speculative): add periodic real accept-length eval by @khazic :: PR: #3037
- feat(speculative): support sequence packing in the Domino and JetSpec drafts by @khazic :: PR: #3023
- feat(speculative): support sequence packing in the DSpark Qwen3 draft by @khazic :: PR: #3005
- fix(models): temporarily disable TE fused RoPE globally by @HuiyingLi :: PR: #3028
- fix(glm_moe_dsa): clamp DSA indexer top-k for short sequences by @HuiyingLi :: PR: #3035
- feat(eagle3): on-policy regeneration loop with step-cadence dataloader swap by @khazic :: PR: #3042
- fix(docs): update technical accuracy by @akoumpa :: PR: #3031
- fix(checkpoint): compare torch versions semantically by @HuiyingLi :: PR: #3043
- docs: add uv run prefix by @akoumpa :: PR: #3057
- fix(recipe): add generation-marked chat templates by @khazic :: PR: #3051
- ci: include cutlass deps by @akoumpa :: PR: #3056
- fix(dspark): omit unmeasured positional acceptance rates by @khazic :: PR: #3050
- test(transformers): add source-load parity coverage by @yuhezhang-ai :: PR: #2960
- fix(distributed): resolve VLM AC layer-group paths structurally, not by version by @yuhezhang-ai :: PR: #3025
- fix(vlm): load Tulu-3 directly from HF Hub instead of local meta JSON by @athitten :: PR: #3001
- fix(optim): drop zero-numel local DTensor shards before TE FusedAdam by @yuhezhang-ai :: PR: #2997
- fix(moe): checkpoint trainable vision towers on the expert-parallel path by @yuhezhang-ai :: PR: #2993
- feat(datasets): add optional parallel pre-tokenization before packing by @khazic :: PR: #3045
- fix(models): handle checkpoint-wrapped fp32 buffers by @akoumpa :: PR: #3059
- feat(optim): support per-parameter-group learning rate by @khazic :: PR: #3046
- fix(cp): preserve gradients when sharding VLM inputs by @HuiyingLi :: PR: #2931
- refactor(vlm): gate build_model via recipe-side model-target allowlist by @athitten :: PR: #2357
- fix(ci): use HybridEP for Step 3.5 benchmark by @HuiyingLi :: PR: #3069
- fix: Kimi K2 config loading without remote code by @HuiyingLi :: PR: #3065
- feat(speculative): add Gemma4 EAGLE-3 target support by @khazic :: PR: #3071
- fix(dllm): remove obsolete DiffusionGemma HF compatibility by @zyzhou5 :: PR: #3067
- feat(speculative): add Gemma4-E4B EAGLE-3 example config by @khazic :: PR: #3073
- fix(test): use SDPA for Nemotron-H HF reload by @yuhezhang-ai :: PR: #3060
- feat(speculative): add Gemma4-31B EAGLE-3 example config by @khazic :: PR: #3077
- feat(speculative): add Gemma4-26B-A4B MoE EAGLE-3 example config by @khazic :: PR: #3079
- chore(ci): AUT-852 bump claude review template to v1.8.4 by @svcnemo-autobot :: PR: #3087
- feat: add Ministral3 embedding distillation recipe by @vinay-raman :: PR: #3058
- fix(perf): use GPT-OSS head dimension in FLOPs accounting by @yaoyu-33 :: PR: #3091
- test(speculative): add EAGLE-3 fp8 draft convergence smoke for SM89+ by @khazic :: PR: #3081
- fix(speculative): keep DSpark target_layer_ids in [0, N-2] for SGLang by @khazic :: PR: #3083
- chore(skills): add Regent Open Plugin manifest by @ko3n1g :: PR: #3097
- chore(skills): remove Open Plugin manifest (superseded) by @ko3n1g :: PR: #3099
- test: clean up KD test process group by @yuhezhang-ai :: PR: #3090
- refactor(models): single tie_word_embeddings guard via TieSupport by @Achyuthan-S :: PR: #2998
- feat(speculative): add DFlash validation metrics by @khazic :: PR: #3072
- refactor(datasets): typed Config + build per dataset, config-driven dataloader by @akoumpa :: PR: #2390
- fix(bagel): enable periodic garbage collection by @zyzhou5 :: PR: #3105
- fix(ci): pass PR number and commit to Codecov for fork PRs by @svcnemo-autobot :: PR: #3101
- fix(test): use eager attention for Nemotron-H HF loads by @yuhezhang-ai :: PR: #3100
- feat(kd): support separate student and teacher meshes by @akoumpa :: PR: #2954
- fix(moe): safe TP and EP/TP gradient correctness for custom MoE models by @yuhezhang-ai :: PR: #2995
- test(ci): remove deprecated 26.10 models from nightly and release CI by @athitten :: PR: #2884
- feat(training): add opt-in setup-time prewarms by @yuhezhang-ai :: PR: #2992
- fix: bound chunked CE memory and preserve packing masks by @yuhezhang-ai :: PR: #2996
- fix(ci): restore Ministral3 checkpoint robustness by @yuhezhang-ai :: PR: #3111
- ci: add HF hub cache preflight check by @thomasdhc :: PR: #3119
- fix(vlm): resolve get_rope_index from the base model for packed mRoPE by @khazic :: PR: #3113
- feat(models): add Inkling VLM MoE support by @hemildesai :: PR: #3095
- ci(dllm): add dLLM SFT nightly train-to-generate launcher and recipes by @zyzhou5 :: PR: #2783
- feat(vlm): add THD packed-sequence support for Qwen3-VL-MoE by @khazic :: PR: #3052
- docs: Add Inkling README news by @HuiyingLi :: PR: #3129
- fix(checkpoint): gather PEFT adapter across PP stages by @hyfine :: PR: #3096
- build(deps): bump base container to 26.06-cuda13.3 by @thomasdhc :: PR: #2983
- fix(transformers): keep custom M3 config when transformers ships its own by @HuiyingLi :: PR: #3134
- ci: AUT-911 shard GPU unit tests into 5 pytest-shard chunks by @svcnemo-autobot :: PR: #3138
- fix(checkpoint): isolate consolidation timeout from NCCL by @yuhezhang-ai :: PR: #3108
- test(vlm): add checkpoint robustness coverage by @yuhezhang-ai :: PR: #3112
- fix(ci): preserve retrieval evaluation schedule by @yuhezhang-ai :: PR: #3139
- feat(bagel): add TE support by @zyzhou5 :: PR: #2895
- fix(moe): preserve HSDP replica gradient synchronization by @wangzhxg :: PR: #3135
- docs: add Laguna SFT docs and recipe by @HuiyingLi :: PR: #3146
- ci: AUT-895 gate GB200 tests on DISABLE_GB200_TESTS variable by @svcnemo-autobot :: PR: #3118
- feat(dllm): add LLaDA2 generation support by @zyzhou5 :: PR: #3092
- feat(models): add Laguna model implementation by @HuiyingLi :: PR: #3148
- fix: prevent damaged token embeddings from dominating grad clipping by @akoumpa :: PR: #3136
- fix(checkpoint): add bounded retention window by @oliverholworthy :: PR: #2416
- fix(kd): use fp32 master weight copy by @akoumpa :: PR: #3019
- fix(checkpoint): restore adapters through DDP wrappers by @yuhezhang-ai :: PR: #3150
- feat(distributed): block-diagonal varlen CP for packed sequences by @yuhezhang-ai :: PR: #2989
- fix(distributed): preserve canonical activation checkpoint keys by @yuhezhang-ai :: PR: #3152
- feat(examples): Gemma4-31B CoderForge data pipeline + CP SFT recipes by @athitten :: PR: #3151
- feat(attention): support flash_attention_3 and flash_attention_4 by @HuiyingLi :: PR: #2929
- feat(checkpoint): add DCP CPU offload option by @khazic :: PR: #3130
- test(packing): fix stale unsupported-backend test after fa3/fa4 support by @khazic :: PR: #3199
- test(ci): add DeepSeek V4 Flash pretrain coverage by @akoumpa :: PR: #3128
- refactor(distributed): unify CP input prep and dispatch across models by @HuiyingLi :: PR: #2937
- docs(fern): add legacy URL redirects by @akoumpa :: PR: #3203
- ci: reduce L0 GPU tests to two shards by @akoumpa :: PR: #3201
- feat: add checkpoint staging wait option by @khazic :: PR: #3131
- fix(datasets): keep system turns in the sharegpt conversation converter by @khazic :: PR: #3144
- fix(peft): honor memory-efficient LoRA opt-out by @akoumpa :: PR: #3126
- feat: enable text inclusion alongside images in retrieval training by @rnyak :: PR: #3189
- feat: Support preemption checkpointing by @edjson :: PR: #3007
- fix(ci): AUT-964 restore Codecov after skipped GB200 jobs by @svcnemo-autobot :: PR: #3205
- fix(ci): preserve HF meta init for device-mapped loads by @yuhezhang-ai :: PR: #3188
- feat(dllm): add lora to dllm by @zyzhou5 :: PR: #3163
- fix(gemma4): run SDPA in fp32 to avoid #2208 NaN on Hopper by @aminehd :: PR: #3141
- docs(retrieval): complete fine-tuning guide integration by @oliverholworthy :: PR: #2306
- fix(glm): apply LoRA to TileLang MLA KV projection by @shahafwa :: PR: #3176
- fix: prefer Automodel config registry lookup by @akoumpa :: PR: #3202
- feat(distributed): shape TPLinear/LinearLoRA graphs for async-TP fusion by @yuhezhang-ai :: PR: #2987
- fix(distributed): Megatron-FSDP 0.5.0 compatibility by @yuhezhang-ai :: PR: #2986
- feat(checkpoint): preserve intrinsic fp32 during offline consolidation by @yuhezhang-ai :: PR: #3193
- feat(kernels): add QuACK backend by @akoumpa :: PR: #3115
- feat(dllm): add DiffusionGemma generation via the built-in HF sampler by @zyzhou5 :: PR: #3161
- refactor(diffusion): migrate recipe onto typed RecipeConfig build() path by @pthombre :: PR: #3122
- fix(dist): keep profiler record-function ops out of SAC replay accounting by @HuiyingLi :: PR: #3133
- fix(vlm): size qwen3.6 medpix CI configs by @HuiyingLi :: PR: #3209
- feat(glm_moe_dsa): expose update_moe_gate_bias on GLM MoE DSA models by @jQizhang :: PR: #3207
- fix(ci): shard Nemotron Super vLLM deploy across 8 GPUs by @yuhezhang-ai :: PR: #3061
- fix(glm): size GLM5.2 release CI jobs by @HuiyingLi :: PR: #3210
- fix(distributed): activation-checkpoint Qwen3-Next linear_attn layers by @amolkhanna :: PR: #3192
- fix(distributed): support packed CP for Llama, Qwen2, and Qwen3 by @akoumpa :: PR: #2999
- fix(dllm): recipe defaults, corruption seeding, sampler and docs fixes by @zyzhou5 :: PR: #3162
- perf(bagel): grouped MoT routing + fused SwiGLU/RoPE by @zyzhou5 :: PR: #3214
- feat(retrieval): add normalized Arrow dataset tooling by @yuhezhang-ai :: PR: #2596
- feat(diffusion): context parallelism via diffusers ContextParallelConfig by @pthombre :: PR: #3157
- docs: explain embedding row repair by @akoumpa :: PR: #3175
- fix(transformers): custom configs override builtin ones by default by @HuiyingLi :: PR: #3221
- build(deps): bump ffpa-attn to 0.2.2 by @akoumpa :: PR: #3225
- fix(distributed): checkpoint linear attention with compile by @yuhezhang-ai :: PR: #3213
- fix(glm): own the GLM MoE DSA config so qk_rope_head_dim survives by @HuiyingLi :: PR: #3222
- feat(speculative): add the multimodal speculative decoding (MSD) core by @khazic :: PR: #3166
- fix(speculative): average DFlash train metrics over the log window by @khazic :: PR: #3234
- fix(speculative): honor --dflash-causal on the vLLM remapping export by @khazic :: PR: #3235
- fix(speculative): align the decode-eval cadence on resume by @khazic :: PR: #3237
- fix(speculative): validate mask_token_id when resuming DSpark by @khazic :: PR: #3238
- build: cache CUDA extension source builds by @akoumpa :: PR: #3204
- docs(fern): cut v0.5.0 version train by @lbliii :: PR: #2970
- fix(checkpoint): infer Qwen MTP layout from checkpoint keys by @HuiyingLi :: PR: #3229
- docs: fix remaining broken link sources by @akoumpa :: PR: #3249
- fix(distributed): control frozen multimodal FSDP sharding by @yuhezhang-ai :: PR: #2763
- fix(ci): preserve main wheelhouse cache builds by @akoumpa :: PR: #3253
- feat(diffusion): add Qwen-Image-Edit-2511 training support by @pthombre :: PR: #3217
- fix(ci): omit token type IDs from vLLM parity prompts by @yuhezhang-ai :: PR: #3198
- fix(deepseek-v4): accept cu_seqlens for THD packing by @HuiyingLi :: PR: #3231
- fix(deepseek-v4): preserve fp32 model dtype by @HuiyingLi :: PR: #3227
- fix(thd): derive packed padding mask from the pack layout, not token ids by @HuiyingLi :: PR: #3223
- fix(nemotron-v3): recompute deterministic LoRA router by @HuiyingLi :: PR: #3258
- fix(moe): preserve top-k routing under full AC by @yuhezhang-ai :: PR: #3140
- fix(gpt-oss): route packed attention through THD by @akoumpa :: PR: #3226
- perf(benchmark): tune Qwen3-VL LoRA pipeline batches by @akoumpa :: PR: #3230
- fix(moe): preserve gate load across activation recompute by @akoumpa :: PR: #3247
- fix(ci): narrow CUDA wheelhouse cache keys by @akoumpa :: PR: #3269
- feat(distributed): frame-level context-parallel vision-tower sharding by @yuhezhang-ai :: PR: #2990
- fix(vlm): preserve mRoPE axes in PP chunking by @HuiyingLi :: PR: #3208
- feat(diffusion): add LTX-2.3 video+audio finetuning (training + prepr… by @linnanwang :: PR: #3165
- fix: harden checkpoint torch loads by @HuiyingLi :: PR: #3240
- build: add MagiAttention optional dependency by @HuiyingLi :: PR: #3070
- fix(checkpoint): harden PP safetensors consolidation by @yuhezhang-ai :: PR: #3248
- fix(checkpoint): save all PEFT EP optimizer parts by @yuhezhang-ai :: PR: #3250
- fix(ci): install CUDA torchvision with CUDA torch by @akoumpa :: PR: #3279
- fix(docker): cap uv install concurrency to avoid ARM FD exhaustion by @thomasdhc :: PR: #3275
- feat: support pre-extracted frame sequence of video sample by @GITsologun :: PR: #3211
- fix(inkling): support HF 5.14 attention fields and embed norm FSDP by @HuiyingLi :: PR: #3281
- fix(KD): Fix variable-length teacher PP batches on separate meshes by @HuiyingLi :: PR: #3280
- feat: add Kimi K3 model support by @HuiyingLi :: PR: #3259
- docs(models): add Kimi K3 coverage by @HuiyingLi :: PR: #3283
- test(vlm): stabilize Qwen3.5 checkpoint robustness by @HuiyingLi :: PR: #3233
- feat(vlm): integrate packed CP and vision sharding for Qwen3.5-MoE by @yuhezhang-ai :: PR: #3186
- feat(checkpoint): defer distributed async consolidation by @khazic :: PR: #3125
- test(retrieval): add Nemotron VL checkpoint coverage by @yuhezhang-ai :: PR: #3276
- chore(ci): AUT-1135 pin GitHub Actions to commit SHAs by @svcnemo-autobot :: PR: #3277
- feat: add ViSpec VLM draft training by @khazic :: PR: #3173
- fix(docs): hide Kimi tokenizer regex from autodoc (3288) by @svcnvidia-nemo-ci :: PR: #3304
- docs(distributed): review frozen multimodal FSDP guidance (3272) by @svcnvidia-nemo-ci :: PR: #3306
- fix(vlm): expandable_segments:True for gemma4 31B FFPA 8k recipe (AMINT-203) (3273) by @svcnvidia-nemo-ci :: PR: #3307
- fix(pp): use static metadata with PyTorch 2.13 (3290) by @svcnvidia-nemo-ci :: PR: #3305
- ci: add Tulu-3 convergence + eval flow (3171) by @svcnvidia-nemo-ci :: PR: #3312
- fix(distributed): preserve ERNIE router and dense Qwen3.5 SSM precision (3255) by @svcnvidia-nemo-ci :: PR: #3310
- perf(moe): add scoped partial CUDA graphs (2917) by @svcnvidia-nemo-ci :: PR: #3314
- fix(docs): avoid literal ampersand in Kimi autodoc (3317) by @svcnvidia-nemo-ci :: PR: #3320
- fix(config): repair GLM-5.2 LoRA recipe (3313) by @svcnvidia-nemo-ci :: PR: #3323
- fix(kd): preserve tensor-valued hidden states (3324) by @svcnvidia-nemo-ci :: PR: #3332
- fix(kd): use mesh-safe gradient clipping (3302) by @svcnvidia-nemo-ci :: PR: #3334
- refactor(docker): build torchao & FlashAttention as isolated wheel stages (3329) by @svcnvidia-nemo-ci :: PR: #3343
- fix(distributed): restore Nemotron Flash TP2 training (3345) by @svcnvidia-nemo-ci :: PR: #3349
- fix(nemotron-parse): sync RADIO preprocessing (3331) by @svcnvidia-nemo-ci :: PR: #3341
- fix(distributed): reuse default group for world-sized meshes (3319) by @svcnvidia-nemo-ci :: PR: #3347
- fix(training): prewarm Mamba SSD autotune kernels (3296) by @svcnvidia-nemo-ci :: PR: #3338
- fix(docs): make Fern autodoc metadata MDX-safe (3355) by @akoumpa :: PR: #3357
- fix(pp): preserve VLM media cursor with static metadata (3344) by @svcnvidia-nemo-ci :: PR: #3361
- fix(deps): resolve 26.08 rc2 container CVEs (3346) by @svcnvidia-nemo-ci :: PR: #3363
- fix(pp): route Qwen3.5 MoE pre-embedded inputs (3294) by @svcnvidia-nemo-ci :: PR: #3364
- fix(kimi_k25_vl): don't int4 quantize LoRA adapter keys on save (3295) by @svcnvidia-nemo-ci :: PR: #3362
- fix(diffusion): use spawn start method for GPU preprocessing pools (3339) by @svcnvidia-nemo-ci :: PR: #3366
- fix(tokenizer): make special token insertion opt-in (3337) by @svcnvidia-nemo-ci :: PR: #3375
- fix(distributed): stabilize SAC replay with TE and FSDP (3330) by @svcnvidia-nemo-ci :: PR: #3376
- fix(ci): use a single container cache donor (3377) by @svcnvidia-nemo-ci :: PR: #3381
- fix(fsdp): resolve fp32 master-weight compute dtype per parameter (3328) by @svcnvidia-nemo-ci :: PR: #3384
- perf(checkpoint): reduce distributed save overhead (3369) by @svcnvidia-nemo-ci :: PR: #3389
- fix(ci): enable LTX-2.3 diffusion finetuning (3372) by @svcnvidia-nemo-ci :: PR: #3391
- ci: tune Nemotron single-GPU model load threads (3370) by @svcnvidia-nemo-ci :: PR: #3395
- fix(model): correct Mistral4 attention and distributed MoE routing (3348) by @svcnvidia-nemo-ci :: PR: #3393
- fix(checkpoint): survive an interrupted save instead of hanging all ranks (3261) by @svcnvidia-nemo-ci :: PR: #3394
- perf: remove Python overhead from model hot paths (3374) by @svcnvidia-nemo-ci :: PR: #3401
- fix(deps): resolve OSS CVE findings (3398) by @svcnvidia-nemo-ci :: PR: #3412
- fix(moe): support EP-free DTensor state-dict conversion (3397) by @svcnvidia-nemo-ci :: PR: #3409
- fix(checkpoint): harden PEFT PP checkpointing and Step-3.7 coverage (3316) by @svcnvidia-nemo-ci :: PR: #3403
- docs(training): mark Mamba prewarm sections for review (3340) by @svcnvidia-nemo-ci :: PR: #3399
- fix(retrieval): guard optional W&B imports (3380) by @svcnvidia-nemo-ci :: PR: #3402
- fix(gemma4): enable expandable_segments on the CP tulu3 recipes (3408) by @svcnvidia-nemo-ci :: PR: #3410
- feat: Add {% generation %} chat template for DiffusionGemma SFT/LoRA examples (3353) by @svcnvidia-nemo-ci :: PR: #3417
- fix(docs): AUT-1324 qualify LTX model coverage slug (3418) by @svcnvidia-nemo-ci :: PR: #3419
- fix(ci): run Kimi K3 HellaSwag on GB200 (3420) by @svcnvidia-nemo-ci :: PR: #3421
- fix(transformers): default missing THD capability (3406) by @svcnvidia-nemo-ci :: PR: #3429
- ci(convergence): (3416) by @svcnvidia-nemo-ci :: PR: #3432
- cp: backport MoE checkpoint reload parity fixes to r0.6.0 by @yuhezhang-ai :: PR: #3438
- fix(retrieval): use stock Ministral embedding backbone (3103) by @svcnvidia-nemo-ci :: PR: #3433
- ci(minimax): extend M2.7 LoRA timeout (3428) by @svcnvidia-nemo-ci :: PR: #3430
- fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (3411) by @svcnvidia-nemo-ci :: PR: #3445
- fix(kimi_k25_vl): keep PEFT expert LoRA keys under the base_model.model. prefix (3431) by @svcnvidia-nemo-ci :: PR: #3451
- ci(vlm): raise Slurm wall time for MiniMax-M3 and Gemma4 recipes (3448) by @svcnvidia-nemo-ci :: PR: #3453
- fix: Gemma4-31B CoderForge CP8/64K data processing and recipe (3446) by @svcnvidia-nemo-ci :: PR: #3452
- cp: backport PEFT v5 adapter output fixes to r0.6.0 by @yuhezhang-ai :: PR: #3458
- fix(retrieval): repair Ministral3 training recipe (3441) by @svcnvidia-nemo-ci :: PR: #3454
- test(checkpoint): redesign resume robustness around a shared trajectory (3427) by @svcnvidia-nemo-ci :: PR: #3457
- test(checkpoint): stabilize MoE checkpoint parity gates (3440) by @svcnvidia-nemo-ci :: PR: #3460
- fix(gemma4): exact router scalar and fp32 reference routing for MoE parity (3456) by @svcnvidia-nemo-ci :: PR: #3466
- cp: backport Sentence Transformers metadata export to r0.6.0 by @yuhezhang-ai :: PR: #3464
- fix: normalize MoE auxiliary loss during gradient accumulation (3359) by @svcnvidia-nemo-ci :: PR: #3459
- fix(bagel): type SFT backend before AutoModel init (3449) by @svcnvidia-nemo-ci :: PR: #3469
- fix(peft): enable v5 expert adapters for Nemotron and MiniMax (3439) by @svcnvidia-nemo-ci :: PR: #3468
- fix(checkpoint): save ranked RNG state per global rank (3437) by @svcnvidia-nemo-ci :: PR: #3471
- fix(peft): align frozen LoRA tensors with compute dtype (3470) by @svcnvidia-nemo-ci :: PR: #3480
- cp: backport VLM CP gradient and vision sharding fixes to r0.6.0 by @yuhezhang-ai :: PR: #3479
- fix(training): all-reduce grad-norm scalars on the mesh device (3461) by @svcnvidia-nemo-ci :: PR: #3481
- fix(recipe): shard GPT-OSS 120B across 64 experts (3483) by @svcnvidia-nemo-ci :: PR: #3489
- fix(ci): register flux2/wan2.2/qwen-image-edit diffusion recipes in CI (3485) by @svcnvidia-nemo-ci :: PR: #3487
- perf(benchmarks): recompute deterministic MoE routers under AC (3474) by @svcnvidia-nemo-ci :: PR: #3490
- chore: Bump gitpython to >= 3.1.59 (3482) by @svcnvidia-nemo-ci :: PR: #3488
- perf(recipes): run GPT-OSS 120B with EP64 and no activation checkpointing by @akoumpa :: PR: #3492
- cp: feat(models): make Inkling standalone (#3358) into r0.6.0 by @HuiyingLi :: PR: #3498
- fix(models): preserve lm-head dtype boundaries (3491) by @svcnvidia-nemo-ci :: PR: #3502
- fix(checkpoint): validate PEFT adapter-only state (3501) by @svcnvidia-nemo-ci :: PR: #3503
- ci(convergence): fix gemma4 eval setup and re-baseline Qwen3-MoE (3493) by @svcnvidia-nemo-ci :: PR: #3505
- fix(deepseek_v4): avoid TileLang boolx8 backward codegen (3467) by @svcnvidia-nemo-ci :: PR: #3506
- fix(checkpoint): preserve FSDP2 mixed precision during recompute (3513) by @svcnvidia-nemo-ci :: PR: #3515
- fix(deps): resolve Starlette and GitPython CVEs (3523) by @svcnvidia-nemo-ci :: PR: #3524
- fix(minimax): declare packed-sequence and CP attention ownership by @athitten :: PR: #3548
- feat(gemma4): add E-series tensor parallelism (3512) by @svcnvidia-nemo-ci :: PR: #3552
- fix(retrieval): support canonical Sentence Transformers metadata (3546) by @svcnvidia-nemo-ci :: PR: #3551
- ci: route DGX Spark recipes to GB10 (3538) by @svcnvidia-nemo-ci :: PR: #3556
- fix(docker): build bitsandbytes for SM121 (3553) by @svcnvidia-nemo-ci :: PR: #3557
- cp: backport packed THD VLM context parallelism to r0.6.0 by @qiaochuz-nv :: PR: #3558
- fix(fsdp): uniform reduce dtype and EP-local expert gradients (3540) by @svcnvidia-nemo-ci :: PR: #3561
- fix(deps): resolve msgpack, wandb, and mistune container CVEs (3563) by @svcnvidia-nemo-ci :: PR: #3566
- fix(deps): resolve 26.08 rc9 container CVEs (3607) by @thomasdhc :: PR: #3611
- fix(deps): resolve 26.08 rc10 container CVEs (3647) by @thomasdhc :: PR: #3658
- fix(peft): support mixed-dtype memory-efficient LoRA backward (3675) by @svcnvidia-nemo-ci :: PR: #3682
- beep boop 🤖: Bumping NeMo-Automodel to v0.5.1 by @nemo-automation-bot[bot] :: PR: #3687
- fix(release): set version to 0.6.0 by @thomasdhc :: PR: #3692
- ci: update package version to 0.5.0 (#2472) by @thomasdhc
- feat: make mesh accept meshcontext (#2266) by @adil-a
- docs(contributing): fix perk typo (#2471) by @grgkovac
- feat(vlm): enable Qwen3.5 MoE VLM CP (#2432) by @HuiyingLi
- ci: ask claude-review to flag stale ModelCapabilities (#2480) by @athitten
- feat(speculative): add activation checkpointing for P-EAGLE draft layers (#2458) by @khazic
- fix(speculative): release the EAGLE-3 target/W&B on any training exit (#2460) by @khazic
- feat(model): flux2 (#2145) by @linnanwang
- refactor(speculative): share grad-accum + LR utils across EAGLE/DFlash recipes (#2463) by @khazic
- fix(speculative): require an explicit DFlash mask_token_id (#2461) by @khazic
- fix(speculative): re-sync the target vocab mapping on EAGLE-3 resume (#2467) by @khazic
- fix(speculative): drop position 0 from P-EAGLE depth-1 candidate pool (#2464) by @khazic
- fix(speculative): finalize the last async checkpoint on training exit (#2469) by @khazic
- fix(speculative): make regenerate clobber guard see existing shards (#2477) by @khazic
- fix(speculative): validate config on precompute_eagle3 --resume (#2478) by @khazic
- fix(speculative): keep DFlash DDP ranks in lockstep on no-valid-anchor skips (#2468) by @khazic
- fix(speculative): serialize remote EAGLE-3 /generate end to end (#2479) by @khazic
- feat(speculative): add EAGLE feature-noise augmentation to EAGLE-1/2 training (#2470) by @khazic
- fix(speculative): gate unsupported P-EAGLE backend / cache / packing combos (#2466) by @khazic
- fix(dflash): skip short validation micro-batches in eval loop (#2476) by @khazic
- fix(speculative): retry the remote EAGLE-3 /generate wire path (#2459) by @khazic
- docs(dllm): add DiffusionGemma 26B-A4B SFT/LoRA guide (#2504) by @zyzhou5
- feat(dllm): add DiffusionGemma 26B-A4B block-diffusion SFT (full + LoRA) (#2506) by @zyzhou5
- feat: add MSC cloud storage support for dcp checkpoints (#1709) by @edjson
- feat(checkpoint): export torch_save checkpoints to hf (#2487) by @khazic
- test(checkpoint): fix TestFormatLoad directory-based reader selection (#2520) by @khazic
- fix(gemma4): cast dense params without casting buffers (#2359) by @DOGEUNNKIM
- fix(config): glm4.7 yaml (#2527) by @akoumpa
- docs(fern): relocate Fern under docs/ and remove legacy Sphinx tree (#2391) by @lbliii
- fix: unwrap ModelOutput to extract logits (#2523) by @akoumpa
- feat(speculative): add tqdm training progress bar to EAGLE and DFlash recipes (#2522) by @khazic
- feat(speculative): add context parallelism for the EAGLE-3 target model (#2465) by @khazic
- fix(speculative): compile variable-shape flex callables with dynamic shapes (#2532) by @khazic
- fix(checkpoint): map BOOL in the safetensors backport DTYPE_MAP (#2533) by @khazic
- docs: add MiniMax M3 VL user guide (#2536) by @athitten
- feat: add MiniMax M3 VL (#2538) by @athitten
- fix(docs): restore old .md (#2540) by @akoumpa
- docs(fern): point restored guide pages to .md in nightly nav (#2541) by @lbliii
- docs(dllm): use absolute image URLs in the DiffusionGemma guide (#2543) by @zyzhou5
- fix(fern): generate autodoc library reference before publish (#2537) by @lbliii
- fix(docs): restore previous md (#2544) by @akoumpa
- feat(examples): add Nemotron-3-Ultra-550B benchmark and full-SFT recipes (#2539) by @adil-a
- feat: Add MagiAttention (FFA / context-parallel) attention backend (#2384) by @HuiyingLi
- fix(qwen3_5): make dense VLM pipeline-parallel safe (#2524) by @HuiyingLi
- feat(vlm): enable Qwen3.5 dense VLM CP (#2505) by @HuiyingLi
- fix(models): keep RoPE frequency buffers fp32 under bf16 model cast (#2549) by @akoumpa
- ci: schedule ep-parallel finetune recipes at documented node counts (#2546) by @akoumpa
- feat: add gemma4 moe ring CP (#1914) by @khazic
- ci(fern): run docs check on every /ok to test (drop paths filter) (#2547) by @akoumpa
- feat(qwen3_5): port dense Qwen3.5 to a native custom-model implementation (#2557) by @HuiyingLi
- ci: flag RoPE/precision-buffer dtype hazards in automated PR review (#2552) by @akoumpa
- fix(diffusion): resolve flux nightly CI failures (#2529) by @pthombre
- docs: announce Gemma diffusion support (#2568) by @pthombre
- docs: Update diffusiongemma.mdx (#2542) by @zyzhou5
- feat(moe): MTP FLOPs accounting, inline shared experts, checkpoint warning fix (#2486) by @adil-a
- fix(checkpoint): preserve tied lm_head on resume (#2511) by @yuhezhang-ai
- test: fix all 5 vllm_deploy tests (token drift, nemotron OOM + mamba merge) (#2559) by @adil-a
- fix(ci): bump ling_1t_lora_pp local_batch_size to satisfy PP assert (#2575) by @akoumpa
- fix(ci): set node counts for multi-node VLM finetune recipes (#2574) by @akoumpa
- fix(recipe): reshard MoE experts after forward in nemotron_nano_v3_cp_test (#2577) by @akoumpa
- ci: use digits for spark recipes (#2581) by @akoumpa
- ci: Enable activation checkpointing for gemma_2_9b_it_squad (AM-464) (#2585) by @akoumpa
- fix(test): load checkpoint-robustness HF reference via device_map (#2582) by @akoumpa
- fix(distributed): register Falcon-H1 TP plan to fix 34B PEFT OOM (#2589) by @akoumpa
- feat(mtp): enable MTP to accept pre-fused input embeddings for multimodal models (#2510) by @Slyne
- fix(peft): LoRA MLP QLoRA/PP/gemma3n fixes (AM-435, AM-447, AM-453) (#2584) by @akoumpa
- fix(gemma4): FSDP2-safe kv-sharing + skip frozen audio tower on grad-accum (#2566) by @athitten
- fix(vlm): enable activation checkpointing for 35B Qwen3.5/3.6 VLM recipes (#2600) by @akoumpa
- fix(vlm): use FusedLinearCrossEntropy for qwen3_5_9b to avoid logits OOM (#2603) by @akoumpa
- perf(distributed): add retrieval tuning knobs (#2452) by @yuhezhang-ai
- fix(moe): preserve fp32 A_log in Qwen3.5-MoE and Qwen3-Next GatedDeltaNet (#2484) by @yuhezhang-ai
- fix(moe): weight GroupedExpertsTE down-projection bias by routing probability (#2591) by @akoumpa
- fix: use TE attention for gpt_oss packed-sequence recipe (AM-438) (#2587) by @akoumpa
- fix(oom): use FusedLinearCrossEntropy in qwen3 tulu3 configs to avoid OOM (#2609) by @akoumpa
- fix(bagel): distributed setup init (#2608) by @zyzhou5
- feat(deepseek-v4): support context parallel training (#2590) by @HuiyingLi
- fix(qwen3_5_moe): convert MTP experts as grouped tensors (AM-442) (#2595) by @HuiyingLi
- fix(transformers): keep gemma3n KV sharing working under FSDP2 (AM-454) (#2594) by @HuiyingLi
- feat(vlm): add Gemma 4 31B joint drafter example config (#2620) by @khazic
- feat(datasets): add cp=1 shared-prefix prefix-tree attention for rollouts (#2564) by @khazic
- feat(gemma4): context parallelism for dense 31B (#2592) by @HuiyingLi
- fix(qwen3_moe): keep native forward under PP so CP+THD works (#2625) by @akoumpa
- fix(docker): build DeepEP against the NVSHMEM wheel matching the apt runtime (#2614) by @akoumpa
- fix(config): require pp_size and distributed.pipeline to agree (#2616) by @akoumpa
- feat(gemma4): context parallelism for dense E2B/E4B (#2621) by @HuiyingLi
- fix(moe): default ignore_router_for_ac=True for activation checkpointing (#2635) by @akoumpa
- ci: raise ci.time for slow finetune recipes hitting 10-min default (#2637) by @akoumpa
- ci: cap MAX_STEPS to 10 for slow vlm_finetune recipes (#2639) by @akoumpa
- fix(examples): enable activation checkpointing for phi_4_squad (#2634) by @akoumpa
- fix(parallelizer): resolve NemotronH decoder blocks for Nemotron-V3 (#2638) by @akoumpa
- refactor(moe): remove enable_deepep, switch failing ep recipes to hybridep (#2630) by @akoumpa
- test: run slow CP unit-test files GPU-only (gemma4, deepseek_v4) (#2647) by @akoumpa
- fix(datasets): decode MedPix images on demand instead of up front (#2645) by @akoumpa
- feat(config): add wandb.enable flag; disable W&B in example configs (#2643) by @akoumpa
- fix(datasets): nest processing kwargs under processor_kwargs in default_collate_fn (#2649) by @akoumpa
- fix(devstral2,ministral3): load FP8 checkpoints via custom mistral3_vlm path (drop HF FineGrainedFP8) (#2654) by @akoumpa
- fix(merge_lora): save tokenizer faithfully via NeMoAutoTokenizer (#2653) by @akoumpa
- fix(distributed): make fully_shard_by_dtype produce storage-uniform FSDP2 groups (#2655) by @akoumpa
- fix(models): default yarn original_max_position_embeddings in Step3p7Config (#2652) by @akoumpa
- fix(model): apply dict config overrides with custom dispatch (#2657) by @akoumpa
- fix(recipe): disable fused RoPE for MLA packed-sequence MoE recipes (#2675) by @akoumpa
- fix(training): flag final step when epochs exhaust before max_steps (#2672) by @akoumpa
- fix(docker): bump DeepEP to
42144303to pad HybridEP token capacity (#2678) by @akoumpa - fix(llama3_3): reduce pp_microbatch_size to fit 70b squad on h100 (#2673) by @akoumpa
- fix(checkpoint): fp8 MoE expert weights silently not loaded at ep_shard=1 (random init → garbage loss) (#2682) by @akoumpa
- fix(qwen3_moe): disable fused RoPE for the packed-sequence LoRA recipe (#2687) by @akoumpa
- feat(glm_moe_dsa): GLM-5.2 IndexShare DSA support (#2633) by @HuiyingLi
- fix(vlm): bump mistral3p5_128b_medpix max_length 1024->2048 (#2689) by @akoumpa
- fix(moe): free DeepEP buffer before process-group destroy to avoid hang (#2686) by @akoumpa
- test(models): speed up qwen3.5 moe/vl-moe from_pretrained unit tests (#2698) by @akoumpa
- perf(checkpoint): mmap HF DCP read_data to avoid host-RAM OOM on large loads (#2690) by @akoumpa
- fix(mistral3): remap FP8 VLM checkpoint prefixes (#2692) by @akoumpa
- fix(loss): reuse LM head gather for MTP loss (#2694) by @akoumpa
- fix(gemma4_moe): re-tie lm_head to active embed_tokens on MoE path (#2601) by @Achyuthan-S
- ci: fix 26.06 release cves (#2705) by @thomasdhc
- fix(moe): handle non-EP expert weight DTensors (#2697) by @akoumpa
- ci: add cluster_tag to gb200 benchmarks (#2714) by @thomasdhc
- build: install tilelang + tile_kernels for DeepSeek-V4 recipes (#2683) by @akoumpa
- fix(checkpoint): super-49B consolidated reload and vllm_deploy (#2626) by @adil-a
- fix(models): use bool sparse masks for sdpa (#2624) by @yuhezhang-ai
- fix(distributed): register DeciLM Nemotron TP plan (#2703) by @akoumpa
- ci: add time budgets for 12 new timeout failures (#2707) by @thomasdhc
- fix(qwen3_5): handle packed MTP attention (#2727) by @akoumpa
- docs: update container references to 26.06 and fix mount instructions (#2716) by @adil-a
- perf(DSV4): use generic checkpoint wrapper for activation checkpointing (#2704) by @HuiyingLi
- feat(glm_moe_dsa): add TileLang DSA kernels (#2691) by @HuiyingLi
- ci: add gb200 cluster specification for nemotron_ultra recipe (#2733) by @thomasdhc
- fix(ci): pin qwen3_moe_30b mxfp8 finetune to gb200 (#2735) by @thomasdhc
- ci: address diffusers cve (#2706) by @thomasdhc
- feat(glm_moe_dsa): add GLM5.2 context parallel support (#2695) by @HuiyingLi
- ci: address thrift cve bump to 0.23.0 (#2736) by @thomasdhc
- fix(deepseek-v4): restore batch axis for packed-sequence (THD) forward (#2651) by @akoumpa
- fix(deepseek-v4): avoid bf16 -inf overflow in additive attention mask (#2658) by @akoumpa
- build: install TileKernels for DeepSeek V4 (#2740) by @akoumpa
- fix(ci): reduce mixtral release smoke batch (#2728) by @akoumpa
- ci: Add gemma4 e4b to nightly test (#2749) by @athitten
- fix(models): audit fp32 protected tensors (#2598) by @yuhezhang-ai
- fix(fsdp2): guard uninitialized accumulated grads (#2744) by @akoumpa
- fix(qwen3_moe): step-0 NaN in MXFP8 packed finetune — expert unload + fused RoPE (#2722) by @hemildesai
- fix(diffusion): reuse warm HF cache instead of re-downloading models (#2747) by @pthombre
- fix(diffusion): raise qwen-image dist timeout for checkpoint consolidation (#2748) by @pthombre
- ci: bump benchmark glm_4.7_flash_te_deepep time (#2757) by @thomasdhc
- fix(mistral3): preserve medium VLM checkpoint layout (#2758) by @akoumpa
- ci: avoid remote config load for Nemotron Nano test (#2764) by @akoumpa
- fix(sdpa): apply resolved backend constraints to custom models (#2761) by @akoumpa
- fix(wandb): log different val datasets separately in wandb (#2526) by @grgkovac
- docs(fern): document DistributedSetup Python API for NeMoAutoModel loaders (#2766) by @akoumpa
- fix(distributed): use flattened CP FSDP mesh (#2768) by @HuiyingLi
- fix: Remove dali from container (#2770) by @chtruong814
- ci: run dsv32_lora, kimi_k2 and qwen3_moe_235b deepep benchmarks online (#2773) by @thomasdhc
- fix: skip fused LoRA MLP install for meta weights (#2775) by @akoumpa
- fix(checkpoint): resolve tie_word_embeddings top-level-first to match HF tying (#2732) by @Achyuthan-S
- fix(benchmark): skip unsupported MTP flops (#2767) by @akoumpa
- fix: Qwen3.5 MedPix EP32 NCCL timeout (#2777) by @akoumpa
- fix(ci): address go-git/go-billy and rustls-webpki CVEs (#2780) by @thomasdhc
- fix(ci): stabilize diffusion finetune smoke tests (#2788) by @pthombre
- build(deps): move ffmpeg/opencv deps to opt-in media extra (#2743) by @thomasdhc
- fix(vlm): keep Qwen3.5 media tokens aligned (#2772) by @yuhezhang-ai
- fix(ci): HybridEP for multi-node MoE benchmarks + LoRA OOM fixes (#2789) by @hemildesai
- fix: qwen3.5 and 3.6 mtp expert checkpoint layout (#2778) by @HuiyingLi
- feat: CP support for MiniMax M3 (#2551) by @athitten
- docs(speculative): fix EAGLE drafter layer count and EAGLE-1 loss (#2796) by @khazic
- fix(ci): drop base-image uv/wandb copies flagged for CVEs (#2800) by @thomasdhc
- chore(skills): refresh automodel skill signatures (#2804) by @akoumpa
- feat(distributed): enable selective checkpointing for DDP (#2786) by @yuhezhang-ai
- docs: align nightly navigation routes (#2801) by @akoumpa
- feat: vision biencoder finetuning + Nemotron VL 1B finetuning support (#2354) by @gabrielspmoreira
- fix(docs): catch Fern MDX syntax errors (#2806) by @akoumpa
- feat(moe): add Rollout Routing Replay (R3) for MoE RL training (#2797) by @khazic
- fix(optim): align Dion mesh with FSDP sharding (#2808) by @akoumpa
- fix(ci): stabilize failed benchmark recipes (#2817) by @akoumpa
- docs: document opt-in media extras (vlm-media/diffusion-media) (#2799) by @thomasdhc
- feat(speculative): add SGLang target backend for EAGLE-3 training (#2449) by @khazic
- feat(speculative): add Domino online training path on top of DFlash (#2819) by @khazic
- feat(speculative): add dspark draft model and training objective (#2810) by @kashif
- docs: add GLM-5.2 and speculative decoding (DSpark) updates (#2828) by @khazic
- feat(diffusion): support Hugging Face datasets (#2816) by @pthombre
- feat(speculative): add vLLM target backend for EAGLE-3 training (#2798) by @khazic
- fix(speculative): gather sharded target lm_head in EAGLE-1/2 token loss (#2823) by @khazic
- fix(model): honor yarn/linear/dynamic RoPE in LlamaRotaryEmbedding (#2825) by @khazic
- fix(speculative): keep DFlash padding blocks self-attending to avoid NaN (#2826) by @khazic
- feat(speculative): support target tensor parallelism in EAGLE-3 colocated path (#2827) by @khazic
- feat(speculative): support target tensor parallelism in EAGLE-1/2 (#2829) by @khazic
- feat(speculative): support target tensor parallelism in DFlash (#2830) by @khazic
- fix(speculative): validate DFlash mask_token_id on checkpoint resume (#2824) by @khazic
- docs: enable Fern multi-source (#2845) by @lbliii
- fix(distributed): propagate NCCL timeout to derived device meshes (#2846) by @HuiyingLi
- feat(speculative): add serve_vllm for EAGLE-3 / P-EAGLE drafts (#2841) by @khazic
- fix(speculative): keep DSpark draft RoPE inv_freq in fp32 under bf16 (#2859) by @khazic
- feat(speculative): compress EAGLE-3 offline cache target_probs via top-k (#2847) by @khazic
- perf: FFPA D=512 attention backend for Gemma4 (3× fwd / 6× bwd vs SDPA) (#2436) by @Butterfingrz
- feat(speculative): add JetSpec causal parallel drafting training (#2867) by @khazic
- feat(models): reject tie_word_embeddings=True on separate-head model families (#2805) by @Achyuthan-S
- ci: Ensure diffusion-media extra installed for diffusion tests (#2858) by @chtruong814
- feat(speculative): add DeepSeek-V3 (MLA) EAGLE-3 draft model (#2849) by @khazic
- feat(speculative): add DeepSeek V4 DSpark drafter and V4-Flash training (#2866) by @khazic
- chore(models): add 26.10 deprecation warnings for custom model classes (#2807) by @athitten
- feat(speculative): add DSpark draft for MiniMax M3 VL (text + multimodal) (#2877) by @khazic
- ci: use NVIDIA inference for Claude review (#2882) by @chtruong814
- feat(speculative): add GLM-5.2 DSpark draft model and training (#2885) by @khazic
- fix(checkpoint): write consolidated safetensors without append (#2627) by @huahuajhu
- ci(automodel): set AM-576 release timeouts (#2871) by @yuhezhang-ai
- fix(diffusion_gemma): make parity test work on transformers >= 5.11 (#2883) by @zyzhou5
- fix(bagel): torch2.12 DTensor sharding-prop leak (#2888) by @zyzhou5
- fix(speculative): plumb EAGLE-3.1 fc_norm/norm_output into the draft config (#2897) by @khazic
- fix(speculative): make save_consolidated final export the last EAGLE checkpoint (#2898) by @khazic
- fix(speculative): gather TP target outputs in the Domino trainer (#2899) by @khazic
- fix(speculative): default DFlash/Domino loss_decay_gamma to the paper value (#2908) by @khazic
- refactor(speculative): dispatch all DSpark drafts through the registry (#2909) by @khazic
- feat(speculative): serve DFlash and JetSpec drafts on vLLM (#2913) by @khazic
- docs(speculative): document DSpark, Domino, JetSpec, and the vLLM serve path (#2914) by @khazic
- feat(speculative): add bench_vllm acceptance/speedup benchmark (#2912) by @khazic
- fix(speculative): read DeepSeek EAGLE-3 draft rope_theta from rope_parameters (#2920) by @khazic
- fix(speculative): send chat_template_kwargs as a top-level request field (#2921) by @khazic
- ci(deps): add albumentations to vlm-media for nemotron-parse processor (#2933) by @thomasdhc
- fix(glm): select HybridEP for GLM-5.2 recipes (#2934) by @HuiyingLi
- fix(distributed): omit unset FSDP reshard argument (#2926) by @zhiqi-li
- feat(speculative): add DSpark offline cache path (#2924) by @fkuner
- chore: update readme for 26.06-26.08 (#2886) by @akoumpa
- ci: validate ci section on newly added recipes (#2936) by @thomasdhc
- fix(bagel): make non-backbone init topology independent (#2939) by @zyzhou5
- feat(speculative): add fp8 draft training and LoRA draft adaptation (#2963) by @khazic
- feat(speculative): report simulated accept length during EAGLE-3 training (#2959) by @khazic
- fix(distributed): fix and scope vlm activation checkpointing (#2840) by @yuhezhang-ai
- feat(models): complete tie_word_embeddings guards for #2512 families (#2896) by @Achyuthan-S
- feat(speculative): add distributed DSpark offline precompute (#2953) by @khazic
- ci: route gb200 L2_HF_DCP to the IMEX runner for DeepEP NVLink P2P (#2968) by @ko3n1g
- docs: add 0.5.0 release notes (#2881) by @lbliii
- fix(vlm): gemma4 FFPA mock + e4b cp16 CI recipe fixes (AM-626, AM-627) (#2949) by @athitten
- build(deps): upgrade transformers to 5.12.1 (#2873) by @athitten
- ci: update claude review guidelines (#2739) by @akoumpa
- fix(moe): preserve MLP dispatch through checkpoint wrappers (#2955) by @akoumpa
- feat(speculative): add multi-dataset acceptance-length benchmark sweep (#2966) by @khazic
- fix(transformers): gate _tie_weights_nemo on tie_word_embeddings flag (#2942) by @yuhezhang-ai
- fix(loss): avoid pkg_resources in linear CE (#2700) by @akoumpa
- fix(deepseek_v4): support packed THD with context parallel (#2731) by @HuiyingLi
- feat(speculative): FSDP2-shard dense DSpark target + add Gemma4-31B config (#2976) by @khazic
- fix(attention): initialize FlexAttention masks (#2982) by @akoumpa
- fix(deepseek-v4): initialize random training state (#2991) by @akoumpa
- ci(review): strengthen model review invariants (#2994) by @akoumpa
- ci: Bump claude_review to v1.8.2 (#3012) by @chtruong814
- ci: AUT-779 repin _claude_review.yml to v1.8.3 (#3013) by @svcnemo-autobot
- test(ci): accept SHA-pinned claude-review template refs in the policy test (#3022) by @khazic
- feat(speculative): context parallelism for draft-model training (#2918) by @kashif
- feat(pp): support FusedLinearCrossEntropy under pipeline parallelism (#2927) by @beccohov
- feat(speculative): support sequence packing in the DeepSeek MLA EAGLE-3 draft (#3002) by @khazic
- fix(datasets): reject non-prefix multiturn mask templates (#3024) by @khazic
- feat(speculative): support sequence packing in the EAGLE-1/2 draft (#3003) by @khazic
- feat(speculative): add variable-prefix DFlash and LK EAGLE-3 training losses (#3032) by @khazic
- feat(dspark): log acceptance and confidence metrics during training (#2957) by @kashif
- feat(speculative): support sequence packing in the DFlash draft (#3004) by @khazic
- feat(speculative): add periodic real accept-length eval (#3037) by @khazic
- feat(speculative): support sequence packing in the Domino and JetSpec drafts (#3023) by @khazic
- feat(speculative): support sequence packing in the DSpark Qwen3 draft (#3005) by @khazic
- fix(models): temporarily disable TE fused RoPE globally (#3028) by @HuiyingLi
- fix(glm_moe_dsa): clamp DSA indexer top-k for short sequences (#3035) by @HuiyingLi
- feat(eagle3): on-policy regeneration loop with step-cadence dataloader swap (#3042) by @khazic
- fix(docs): update technical accuracy (#3031) by @akoumpa
- fix(checkpoint): compare torch versions semantically (#3043) by @HuiyingLi
- docs: add uv run prefix (#3057) by @akoumpa
- fix(recipe): add generation-marked chat templates (#3051) by @khazic
- ci: include cutlass deps (#3056) by @akoumpa
- fix(dspark): omit unmeasured positional acceptance rates (#3050) by @khazic
- test(transformers): add source-load parity coverage (#2960) by @yuhezhang-ai
- fix(distributed): resolve VLM AC layer-group paths structurally, not by version (#3025) by @yuhezhang-ai
- fix(vlm): load Tulu-3 directly from HF Hub instead of local meta JSON (#3001) by @athitten
- fix(optim): drop zero-numel local DTensor shards before TE FusedAdam (#2997) by @yuhezhang-ai
- fix(moe): checkpoint trainable vision towers on the expert-parallel path (#2993) by @yuhezhang-ai
- feat(datasets): add optional parallel pre-tokenization before packing (#3045) by @khazic
- fix(models): handle checkpoint-wrapped fp32 buffers (#3059) by @akoumpa
- feat(optim): support per-parameter-group learning rate (#3046) by @khazic
- fix(cp): preserve gradients when sharding VLM inputs (#2931) by @HuiyingLi
- refactor(vlm): gate build_model via recipe-side model-target allowlist (#2357) by @athitten
- fix(ci): use HybridEP for Step 3.5 benchmark (#3069) by @HuiyingLi
- fix: Kimi K2 config loading without remote code (#3065) by @HuiyingLi
- feat(speculative): add Gemma4 EAGLE-3 target support (#3071) by @khazic
- fix(dllm): remove obsolete DiffusionGemma HF compatibility (#3067) by @zyzhou5
- feat(speculative): add Gemma4-E4B EAGLE-3 example config (#3073) by @khazic
- fix(test): use SDPA for Nemotron-H HF reload (#3060) by @yuhezhang-ai
- feat(speculative): add Gemma4-31B EAGLE-3 example config (#3077) by @khazic
- feat(speculative): add Gemma4-26B-A4B MoE EAGLE-3 example config (#3079) by @khazic
- chore(ci): AUT-852 bump claude review template to v1.8.4 (#3087) by @svcnemo-autobot
- feat: add Ministral3 embedding distillation recipe (#3058) by @vinay-raman
- fix(perf): use GPT-OSS head dimension in FLOPs accounting (#3091) by @yaoyu-33
- test(speculative): add EAGLE-3 fp8 draft convergence smoke for SM89+ (#3081) by @khazic
- fix(speculative): keep DSpark target_layer_ids in [0, N-2] for SGLang (#3083) by @khazic
- chore(skills): add Regent Open Plugin manifest (#3097) by @ko3n1g
- chore(skills): remove Open Plugin manifest (superseded) (#3099) by @ko3n1g
- test: clean up KD test process group (#3090) by @yuhezhang-ai
- refactor(models): single tie_word_embeddings guard via TieSupport (#2998) by @Achyuthan-S
- feat(speculative): add DFlash validation metrics (#3072) by @khazic
- refactor(datasets): typed Config + build per dataset, config-driven dataloader (#2390) by @akoumpa
- fix(bagel): enable periodic garbage collection (#3105) by @zyzhou5
- fix(ci): pass PR number and commit to Codecov for fork PRs (#3101) by @svcnemo-autobot
- fix(test): use eager attention for Nemotron-H HF loads (#3100) by @yuhezhang-ai
- feat(kd): support separate student and teacher meshes (#2954) by @akoumpa
- fix(moe): safe TP and EP/TP gradient correctness for custom MoE models (#2995) by @yuhezhang-ai
- test(ci): remove deprecated 26.10 models from nightly and release CI (#2884) by @athitten
- feat(training): add opt-in setup-time prewarms (#2992) by @yuhezhang-ai
- fix: bound chunked CE memory and preserve packing masks (#2996) by @yuhezhang-ai
- fix(ci): restore Ministral3 checkpoint robustness (#3111) by @yuhezhang-ai
- ci: add HF hub cache preflight check (#3119) by @thomasdhc
- fix(vlm): resolve get_rope_index from the base model for packed mRoPE (#3113) by @khazic
- feat(models): add Inkling VLM MoE support (#3095) by @hemildesai
- ci(dllm): add dLLM SFT nightly train-to-generate launcher and recipes (#2783) by @zyzhou5
- feat(vlm): add THD packed-sequence support for Qwen3-VL-MoE (#3052) by @khazic
- docs: Add Inkling README news (#3129) by @HuiyingLi
- fix(checkpoint): gather PEFT adapter across PP stages (#3096) by @hyfine
- build(deps): bump base container to 26.06-cuda13.3 (#2983) by @thomasdhc
- fix(transformers): keep custom M3 config when transformers ships its own (#3134) by @HuiyingLi
- ci: AUT-911 shard GPU unit tests into 5 pytest-shard chunks (#3138) by @svcnemo-autobot
- fix(checkpoint): isolate consolidation timeout from NCCL (#3108) by @yuhezhang-ai
- test(vlm): add checkpoint robustness coverage (#3112) by @yuhezhang-ai
- fix(ci): preserve retrieval evaluation schedule (#3139) by @yuhezhang-ai
- feat(bagel): add TE support (#2895) by @zyzhou5
- fix(moe): preserve HSDP replica gradient synchronization (#3135) by @wangzhxg
- docs: add Laguna SFT docs and recipe (#3146) by @HuiyingLi
- ci: AUT-895 gate GB200 tests on DISABLE_GB200_TESTS variable (#3118) by @svcnemo-autobot
- feat(dllm): add LLaDA2 generation support (#3092) by @zyzhou5
- feat(models): add Laguna model implementation (#3148) by @HuiyingLi
- fix: prevent damaged token embeddings from dominating grad clipping (#3136) by @akoumpa
- fix(checkpoint): add bounded retention window (#2416) by @oliverholworthy
- fix(kd): use fp32 master weight copy (#3019) by @akoumpa
- fix(checkpoint): restore adapters through DDP wrappers (#3150) by @yuhezhang-ai
- feat(distributed): block-diagonal varlen CP for packed sequences (#2989) by @yuhezhang-ai
- fix(distributed): preserve canonical activation checkpoint keys (#3152) by @yuhezhang-ai
- feat(examples): Gemma4-31B CoderForge data pipeline + CP SFT recipes (#3151) by @athitten
- feat(attention): support flash_attention_3 and flash_attention_4 (#2929) by @HuiyingLi
- feat(checkpoint): add DCP CPU offload option (#3130) by @khazic
- test(packing): fix stale unsupported-backend test after fa3/fa4 support (#3199) by @khazic
- test(ci): add DeepSeek V4 Flash pretrain coverage (#3128) by @akoumpa
- refactor(distributed): unify CP input prep and dispatch across models (#2937) by @HuiyingLi
- docs(fern): add legacy URL redirects (#3203) by @akoumpa
- ci: reduce L0 GPU tests to two shards (#3201) by @akoumpa
- feat: add checkpoint staging wait option (#3131) by @khazic
- fix(datasets): keep system turns in the sharegpt conversation converter (#3144) by @khazic
- fix(peft): honor memory-efficient LoRA opt-out (#3126) by @akoumpa
- feat: enable text inclusion alongside images in retrieval training (#3189) by @rnyak
- feat: Support preemption checkpointing (#3007) by @edjson
- fix(ci): AUT-964 restore Codecov after skipped GB200 jobs (#3205) by @svcnemo-autobot
- fix(ci): preserve HF meta init for device-mapped loads (#3188) by @yuhezhang-ai
- feat(dllm): add lora to dllm (#3163) by @zyzhou5
- fix(gemma4): run SDPA in fp32 to avoid #2208 NaN on Hopper (#3141) by @aminehd
- docs(retrieval): complete fine-tuning guide integration (#2306) by @oliverholworthy
- fix(glm): apply LoRA to TileLang MLA KV projection (#3176) by @shahafwa
- fix: prefer Automodel config registry lookup (#3202) by @akoumpa
- feat(distributed): shape TPLinear/LinearLoRA graphs for async-TP fusion (#2987) by @yuhezhang-ai
- fix(distributed): Megatron-FSDP 0.5.0 compatibility (#2986) by @yuhezhang-ai
- feat(checkpoint): preserve intrinsic fp32 during offline consolidation (#3193) by @yuhezhang-ai
- feat(kernels): add QuACK backend (#3115) by @akoumpa
- feat(dllm): add DiffusionGemma generation via the built-in HF sampler (#3161) by @zyzhou5
- refactor(diffusion): migrate recipe onto typed RecipeConfig build() path (#3122) by @pthombre
- fix(dist): keep profiler record-function ops out of SAC replay accounting (#3133) by @HuiyingLi
- fix(vlm): size qwen3.6 medpix CI configs (#3209) by @HuiyingLi
- feat(glm_moe_dsa): expose update_moe_gate_bias on GLM MoE DSA models (#3207) by @jQizhang
- fix(ci): shard Nemotron Super vLLM deploy across 8 GPUs (#3061) by @yuhezhang-ai
- fix(glm): size GLM5.2 release CI jobs (#3210) by @HuiyingLi
- fix(distributed): activation-checkpoint Qwen3-Next linear_attn layers (#3192) by @amolkhanna
- fix(distributed): support packed CP for Llama, Qwen2, and Qwen3 (#2999) by @akoumpa
- fix(dllm): recipe defaults, corruption seeding, sampler and docs fixes (#3162) by @zyzhou5
- perf(bagel): grouped MoT routing + fused SwiGLU/RoPE (#3214) by @zyzhou5
- feat(retrieval): add normalized Arrow dataset tooling (#2596) by @yuhezhang-ai
- feat(diffusion): context parallelism via diffusers ContextParallelConfig (#3157) by @pthombre
- docs: explain embedding row repair (#3175) by @akoumpa
- fix(transformers): custom configs override builtin ones by default (#3221) by @HuiyingLi
- build(deps): bump ffpa-attn to 0.2.2 (#3225) by @akoumpa
- fix(distributed): checkpoint linear attention with compile (#3213) by @yuhezhang-ai
- fix(glm): own the GLM MoE DSA config so qk_rope_head_dim survives (#3222) by @HuiyingLi
- feat(speculative): add the multimodal speculative decoding (MSD) core (#3166) by @khazic
- fix(speculative): average DFlash train metrics over the log window (#3234) by @khazic
- fix(speculative): honor --dflash-causal on the vLLM remapping export (#3235) by @khazic
- fix(speculative): align the decode-eval cadence on resume (#3237) by @khazic
- fix(speculative): validate mask_token_id when resuming DSpark (#3238) by @khazic
- build: cache CUDA extension source builds (#3204) by @akoumpa
- docs(fern): cut v0.5.0 version train (#2970) by @lbliii
- fix(checkpoint): infer Qwen MTP layout from checkpoint keys (#3229) by @HuiyingLi
- docs: fix remaining broken link sources (#3249) by @akoumpa
- fix(distributed): control frozen multimodal FSDP sharding (#2763) by @yuhezhang-ai
- fix(ci): preserve main wheelhouse cache builds (#3253) by @akoumpa
- feat(diffusion): add Qwen-Image-Edit-2511 training support (#3217) by @pthombre
- fix(ci): omit token type IDs from vLLM parity prompts (#3198) by @yuhezhang-ai
- fix(deepseek-v4): accept cu_seqlens for THD packing (#3231) by @HuiyingLi
- fix(deepseek-v4): preserve fp32 model dtype (#3227) by @HuiyingLi
- fix(thd): derive packed padding mask from the pack layout, not token ids (#3223) by @HuiyingLi
- fix(nemotron-v3): recompute deterministic LoRA router (#3258) by @HuiyingLi
- fix(moe): preserve top-k routing under full AC (#3140) by @yuhezhang-ai
- fix(gpt-oss): route packed attention through THD (#3226) by @akoumpa
- perf(benchmark): tune Qwen3-VL LoRA pipeline batches (#3230) by @akoumpa
- fix(moe): preserve gate load across activation recompute (#3247) by @akoumpa
- fix(ci): narrow CUDA wheelhouse cache keys (#3269) by @akoumpa
- feat(distributed): frame-level context-parallel vision-tower sharding (#2990) by @yuhezhang-ai
- fix(vlm): preserve mRoPE axes in PP chunking (#3208) by @HuiyingLi
- feat(diffusion): add LTX-2.3 video+audio finetuning (training + prepr… (#3165) by @linnanwang
- fix: harden checkpoint torch loads (#3240) by @HuiyingLi
- build: add MagiAttention optional dependency (#3070) by @HuiyingLi
- fix(checkpoint): harden PP safetensors consolidation (#3248) by @yuhezhang-ai
- fix(checkpoint): save all PEFT EP optimizer parts (#3250) by @yuhezhang-ai
- fix(ci): install CUDA torchvision with CUDA torch (#3279) by @akoumpa
- fix(docker): cap uv install concurrency to avoid ARM FD exhaustion (#3275) by @thomasdhc
- feat: support pre-extracted frame sequence of video sample (#3211) by @GITsologun
- fix(inkling): support HF 5.14 attention fields and embed norm FSDP (#3281) by @HuiyingLi
- fix(KD): Fix variable-length teacher PP batches on separate meshes (#3280) by @HuiyingLi
- feat: add Kimi K3 model support (#3259) by @HuiyingLi
- docs(models): add Kimi K3 coverage (#3283) by @HuiyingLi
- test(vlm): stabilize Qwen3.5 checkpoint robustness (#3233) by @HuiyingLi
- feat(vlm): integrate packed CP and vision sharding for Qwen3.5-MoE (#3186) by @yuhezhang-ai
- feat(checkpoint): defer distributed async consolidation (#3125) by @khazic
- test(retrieval): add Nemotron VL checkpoint coverage (#3276) by @yuhezhang-ai
- chore(ci): AUT-1135 pin GitHub Actions to commit SHAs (#3277) by @svcnemo-autobot
- feat: add ViSpec VLM draft training (#3173) by @khazic
- fix(docs): hide Kimi tokenizer regex from autodoc (3288) (#3304) by @svcnvidia-nemo-ci
- docs(distributed): review frozen multimodal FSDP guidance (3272) (#3306) by @svcnvidia-nemo-ci
- fix(vlm): fused linear CE in gemma4 31B FFPA 8k recipe (AMINT-203) (3273) (#3307) by @svcnvidia-nemo-ci
- fix(pp): use static metadata with PyTorch 2.13 (3290) (#3305) by @svcnvidia-nemo-ci
- ci: add Tulu-3 convergence + eval flow (3171) (#3312) by @svcnvidia-nemo-ci
- fix(distributed): preserve ERNIE router and dense Qwen3.5 SSM precision (3255) (#3310) by @svcnvidia-nemo-ci
- perf(moe): add scoped partial CUDA graphs (2917) (#3314) by @svcnvidia-nemo-ci
- fix(docs): avoid literal ampersand in Kimi autodoc (3317) (#3320) by @svcnvidia-nemo-ci
- fix(config): repair GLM-5.2 LoRA recipe (3313) (#3323) by @svcnvidia-nemo-ci
- fix(kd): preserve tensor-valued hidden states (3324) (#3332) by @svcnvidia-nemo-ci
- fix(kd): use mesh-safe gradient clipping (3302) (#3334) by @svcnvidia-nemo-ci
- refactor(docker): build torchao & FlashAttention as isolated wheel stages (3329) (#3343) by @svcnvidia-nemo-ci
- fix(distributed): restore Nemotron Flash TP2 training (3345) (#3349) by @svcnvidia-nemo-ci
- fix(nemotron-parse): sync RADIO preprocessing (3331) (#3341) by @svcnvidia-nemo-ci
- fix(distributed): reuse default group for world-sized meshes (3319) (#3347) by @svcnvidia-nemo-ci
- fix(training): prewarm Mamba SSD autotune kernels (3296) (#3338) by @svcnvidia-nemo-ci
- fix(docs): make Fern autodoc metadata MDX-safe (3355) (#3357) by @akoumpa
- fix(pp): preserve VLM media cursor with static metadata (3344) (#3361) by @svcnvidia-nemo-ci
- fix(deps): resolve 26.08 rc2 container CVEs (3346) (#3363) by @svcnvidia-nemo-ci
- fix(pp): route Qwen3.5 MoE pre-embedded inputs (3294) (#3364) by @svcnvidia-nemo-ci
- fix(kimi_k25_vl): don't int4 quantize LoRA adapter keys on save (3295) (#3362) by @svcnvidia-nemo-ci
- fix(diffusion): use spawn start method for GPU preprocessing pools (3339) (#3366) by @svcnvidia-nemo-ci
- fix(tokenizer): make special token insertion opt-in (3337) (#3375) by @svcnvidia-nemo-ci
- fix(distributed): stabilize SAC replay with TE and FSDP (3330) (#3376) by @svcnvidia-nemo-ci
- fix(ci): use a single container cache donor (3377) (#3381) by @svcnvidia-nemo-ci
- fix(fsdp): resolve fp32 master-weight compute dtype per parameter (3328) (#3384) by @svcnvidia-nemo-ci
- perf(checkpoint): reduce distributed save overhead (3369) (#3389) by @svcnvidia-nemo-ci
- fix(ci): enable LTX-2.3 diffusion finetuning (3372) (#3391) by @svcnvidia-nemo-ci
- ci: tune Nemotron single-GPU model load threads (3370) (#3395) by @svcnvidia-nemo-ci
- fix(model): correct Mistral4 attention and distributed MoE routing (3348) (#3393) by @svcnvidia-nemo-ci
- fix(checkpoint): survive an interrupted save instead of hanging all ranks (3261) (#3394) by @svcnvidia-nemo-ci
- perf: remove Python overhead from model hot paths (3374) (#3401) by @svcnvidia-nemo-ci
- fix(deps): resolve OSS CVE findings (3398) (#3412) by @svcnvidia-nemo-ci
- fix(moe): support EP-free DTensor state-dict conversion (3397) (#3409) by @svcnvidia-nemo-ci
- fix(checkpoint): harden PEFT PP checkpointing and Step-3.7 coverage (3316) (#3403) by @svcnvidia-nemo-ci
- docs(training): mark Mamba prewarm sections for review (3340) (#3399) by @svcnvidia-nemo-ci
- fix(retrieval): guard optional W&B imports (3380) (#3402) by @svcnvidia-nemo-ci
- fix(gemma4): enable expandable_segments on the CP tulu3 recipes (3408) (#3410) by @svcnvidia-nemo-ci
- feat: Add {% generation %} chat template for DiffusionGemma SFT/LoRA examples (3353) (#3417) by @svcnvidia-nemo-ci
- fix(docs): AUT-1324 qualify LTX model coverage slug (3418) (#3419) by @svcnvidia-nemo-ci
- fix(ci): run Kimi K3 HellaSwag on GB200 (3420) (#3421) by @svcnvidia-nemo-ci
- fix(transformers): default missing THD capability (3406) (#3429) by @svcnvidia-nemo-ci
- ci(convergence): (3416) (#3432) by @svcnvidia-nemo-ci
- cp: backport MoE checkpoint reload parity fixes to r0.6.0 (#3438) by @yuhezhang-ai
- fix(retrieval): use stock Ministral embedding backbone (3103) (#3433) by @svcnvidia-nemo-ci
- ci(minimax): extend M2.7 LoRA timeout (3428) (#3430) by @svcnvidia-nemo-ci
- fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (3411) (#3445) by @svcnvidia-nemo-ci
- fix(kimi_k25_vl): keep PEFT expert LoRA keys under the base_model.model. prefix (3431) (#3451) by @svcnvidia-nemo-ci
- ci(vlm): raise Slurm wall time for MiniMax-M3 and Gemma4 recipes (3448) (#3453) by @svcnvidia-nemo-ci
- fix: Gemma4-31B CoderForge CP8/64K data processing and recipe (3446) (#3452) by @svcnvidia-nemo-ci
- cp: backport PEFT v5 adapter output fixes to r0.6.0 (#3458) by @yuhezhang-ai
- fix(retrieval): repair Ministral3 training recipe (3441) (#3454) by @svcnvidia-nemo-ci
- test(checkpoint): redesign resume robustness around a shared trajectory (3427) (#3457) by @svcnvidia-nemo-ci
- test(checkpoint): stabilize MoE checkpoint parity gates (3440) (#3460) by @svcnvidia-nemo-ci
- fix(gemma4): exact router scalar and fp32 reference routing for MoE parity (3456) (#3466) by @svcnvidia-nemo-ci
- cp: backport Sentence Transformers metadata export to r0.6.0 (#3464) by @yuhezhang-ai
- fix: normalize MoE auxiliary loss during gradient accumulation (3359) (#3459) by @svcnvidia-nemo-ci
- fix(bagel): type SFT backend before AutoModel init (3449) (#3469) by @svcnvidia-nemo-ci
- fix(peft): enable v5 expert adapters for Nemotron and MiniMax (3439) (#3468) by @svcnvidia-nemo-ci
- fix(checkpoint): save ranked RNG state per global rank (3437) (#3471) by @svcnvidia-nemo-ci
- fix(peft): align frozen LoRA tensors with compute dtype (3470) (#3480) by @svcnvidia-nemo-ci
- cp: backport VLM CP gradient and vision sharding fixes to r0.6.0 (#3479) by @yuhezhang-ai
- fix(training): all-reduce grad-norm scalars on the mesh device (3461) (#3481) by @svcnvidia-nemo-ci
- fix(recipe): shard GPT-OSS 120B across 64 experts (3483) (#3489) by @svcnvidia-nemo-ci
- fix(ci): register flux2/wan2.2/qwen-image-edit diffusion recipes in CI (3485) (#3487) by @svcnvidia-nemo-ci
- perf(benchmarks): recompute deterministic MoE routers under AC (3474) (#3490) by @svcnvidia-nemo-ci
- chore: Bump gitpython to >= 3.1.59 (3482) (#3488) by @svcnvidia-nemo-ci
- perf(recipes): run GPT-OSS 120B with EP64 and no activation checkpointing (#3492) by @akoumpa
- cp: feat(models): make Inkling standalone (#3358) into r0.6.0 (#3498) by @HuiyingLi
- fix(models): preserve lm-head dtype boundaries (3491) (#3502) by @svcnvidia-nemo-ci
- fix(checkpoint): validate PEFT adapter-only state (3501) (#3503) by @svcnvidia-nemo-ci
- ci(convergence): fix gemma4 eval setup and re-baseline Qwen3-MoE (3493) (#3505) by @svcnvidia-nemo-ci
- fix(deepseek_v4): avoid TileLang boolx8 backward codegen (3467) (#3506) by @svcnvidia-nemo-ci
- fix(checkpoint): preserve FSDP2 mixed precision during recompute (3513) (#3515) by @svcnvidia-nemo-ci
- fix(deps): resolve Starlette and GitPython CVEs (3523) (#3524) by @svcnvidia-nemo-ci
- fix(minimax): declare packed-sequence and CP attention ownership (#3548) by @athitten
- feat(gemma4): add E-series tensor parallelism (3512) (#3552) by @svcnvidia-nemo-ci
- fix(retrieval): support canonical Sentence Transformers metadata (3546) (#3551) by @svcnvidia-nemo-ci
- ci: route DGX Spark recipes to GB10 (3538) (#3556) by @svcnvidia-nemo-ci
- fix(docker): build bitsandbytes for SM121 (3553) (#3557) by @svcnvidia-nemo-ci
- cp: backport packed THD VLM context parallelism to r0.6.0 (#3558) by @qiaochuz-nv
- fix(fsdp): uniform reduce dtype and EP-local expert gradients (3540) (#3561) by @svcnvidia-nemo-ci
- fix(deps): resolve msgpack, wandb, and mistune container CVEs (3563) (#3566) by @svcnvidia-nemo-ci
- fix(deps): resolve 26.08 rc9 container CVEs (3607) (#3611) by @thomasdhc
- fix(deps): resolve 26.08 rc10 container CVEs (3647) (#3658) by @thomasdhc
- fix(peft): support mixed-dtype memory-efficient LoRA backward (3675) (#3682) by @svcnvidia-nemo-ci
- beep boop 🤖: Bumping NeMo-Automodel to v0.5.1 [skip ci] by @github-actions[bot]
- fix(release): set version to 0.6.0 (#3692) by @thomasdhc