Skip to content

NVIDIA NeMo-Automodel 0.6.0

Latest

Choose a tag to compare

@nemo-automation-bot nemo-automation-bot released this 26 Aug 17:54
89c248a
Highlights
  • Frontier-scale MoE and multimodal models. Train Moonshot AI's Kimi K3 (2.8T/104B active), Thinking Machines Lab's Inkling (975B/41B active), MiniMax AI's MiniMax-M3 VL (428B/22B active), Poolside's Laguna, and Zhipu AI's GLM-5.2.
  • Long-context training. Shard packed multi-document sequences with block-diagonal variable-length context parallelism, and shard vision towers by frame to remove the redundant vision compute of a CP fold; recipes reach 128K tokens at EP8/CP32.
  • New attention and kernel backends. Select FlashAttention 3 and 4, MagiAttention, FFPA for headdim=512 layers, QuACK linear/RMSNorm/RoPE kernels, and scoped partial CUDA graphs for MoE.
  • Resilient checkpointing. Checkpoint on preemption signals, write to msc:// cloud directories, consolidate async saves in the background, bound checkpoint retention, and resume past an interrupted save.
  • Diffusion. Fine-tune and generate with LTX-2.3 joint video+audio and Qwen-Image-Edit-2511, and run context parallelism through the diffusers ContextParallelConfig API.
New Hardware and Precision Support
  • FP8 draft training. Extends the SFT recipes' top-level fp8: block to every speculative recipe, swapping the draft's nn.Linear layers to torchao Float8Linear on SM89 and newer, with emulate: true for older GPUs. (#2963)
  • Quantized checkpoint ingestion. Dequantizes MXFP4 dequantization for Kimi K3 and MXFP8 dequantization for MiniMax-M3 for BF16 training, keeps Kimi-K2.5-VL LoRA keys out of INT4 quantization on save, and fixes a TileLang boolx8 backward codegen failure on the DeepSeek-V4 path. (#3259, #3295, #3467)
New and Expanded Model Support
LLM and MoE
VLM
  • MiniMax AI: MiniMax-M3 VL (428B / 22B-active MoE) — MedPix full SFT at EP32/PP4 on 16 nodes, MedPix LoRA at PP4/EP8, a CP2 comparison recipe, and 16K packed Tulu-3 text SFT at CP8. (docs, #2538, #2551)
  • Thinking Machines Lab: Inkling (975B / 41B-active MoE; text, image, video, audio) — MedPix SFT at PP8/EP32 on 32 nodes. (docs, #3095, #3281, #3358)
  • Alibaba: Qwen3.5 and Qwen3.6 VLM (122B-A10B MoE, 27B dense) — 128K packed sequences at EP8/CP32 on 16 nodes with a trainable vision tower, dense context parallelism, and matched CP1/CP2 27B MedPix recipes. (docs, #2505, #3186, #3209)
  • Google: Gemma 4 (31B and E2B/E4B dense, 26B-A4B MoE) — context parallelism with 16K Tulu-3 at CP8, 64K Tulu-3 at CP16, and 4K MedPix at EP8/CP2; plus fp32 SDPA to avoid NaNs on Hopper and an exact fp32 reference router for MoE parity. (docs, #1914, #2592, #2621, #3141)
  • Moonshot AI: Kimi-K2.5 VL (MoE) — LoRA tensors kept out of INT4 expert quantization on save, so adapters reload on resume. (#3295, #3431)
  • NVIDIA: Nemotron-Parse — corrected preprocessing stopping the RADIO encoder from normalizing already-normalized images a second time. (#3331)
  • MTP on multimodal inputs. Pre-rolls pre-fused inputs_embeds per depth so audio and vision positions keep their continuous embedding instead of being re-embedded from a padding token id. (#2510)
  • Pre-extracted video frames. Loads a video field holding a list of image paths directly as a frame sequence with no decoder, padded to the processor's temporal_patch_size. (#3211)
Diffusion
  • Lightricks: LTX-2.3 (video + audio) — full-finetuning, LoRA, and generation recipes that train video and audio jointly and mux the generated waveform into the output mp4; --processor ltx2 preprocessing requires an 8n+1 frame count. (docs, #3165, #3372)
  • Alibaba: Qwen-Image-Edit-2511 — cached, full-parameter instruction-based editing with an offline latent and prompt cache encoder, an eight-GPU BF16 recipe, and an image-edit preprocessing subcommand. (docs, #3217)
  • Diffusion context parallelism. Drives the diffusers ContextParallelConfig API from fsdp.cp_size, reusing the cp axis of the existing FSDP2 mesh; only pure Ulysses sharding is accepted. (#3157)
dLLM
  • Google: DiffusionGemma 26B-A4B (block-diffusion MoE) — a native implementation with GSM8K SFT and LoRA recipes at ep_size: 8. (docs, #2506)
  • Generation and LoRA. Exposes --sampler presets llada, llada2, nemotron, and gemma, an --adapter merge path, and llada_lora, llada2_lora, and nemotron_labs_diffusion_lora recipes. (#3092, #3161, #3163)
  • Reproducible corruption. Seeds corruption noise per sample from that example's global index in the shuffled stream, so noise is independent of parallel topology and resume reproduces exactly. (#3162)
Speculative Decoding
  • DSpark. Introduces a semi-autoregressive draft in which a parallel backbone proposes a whole block, a serial Markov head adds intra-block dependency, and a confidence head predicts acceptance, with configs for Qwen3, Gemma 4, GLM-5.2, MiniMax-M3, and DeepSeek-V4-Flash. (docs, #2810, #2866, #2885, #2909)
  • Domino, JetSpec, and ViSpec. Adds a GRU correction head on the DFlash backbone (Domino), causal in-block attention distilled with a temperature-scaled forward KL (JetSpec), and two-stage VLM draft training (ViSpec). (#2819, #2867, #3173)
  • New EAGLE-3 targets. Registers a DeepSeek-V3 MLA draft class reusing the target's low-rank q/kv projections, and registers Gemma 4 as a target with E2B, E4B, 31B, and 26B-A4B configs. (docs, #2849, #3071, #3077, #3079)
  • Sequence packing, TP, and CP. Enables block-causal packed training for every draft family, target tensor parallelism for EAGLE-1/2/3 and DFlash, and context parallelism for both target and draft through a differentiable ring attention. (#2827, #2918, #3002, #3005)
  • fp8, compile, LoRA, and new objectives. Exposes a compile: block, EAGLE-3 adapter-only training through peft:, DFlash loss_type: variable_prefix, and EAGLE-3 lk_loss_type: alpha/lambda. (#2963, #3032, #3081)
  • Feature-noise augmentation. Perturbs the target features fed to the EAGLE-1/2 draft with uniform noise on the target features fed to the EAGLE-1/2 draft through recipe_args.feature_noise, defaulting to the paper value 0.1. (#2470)
  • On-policy regeneration. Interleaves EAGLE-3 training with online target regeneration through a recipe_args.regen block that interleaves EAGLE-3 training with online target regeneration on a reserved GPU, hot-swapping the dataloader across all ranks in lockstep. (#3042)
  • Serving, benchmarking, and caches. Covers sglang and vllm target backends for training supervision, serve_vllm for trained EAGLE-3, P-EAGLE, and DFlash-family drafts, bench_vllm/bench_sweep acceptance and speedup benchmarks, and distributed DSpark offline precompute for targets too large for one node. (#2798, #2841, #2912, #2953)
  • Metrics and resume. Reports simulated accept length, per-position accept_rate@k, and periodic real accept length through decode_eval, and rejects a mismatched mask_token_id, a stale draft-vocab mapping, and mixed precompute shards on resume. (#2959, #3037, #3237, #3238)
Training, Parallelism, and Performance
  • Block-diagonal variable-length context parallelism. Introduces a blockdiag_cp implementation that shards packed multi-document sequences while keeping masking block-causal per document, driving varlen kernels from precomputed cu_seqlens with a fused K/V all-gather. (#2989, #3223)
  • Unified context-parallel input preparation. Returns a ContextParallelSharder from a single prepare_model_inputs_for_cp hook and replaces cp_utils with components.distributed.context_parallel. (#2937)
  • Frame-level vision-tower sharding. Partitions frame units across the CP group through distributed.multimodal.vision.frame_sharding and reassembles embeddings with a differentiable variable-length all-gather. (docs, #2990, #3186)
  • Attention backends. Selects flash_attention_3 and flash_attention_4 with a fa4 → fa3 → fa2 → sdpa → eager fallback ladder, MagiAttention as both a kernel and a context-parallel backend with CP1/CP2 Qwen3-MoE-30B packed recipes, and a CuTeDSL FFPA backend for Gemma 4's head_dim=512 layers (8K, packed CP8 16K). (#2384, #2929, #3070)
  • QuACK kernels and partial CUDA graphs. Exposes quack options for backend.linear, backend.rms_norm, and backend.rope, and backend.cuda_graph.modules capture of attn, moe_router, and moe_preprocess while dispatch, expert compute, and combine stay eager. (#2917, #3115)
  • Activation checkpointing. Accepts activation_checkpointing: selective for DDP and an activation_checkpointing_scope selecting which layer groups are wrapped, and saves the MoE router and its top-k output so recompute preserves expert assignments. (#2786, #2840, #3140, #3247)
  • Cross-entropy memory. Supports FusedLinearCrossEntropy under pipeline parallelism and bounds ChunkedCrossEntropy memory with a kernel that saves only original-dtype logits and recomputes each softmax chunk in backward. (#2927, #2996)
  • FSDP2 Wraps HF transformer layers with PyTorch's non-reentrant checkpoint_wrapper to remove duplicate prefetch all-gathers, and adds distributed.multimodal.frozen_sharding for fully frozen towers. (#3328, #3411, #3513)
  • Gradient correctness. Normalizes the MoE auxiliary loss by the gradient-accumulation and pipeline-microbatch count, all-reduces grad-norm accumulators on the mesh device, and preserves HSDP replica gradient synchronization for MoE. (#3135, #3359, #3461)
  • Damaged-embedding repair. Replaces, through an optional embedding_row_repair: section, input-embedding rows whose L2 norm is non-finite or below min_norm with the scaled output-embedding direction; it is rejected under pipeline parallelism. (docs, #3136)
  • Setup-time prewarms and mesh timeouts. Initializes lazily created resources through an opt-in prewarm: section covering cublas_backward, fla_gdn_autotune, mamba_ssd_autotune, and comm_groups, and applies dist_env.timeout_minutes to flattened and expert-parallel meshes. (#2846, #2992, #3296)
  • Rollout Routing Replay. Reuses the rollout's discrete top-k expert selection in the training forward through MoEConfig.enable_routing_replay, which reuses the rollout's discrete top-k expert selection in the training forward to remove the rollout/training routing mismatch in on-policy RL. (#2797)
  • Optimizer parameter groups and throughput. Adds optimizer.param_group_overrides with lr_mult and wd_mult multipliers, removes Python overhead from Transformer Engine attention and MoE expert hot paths, and shapes TPLinear/LinearLoRA graphs so async-TP fusion fires. (#2987, #3046, #3374)
  • PEFT under pipeline and expert parallelism. Gathers adapter weights across pipeline stages before writing and saves optimizer state for every PEFT model part namespaced by stage. (#3096, #3250, #3316)
  • PEFT v5 expert adapters. Exports MoE expert LoRA through an opt-in PEFT v0.18+ ParamWrapper export for MoE expert LoRA with fused target_parameters, validated on Qwen3-MoE, MiniMax-M2, and Nemotron v3. (#3439, #3458)
  • LoRA correctness. Makes use_memory_efficient_lora: false also suppress the fused SwiGLU/ReLU² LoRA MLP, restores adapters through DDP wrappers, and reaches the absorbed GLM MLA KV weight through a new materialize_effective_weight(). (#3126, #3150, #3176, #3470)
Checkpointing
  • Preemption checkpointing. Introduces step_scheduler.preemption_signal, which watches one or more signals (default SIGTERM), gathers the flag across ranks, and saves at the next step boundary before exiting cleanly. (docs, #3007)
  • Cloud checkpoint directories. Accepts msc:// paths for sharded DCP state; cloud roots require save_consolidated: false, reject max_recent_checkpoints, and keep RNG, dataloader, and pointer files local. (#1709)
  • Retention and interrupted saves. Prunes older directories through max_recent_checkpoints, and marks in-progress directories until every component is published so resume falls back instead of hanging all ranks. (#2416, #3261)
  • Async consolidation and new knobs. Runs consolidated safetensors export through the same all-rank consolidation as synchronous mode on a background thread, and adds cpu_offload, wait_for_staging, and consolidation_timeout_minutes. (#3108, #3125, #3130, #3131)
  • torchsave export to Hugging Face. Adds scripts/export_llm_dcp_to_hf.py, which rebuilds the original topology from the checkpoint's recorded config.yaml and dispatches the LLM or VLM recipe named there, so a VLM checkpoint keeps its vision tower. (#2487)
  • Export fidelity and safety. Keeps fp32 parameters fp32 even under an explicit cast_dtype, writes safetensors without append, retains later-stage keys in pipeline-parallel consolidation, and loads auxiliary torch.save state with weights_only=True. (#2627, #3193, #3240, #3248)
  • Per-global-rank RNG state. Saves and restores RNG state per global rank rather than per data-parallel rank, so tensor- and pipeline-parallel peers keep distinct streams across a resume. (#3437)
  • Tied-embedding guards. Declares a TieSupport policy per model class and rejects an unsupported tie_word_embeddings value. (#2805, #2896, #2998)
Data
  • Typed dataset configs. Gives every shipped dataset a typed *Config with a build method and resolves dataset: and dataloader: into a single DataloaderConfig; legacy _target_ values still resolve through a compatibility registry. (#2390)
  • Prefix-tree attention for rollouts. Folds, through prefix_tree_collate_fn, one shared prompt and its completions into a single deduplicated sequence with a block-sparse mask so the prompt is encoded once. (#2564)
  • Long-context agent-SFT data pipeline. Tokenizes CoderForge trajectories once in a dedicated stage and drops over-length ones rather than truncating, with a paired Gemma-4-31B recipe at seq_length: 65536 and cp_size: 8. (#3151, #3446)
  • Chat-template masking and pre-tokenization. Validates prefix-built answer-only masks against the full-conversation render, keeps system turns in the ShareGPT converter, and adds packed_sequence.num_proc for parallel pre-tokenization before packing. (#3024, #3045, #3144)
Embedding and Re-ranker
  • Normalized Arrow retrieval datasets. Reads, through NormalizedRetrievalDatasetConfig, a portable Arrow bundle that stores each referenced document or image once, with a CPU-side preparation script and Slurm wrappers. (docs, #2596)
  • NVIDIA: Nemotron VL 1B. Adds a native llama_nemotron_vl model and processor for nvidia/llama-nemotron-embed-vl-1b-v2, trained as a vision bi-encoder with a bi-encoder recipe and a ColPali conversion notebook. (docs, #2354)
  • Cross-encoder re-ranking. use_text_in_document applies to the cross-encoder transform as well as the bi-encoder, so a re-ranker can score an image document together with its text. (#3189)
  • Embedding distillation. Distills a teacher embedding model into a student with EmbeddingDistillRecipe and an 8-GPU example, mixing cosine, MSE, and listwise InfoNCE terms with support for intermediate-layer distillation and cross-tokenizer cached teachers. (#3058)
  • Ministral 3 recipes. Loads the stock bidirectional Ministral backbone instead of a bespoke extraction, repairs the bi-encoder recipe, and guards the optional Weights & Biases import so the retrieval recipes import without it. (#3103, #3380, #3441)
  • Sentence Transformers metadata. Writes modules.json, 1_Pooling/config.json, and config_sentence_transformers.json alongside consolidated bi-encoder exports, which the hard-negative miner reads back for pooling and normalization. (#3464)
Distillation
  • Distillation on separate meshes. Places the teacher on its own ranks through separate_meshes: true and a teacher_distributed: block, with Llama-3.2 1B/3B and Qwen3.5 4B/9B examples under TP2, CP2, PP2, or DP2 teacher meshes. (#2954, #3280)
  • Distillation correctness. Resolves fp32 master weights for the student, runs gradient clipping on the recipe's own mesh with the expert TP replication factor, and passes tensor-valued hidden states through unchanged. (#3019, #3302, #3324)
Packaging and Dependencies
  • Framework pins. Moves transformers from 5.8.1 to 5.12.1 and megatron-fsdp from 0.2.3 to a pinned 0.5.0 resolved from a Megatron-LM fork commit; both force an environment rebuild. (#2873)
  • QuACK is a default dependency. Promotes quack-kernels==0.6.1 to the base dependencies on Linux, which pulls nvidia-cutlass-dsl, apache-tvm-ffi, and torch-c-dlpack-ext into every Linux install. (#3115)
  • New optional dependency sets. Adds an ffpa extra with ffpa-attn and nvidia-cutlass-dsl, and ships MagiAttention as a uv dependency group installed with uv sync --group magi; neither is part of [all]. (#2436, #3070, #3225)
  • Container. Builds FlashAttention 3 from source on x86 by default and leaves FlashAttention 4 opt-in behind INSTALL_FA4=true, moves the optional PyTorch base stage to pytorch:26.06-py3, and moves the deploy image to vllm:26.04-py3. (#2929, #2983, #3061, #3329)
  • Security constraints. Advances aiohttp, cryptography, gitpython, pillow, starlette, mlflow, and thrift, and adds pyarrow and ray constraints. (#3346, #3398, #3482, #3523)
Media dependencies

Media packages remain opt-in. The vlm-media extra adds albumentations for the Nemotron-Parse processor and diffusion-media adds PyAV for LTX-2 audio decoding; the diffusion extra now requires diffusers>=0.39.0:

pip install 'nemo-automodel[media]'

Use nemo-automodel[vlm-media] or nemo-automodel[diffusion-media] when only one of those media stacks is required.

Breaking Changes
  • tie_word_embeddings is validated per model family and raises instead of leaving a randomly initialized lm_head. (#2805, #2896, #2998)
  • Tokenizer BOS/EOS insertion is opt-in; set add_bos_token / add_eos_token to keep the previous tokenization. (#3337)
  • Multi-turn masking raises on chat templates that rewrite earlier turns; affected templates must carry {% generation %} blocks. (#3024, #3051)
  • Unrecognized keys under dataloader: or on a typed dataset config now raise instead of being silently forwarded. (#2390)
  • DFlash requires an explicit mask_token_id, and loss_decay_gamma defaults to 7.0. (#2461, #2908)
  • BaichuanForCausalLM, Qwen2ForCausalLM, and KimiVLForConditionalGeneration are deprecated and scheduled for removal in 26.10. (#2807, #2884)
External Contributors 🎉

This release includes work from the following contributors outside the NVIDIA-NeMo organization. Thank you.

  • @khazic — most of the speculative-decoding stack in this release: the DSpark, Domino, and JetSpec drafts, ViSpec VLM draft training, sequence packing and target tensor/context parallelism across every draft family, the vLLM and SGLang target backends, serving and acceptance benchmarks, and Gemma 4 ring context parallelism (94 PRs). (#2449, #2810, #2819, #2867, #3173)
  • @Achyuthan-S — the TieSupport policy and tie_word_embeddings guards across the registered model families. (#2732, #2805, #2896, #2998)
  • @kashif — the DSpark draft model and training objective, context parallelism for draft training, and DSpark acceptance and confidence metrics. (#2810, #2918, #2957)
  • @edjson — MSC cloud-storage support for DCP checkpoints and preemption checkpointing. (#1709, #3007)
  • @Butterfingrz — the FFPA headdim=512 attention backend for Gemma 4. (#2436)
  • @beccohov — FusedLinearCrossEntropy support under pipeline parallelism. (#2927)
  • @fkuner — the DSpark offline target-cache path. (#2924)
  • @GITsologun — pre-extracted video frame sequences in the VLM data path. (#3211)
  • @hyfine — gathering PEFT adapters across pipeline stages before save. (#3096)
  • @wangzhxg — preserving HSDP replica gradient synchronization for MoE models. (#3135)
  • @shahafwa — applying LoRA to the TileLang MLA KV projection. (#3176)
  • @amolkhanna — activation checkpointing for Qwen3-Next linear-attention layers. (#3192)
  • @aminehd — running Gemma 4 SDPA in fp32 to avoid NaNs on Hopper. (#3141)
  • @huahuajhu — writing consolidated safetensors without append. (#2627)
  • @zhiqi-li — omitting the unset FSDP reshard argument. (#2926)
  • @DOGEUNNKIM — casting Gemma 4 dense parameters without casting buffers. (#2359)
  • @grgkovac — per-validation-dataset logging in Weights & Biases. (#2526)
Known Issues
  • Transformer Engine fused RoPE is force-disabled globally and overrides an explicit rope_fusion: true. (#3028)
  • Inter-node DeepEP dispatch still faults and deepep remains the default whenever DeepEP is importable; use HybridEP for inter-node EP.
  • FlashAttention 4 and the TileLang backend cannot coexist in one image; an INSTALL_FA4=true build loses the TileLang backend used by DeepSeek-V4 and GLM MoE DSA. (#2929)
  • P-EAGLE rejects a remote target backend, a cached target path, and packed sequences. (#2466)
  • Qwen-Image-Edit-2511 training is limited to cached latents on a single node, and diffusion context parallelism accepts only pure Ulysses sharding. (#3157, #3217)
Changelog Details
  • fix(speculative): serialize remote EAGLE-3 /generate end to end by @khazic :: PR: #2479
  • feat(speculative): add EAGLE feature-noise augmentation to EAGLE-1/2 training by @khazic :: PR: #2470
  • fix(speculative): gate unsupported P-EAGLE backend / cache / packing combos by @khazic :: PR: #2466
  • feat(dllm): add DiffusionGemma 26B-A4B block-diffusion SFT (full + LoRA) by @zyzhou5 :: PR: #2506
  • test(checkpoint): fix TestFormatLoad directory-based reader selection by @khazic :: PR: #2520
  • feat(speculative): add tqdm training progress bar to EAGLE and DFlash recipes by @khazic :: PR: #2522
  • feat(vlm): add Gemma 4 31B joint drafter example config by @khazic :: PR: #2620
  • fix: skip fused LoRA MLP install for meta weights by @akoumpa :: PR: #2775
  • fix(checkpoint): resolve tie_word_embeddings top-level-first to match HF tying by @Achyuthan-S :: PR: #2732
  • fix(benchmark): skip unsupported MTP flops by @akoumpa :: PR: #2767
  • fix: Qwen3.5 MedPix EP32 NCCL timeout by @akoumpa :: PR: #2777
  • fix(ci): address go-git/go-billy and rustls-webpki CVEs by @thomasdhc :: PR: #2780
  • fix(ci): stabilize diffusion finetune smoke tests by @pthombre :: PR: #2788
  • build(deps): move ffmpeg/opencv deps to opt-in media extra by @thomasdhc :: PR: #2743
  • fix(vlm): keep Qwen3.5 media tokens aligned by @yuhezhang-ai :: PR: #2772
  • fix(ci): HybridEP for multi-node MoE benchmarks + LoRA OOM fixes by @hemildesai :: PR: #2789
  • fix: qwen3.5 and 3.6 mtp expert checkpoint layout by @HuiyingLi :: PR: #2778
  • feat: CP support for MiniMax M3 by @athitten :: PR: #2551
  • docs(speculative): fix EAGLE drafter layer count and EAGLE-1 loss by @khazic :: PR: #2796
  • fix(ci): drop base-image uv/wandb copies flagged for CVEs by @thomasdhc :: PR: #2800
  • chore(skills): refresh automodel skill signatures by @akoumpa :: PR: #2804
  • feat(distributed): enable selective checkpointing for DDP by @yuhezhang-ai :: PR: #2786
  • docs: align nightly navigation routes by @akoumpa :: PR: #2801
  • feat: vision biencoder finetuning + Nemotron VL 1B finetuning support by @gabrielspmoreira :: PR: #2354
  • fix(docs): catch Fern MDX syntax errors by @akoumpa :: PR: #2806
  • feat(moe): add Rollout Routing Replay (R3) for MoE RL training by @khazic :: PR: #2797
  • fix(optim): align Dion mesh with FSDP sharding by @akoumpa :: PR: #2808
  • fix(ci): stabilize failed benchmark recipes by @akoumpa :: PR: #2817
  • docs: document opt-in media extras (vlm-media/diffusion-media) by @thomasdhc :: PR: #2799
  • feat(speculative): add SGLang target backend for EAGLE-3 training by @khazic :: PR: #2449
  • feat(speculative): add Domino online training path on top of DFlash by @khazic :: PR: #2819
  • feat(speculative): add dspark draft model and training objective by @kashif :: PR: #2810
  • docs: add GLM-5.2 and speculative decoding (DSpark) updates by @khazic :: PR: #2828
  • feat(diffusion): support Hugging Face datasets by @pthombre :: PR: #2816
  • feat(speculative): add vLLM target backend for EAGLE-3 training by @khazic :: PR: #2798
  • fix(speculative): gather sharded target lm_head in EAGLE-1/2 token loss by @khazic :: PR: #2823
  • fix(model): honor yarn/linear/dynamic RoPE in LlamaRotaryEmbedding by @khazic :: PR: #2825
  • fix(speculative): keep DFlash padding blocks self-attending to avoid NaN by @khazic :: PR: #2826
  • feat(speculative): support target tensor parallelism in EAGLE-3 colocated path by @khazic :: PR: #2827
  • feat(speculative): support target tensor parallelism in EAGLE-1/2 by @khazic :: PR: #2829
  • feat(speculative): support target tensor parallelism in DFlash by @khazic :: PR: #2830
  • fix(speculative): validate DFlash mask_token_id on checkpoint resume by @khazic :: PR: #2824
  • docs: enable Fern multi-source by @lbliii :: PR: #2845
  • fix(distributed): propagate NCCL timeout to derived device meshes by @HuiyingLi :: PR: #2846
  • feat(speculative): add serve_vllm for EAGLE-3 / P-EAGLE drafts by @khazic :: PR: #2841
  • fix(speculative): keep DSpark draft RoPE inv_freq in fp32 under bf16 by @khazic :: PR: #2859
  • feat(speculative): compress EAGLE-3 offline cache target_probs via top-k by @khazic :: PR: #2847
  • perf: FFPA D=512 attention backend for Gemma4 (3× fwd / 6× bwd vs SDPA) by @Butterfingrz :: PR: #2436
  • feat(speculative): add JetSpec causal parallel drafting training by @khazic :: PR: #2867
  • feat(models): reject tie_word_embeddings=True on separate-head model families by @Achyuthan-S :: PR: #2805
  • ci: Ensure diffusion-media extra installed for diffusion tests by @chtruong814 :: PR: #2858
  • feat(speculative): add DeepSeek-V3 (MLA) EAGLE-3 draft model by @khazic :: PR: #2849
  • feat(speculative): add DeepSeek V4 DSpark drafter and V4-Flash training by @khazic :: PR: #2866
  • chore(models): add 26.10 deprecation warnings for custom model classes by @athitten :: PR: #2807
  • feat(speculative): add DSpark draft for MiniMax M3 VL (text + multimodal) by @khazic :: PR: #2877
  • ci: use NVIDIA inference for Claude review by @chtruong814 :: PR: #2882
  • feat(speculative): add GLM-5.2 DSpark draft model and training by @khazic :: PR: #2885
  • fix(checkpoint): write consolidated safetensors without append by @huahuajhu :: PR: #2627
  • ci(automodel): set AM-576 release timeouts by @yuhezhang-ai :: PR: #2871
  • fix(diffusion_gemma): make parity test work on transformers >= 5.11 by @zyzhou5 :: PR: #2883
  • fix(bagel): torch2.12 DTensor sharding-prop leak by @zyzhou5 :: PR: #2888
  • fix(speculative): plumb EAGLE-3.1 fc_norm/norm_output into the draft config by @khazic :: PR: #2897
  • fix(speculative): make save_consolidated final export the last EAGLE checkpoint by @khazic :: PR: #2898
  • fix(speculative): gather TP target outputs in the Domino trainer by @khazic :: PR: #2899
  • fix(speculative): default DFlash/Domino loss_decay_gamma to the paper value by @khazic :: PR: #2908
  • refactor(speculative): dispatch all DSpark drafts through the registry by @khazic :: PR: #2909
  • feat(speculative): serve DFlash and JetSpec drafts on vLLM by @khazic :: PR: #2913
  • docs(speculative): document DSpark, Domino, JetSpec, and the vLLM serve path by @khazic :: PR: #2914
  • feat(speculative): add bench_vllm acceptance/speedup benchmark by @khazic :: PR: #2912
  • fix(speculative): read DeepSeek EAGLE-3 draft rope_theta from rope_parameters by @khazic :: PR: #2920
  • fix(speculative): send chat_template_kwargs as a top-level request field by @khazic :: PR: #2921
  • ci(deps): add albumentations to vlm-media for nemotron-parse processor by @thomasdhc :: PR: #2933
  • fix(glm): select HybridEP for GLM-5.2 recipes by @HuiyingLi :: PR: #2934
  • fix(distributed): omit unset FSDP reshard argument by @zhiqi-li :: PR: #2926
  • feat(speculative): add DSpark offline cache path by @fkuner :: PR: #2924
  • chore: update readme for 26.06-26.08 by @akoumpa :: PR: #2886
  • ci: validate ci section on newly added recipes by @thomasdhc :: PR: #2936
  • fix(bagel): make non-backbone init topology independent by @zyzhou5 :: PR: #2939
  • feat(speculative): add fp8 draft training and LoRA draft adaptation by @khazic :: PR: #2963
  • feat(speculative): report simulated accept length during EAGLE-3 training by @khazic :: PR: #2959
  • fix(distributed): fix and scope vlm activation checkpointing by @yuhezhang-ai :: PR: #2840
  • feat(models): complete tie_word_embeddings guards for #2512 families by @Achyuthan-S :: PR: #2896
  • feat(speculative): add distributed DSpark offline precompute by @khazic :: PR: #2953
  • ci: route gb200 L2_HF_DCP to the IMEX runner for DeepEP NVLink P2P by @ko3n1g :: PR: #2968
  • docs: add 0.5.0 release notes by @lbliii :: PR: #2881
  • fix(vlm): gemma4 FFPA mock + e4b cp16 CI recipe fixes (AM-626, AM-627) by @athitten :: PR: #2949
  • build(deps): upgrade transformers to 5.12.1 by @athitten :: PR: #2873
  • ci: update claude review guidelines by @akoumpa :: PR: #2739
  • fix(moe): preserve MLP dispatch through checkpoint wrappers by @akoumpa :: PR: #2955
  • feat(speculative): add multi-dataset acceptance-length benchmark sweep by @khazic :: PR: #2966
  • fix(transformers): gate _tie_weights_nemo on tie_word_embeddings flag by @yuhezhang-ai :: PR: #2942
  • fix(loss): avoid pkg_resources in linear CE by @akoumpa :: PR: #2700
  • fix(deepseek_v4): support packed THD with context parallel by @HuiyingLi :: PR: #2731
  • feat(speculative): FSDP2-shard dense DSpark target + add Gemma4-31B config by @khazic :: PR: #2976
  • fix(attention): initialize FlexAttention masks by @akoumpa :: PR: #2982
  • fix(deepseek-v4): initialize random training state by @akoumpa :: PR: #2991
  • ci(review): strengthen model review invariants by @akoumpa :: PR: #2994
  • ci: Bump claude_review to v1.8.2 by @chtruong814 :: PR: #3012
  • ci: AUT-779 repin _claude_review.yml to v1.8.3 by @svcnemo-autobot :: PR: #3013
  • test(ci): accept SHA-pinned claude-review template refs in the policy test by @khazic :: PR: #3022
  • feat(speculative): context parallelism for draft-model training by @kashif :: PR: #2918
  • feat(pp): support FusedLinearCrossEntropy under pipeline parallelism by @beccohov :: PR: #2927
  • feat(speculative): support sequence packing in the DeepSeek MLA EAGLE-3 draft by @khazic :: PR: #3002
  • fix(datasets): reject non-prefix multiturn mask templates by @khazic :: PR: #3024
  • feat(speculative): support sequence packing in the EAGLE-1/2 draft by @khazic :: PR: #3003
  • feat(speculative): add variable-prefix DFlash and LK EAGLE-3 training losses by @khazic :: PR: #3032
  • feat(dspark): log acceptance and confidence metrics during training by @kashif :: PR: #2957
  • feat(speculative): support sequence packing in the DFlash draft by @khazic :: PR: #3004
  • feat(speculative): add periodic real accept-length eval by @khazic :: PR: #3037
  • feat(speculative): support sequence packing in the Domino and JetSpec drafts by @khazic :: PR: #3023
  • feat(speculative): support sequence packing in the DSpark Qwen3 draft by @khazic :: PR: #3005
  • fix(models): temporarily disable TE fused RoPE globally by @HuiyingLi :: PR: #3028
  • fix(glm_moe_dsa): clamp DSA indexer top-k for short sequences by @HuiyingLi :: PR: #3035
  • feat(eagle3): on-policy regeneration loop with step-cadence dataloader swap by @khazic :: PR: #3042
  • fix(docs): update technical accuracy by @akoumpa :: PR: #3031
  • fix(checkpoint): compare torch versions semantically by @HuiyingLi :: PR: #3043
  • docs: add uv run prefix by @akoumpa :: PR: #3057
  • fix(recipe): add generation-marked chat templates by @khazic :: PR: #3051
  • ci: include cutlass deps by @akoumpa :: PR: #3056
  • fix(dspark): omit unmeasured positional acceptance rates by @khazic :: PR: #3050
  • test(transformers): add source-load parity coverage by @yuhezhang-ai :: PR: #2960
  • fix(distributed): resolve VLM AC layer-group paths structurally, not by version by @yuhezhang-ai :: PR: #3025
  • fix(vlm): load Tulu-3 directly from HF Hub instead of local meta JSON by @athitten :: PR: #3001
  • fix(optim): drop zero-numel local DTensor shards before TE FusedAdam by @yuhezhang-ai :: PR: #2997
  • fix(moe): checkpoint trainable vision towers on the expert-parallel path by @yuhezhang-ai :: PR: #2993
  • feat(datasets): add optional parallel pre-tokenization before packing by @khazic :: PR: #3045
  • fix(models): handle checkpoint-wrapped fp32 buffers by @akoumpa :: PR: #3059
  • feat(optim): support per-parameter-group learning rate by @khazic :: PR: #3046
  • fix(cp): preserve gradients when sharding VLM inputs by @HuiyingLi :: PR: #2931
  • refactor(vlm): gate build_model via recipe-side model-target allowlist by @athitten :: PR: #2357
  • fix(ci): use HybridEP for Step 3.5 benchmark by @HuiyingLi :: PR: #3069
  • fix: Kimi K2 config loading without remote code by @HuiyingLi :: PR: #3065
  • feat(speculative): add Gemma4 EAGLE-3 target support by @khazic :: PR: #3071
  • fix(dllm): remove obsolete DiffusionGemma HF compatibility by @zyzhou5 :: PR: #3067
  • feat(speculative): add Gemma4-E4B EAGLE-3 example config by @khazic :: PR: #3073
  • fix(test): use SDPA for Nemotron-H HF reload by @yuhezhang-ai :: PR: #3060
  • feat(speculative): add Gemma4-31B EAGLE-3 example config by @khazic :: PR: #3077
  • feat(speculative): add Gemma4-26B-A4B MoE EAGLE-3 example config by @khazic :: PR: #3079
  • chore(ci): AUT-852 bump claude review template to v1.8.4 by @svcnemo-autobot :: PR: #3087
  • feat: add Ministral3 embedding distillation recipe by @vinay-raman :: PR: #3058
  • fix(perf): use GPT-OSS head dimension in FLOPs accounting by @yaoyu-33 :: PR: #3091
  • test(speculative): add EAGLE-3 fp8 draft convergence smoke for SM89+ by @khazic :: PR: #3081
  • fix(speculative): keep DSpark target_layer_ids in [0, N-2] for SGLang by @khazic :: PR: #3083
  • chore(skills): add Regent Open Plugin manifest by @ko3n1g :: PR: #3097
  • chore(skills): remove Open Plugin manifest (superseded) by @ko3n1g :: PR: #3099
  • test: clean up KD test process group by @yuhezhang-ai :: PR: #3090
  • refactor(models): single tie_word_embeddings guard via TieSupport by @Achyuthan-S :: PR: #2998
  • feat(speculative): add DFlash validation metrics by @khazic :: PR: #3072
  • refactor(datasets): typed Config + build per dataset, config-driven dataloader by @akoumpa :: PR: #2390
  • fix(bagel): enable periodic garbage collection by @zyzhou5 :: PR: #3105
  • fix(ci): pass PR number and commit to Codecov for fork PRs by @svcnemo-autobot :: PR: #3101
  • fix(test): use eager attention for Nemotron-H HF loads by @yuhezhang-ai :: PR: #3100
  • feat(kd): support separate student and teacher meshes by @akoumpa :: PR: #2954
  • fix(moe): safe TP and EP/TP gradient correctness for custom MoE models by @yuhezhang-ai :: PR: #2995
  • test(ci): remove deprecated 26.10 models from nightly and release CI by @athitten :: PR: #2884
  • feat(training): add opt-in setup-time prewarms by @yuhezhang-ai :: PR: #2992
  • fix: bound chunked CE memory and preserve packing masks by @yuhezhang-ai :: PR: #2996
  • fix(ci): restore Ministral3 checkpoint robustness by @yuhezhang-ai :: PR: #3111
  • ci: add HF hub cache preflight check by @thomasdhc :: PR: #3119
  • fix(vlm): resolve get_rope_index from the base model for packed mRoPE by @khazic :: PR: #3113
  • feat(models): add Inkling VLM MoE support by @hemildesai :: PR: #3095
  • ci(dllm): add dLLM SFT nightly train-to-generate launcher and recipes by @zyzhou5 :: PR: #2783
  • feat(vlm): add THD packed-sequence support for Qwen3-VL-MoE by @khazic :: PR: #3052
  • docs: Add Inkling README news by @HuiyingLi :: PR: #3129
  • fix(checkpoint): gather PEFT adapter across PP stages by @hyfine :: PR: #3096
  • build(deps): bump base container to 26.06-cuda13.3 by @thomasdhc :: PR: #2983
  • fix(transformers): keep custom M3 config when transformers ships its own by @HuiyingLi :: PR: #3134
  • ci: AUT-911 shard GPU unit tests into 5 pytest-shard chunks by @svcnemo-autobot :: PR: #3138
  • fix(checkpoint): isolate consolidation timeout from NCCL by @yuhezhang-ai :: PR: #3108
  • test(vlm): add checkpoint robustness coverage by @yuhezhang-ai :: PR: #3112
  • fix(ci): preserve retrieval evaluation schedule by @yuhezhang-ai :: PR: #3139
  • feat(bagel): add TE support by @zyzhou5 :: PR: #2895
  • fix(moe): preserve HSDP replica gradient synchronization by @wangzhxg :: PR: #3135
  • docs: add Laguna SFT docs and recipe by @HuiyingLi :: PR: #3146
  • ci: AUT-895 gate GB200 tests on DISABLE_GB200_TESTS variable by @svcnemo-autobot :: PR: #3118
  • feat(dllm): add LLaDA2 generation support by @zyzhou5 :: PR: #3092
  • feat(models): add Laguna model implementation by @HuiyingLi :: PR: #3148
  • fix: prevent damaged token embeddings from dominating grad clipping by @akoumpa :: PR: #3136
  • fix(checkpoint): add bounded retention window by @oliverholworthy :: PR: #2416
  • fix(kd): use fp32 master weight copy by @akoumpa :: PR: #3019
  • fix(checkpoint): restore adapters through DDP wrappers by @yuhezhang-ai :: PR: #3150
  • feat(distributed): block-diagonal varlen CP for packed sequences by @yuhezhang-ai :: PR: #2989
  • fix(distributed): preserve canonical activation checkpoint keys by @yuhezhang-ai :: PR: #3152
  • feat(examples): Gemma4-31B CoderForge data pipeline + CP SFT recipes by @athitten :: PR: #3151
  • feat(attention): support flash_attention_3 and flash_attention_4 by @HuiyingLi :: PR: #2929
  • feat(checkpoint): add DCP CPU offload option by @khazic :: PR: #3130
  • test(packing): fix stale unsupported-backend test after fa3/fa4 support by @khazic :: PR: #3199
  • test(ci): add DeepSeek V4 Flash pretrain coverage by @akoumpa :: PR: #3128
  • refactor(distributed): unify CP input prep and dispatch across models by @HuiyingLi :: PR: #2937
  • docs(fern): add legacy URL redirects by @akoumpa :: PR: #3203
  • ci: reduce L0 GPU tests to two shards by @akoumpa :: PR: #3201
  • feat: add checkpoint staging wait option by @khazic :: PR: #3131
  • fix(datasets): keep system turns in the sharegpt conversation converter by @khazic :: PR: #3144
  • fix(peft): honor memory-efficient LoRA opt-out by @akoumpa :: PR: #3126
  • feat: enable text inclusion alongside images in retrieval training by @rnyak :: PR: #3189
  • feat: Support preemption checkpointing by @edjson :: PR: #3007
  • fix(ci): AUT-964 restore Codecov after skipped GB200 jobs by @svcnemo-autobot :: PR: #3205
  • fix(ci): preserve HF meta init for device-mapped loads by @yuhezhang-ai :: PR: #3188
  • feat(dllm): add lora to dllm by @zyzhou5 :: PR: #3163
  • fix(gemma4): run SDPA in fp32 to avoid #2208 NaN on Hopper by @aminehd :: PR: #3141
  • docs(retrieval): complete fine-tuning guide integration by @oliverholworthy :: PR: #2306
  • fix(glm): apply LoRA to TileLang MLA KV projection by @shahafwa :: PR: #3176
  • fix: prefer Automodel config registry lookup by @akoumpa :: PR: #3202
  • feat(distributed): shape TPLinear/LinearLoRA graphs for async-TP fusion by @yuhezhang-ai :: PR: #2987
  • fix(distributed): Megatron-FSDP 0.5.0 compatibility by @yuhezhang-ai :: PR: #2986
  • feat(checkpoint): preserve intrinsic fp32 during offline consolidation by @yuhezhang-ai :: PR: #3193
  • feat(kernels): add QuACK backend by @akoumpa :: PR: #3115
  • feat(dllm): add DiffusionGemma generation via the built-in HF sampler by @zyzhou5 :: PR: #3161
  • refactor(diffusion): migrate recipe onto typed RecipeConfig build() path by @pthombre :: PR: #3122
  • fix(dist): keep profiler record-function ops out of SAC replay accounting by @HuiyingLi :: PR: #3133
  • fix(vlm): size qwen3.6 medpix CI configs by @HuiyingLi :: PR: #3209
  • feat(glm_moe_dsa): expose update_moe_gate_bias on GLM MoE DSA models by @jQizhang :: PR: #3207
  • fix(ci): shard Nemotron Super vLLM deploy across 8 GPUs by @yuhezhang-ai :: PR: #3061
  • fix(glm): size GLM5.2 release CI jobs by @HuiyingLi :: PR: #3210
  • fix(distributed): activation-checkpoint Qwen3-Next linear_attn layers by @amolkhanna :: PR: #3192
  • fix(distributed): support packed CP for Llama, Qwen2, and Qwen3 by @akoumpa :: PR: #2999
  • fix(dllm): recipe defaults, corruption seeding, sampler and docs fixes by @zyzhou5 :: PR: #3162
  • perf(bagel): grouped MoT routing + fused SwiGLU/RoPE by @zyzhou5 :: PR: #3214
  • feat(retrieval): add normalized Arrow dataset tooling by @yuhezhang-ai :: PR: #2596
  • feat(diffusion): context parallelism via diffusers ContextParallelConfig by @pthombre :: PR: #3157
  • docs: explain embedding row repair by @akoumpa :: PR: #3175
  • fix(transformers): custom configs override builtin ones by default by @HuiyingLi :: PR: #3221
  • build(deps): bump ffpa-attn to 0.2.2 by @akoumpa :: PR: #3225
  • fix(distributed): checkpoint linear attention with compile by @yuhezhang-ai :: PR: #3213
  • fix(glm): own the GLM MoE DSA config so qk_rope_head_dim survives by @HuiyingLi :: PR: #3222
  • feat(speculative): add the multimodal speculative decoding (MSD) core by @khazic :: PR: #3166
  • fix(speculative): average DFlash train metrics over the log window by @khazic :: PR: #3234
  • fix(speculative): honor --dflash-causal on the vLLM remapping export by @khazic :: PR: #3235
  • fix(speculative): align the decode-eval cadence on resume by @khazic :: PR: #3237
  • fix(speculative): validate mask_token_id when resuming DSpark by @khazic :: PR: #3238
  • build: cache CUDA extension source builds by @akoumpa :: PR: #3204
  • docs(fern): cut v0.5.0 version train by @lbliii :: PR: #2970
  • fix(checkpoint): infer Qwen MTP layout from checkpoint keys by @HuiyingLi :: PR: #3229
  • docs: fix remaining broken link sources by @akoumpa :: PR: #3249
  • fix(distributed): control frozen multimodal FSDP sharding by @yuhezhang-ai :: PR: #2763
  • fix(ci): preserve main wheelhouse cache builds by @akoumpa :: PR: #3253
  • feat(diffusion): add Qwen-Image-Edit-2511 training support by @pthombre :: PR: #3217
  • fix(ci): omit token type IDs from vLLM parity prompts by @yuhezhang-ai :: PR: #3198
  • fix(deepseek-v4): accept cu_seqlens for THD packing by @HuiyingLi :: PR: #3231
  • fix(deepseek-v4): preserve fp32 model dtype by @HuiyingLi :: PR: #3227
  • fix(thd): derive packed padding mask from the pack layout, not token ids by @HuiyingLi :: PR: #3223
  • fix(nemotron-v3): recompute deterministic LoRA router by @HuiyingLi :: PR: #3258
  • fix(moe): preserve top-k routing under full AC by @yuhezhang-ai :: PR: #3140
  • fix(gpt-oss): route packed attention through THD by @akoumpa :: PR: #3226
  • perf(benchmark): tune Qwen3-VL LoRA pipeline batches by @akoumpa :: PR: #3230
  • fix(moe): preserve gate load across activation recompute by @akoumpa :: PR: #3247
  • fix(ci): narrow CUDA wheelhouse cache keys by @akoumpa :: PR: #3269
  • feat(distributed): frame-level context-parallel vision-tower sharding by @yuhezhang-ai :: PR: #2990
  • fix(vlm): preserve mRoPE axes in PP chunking by @HuiyingLi :: PR: #3208
  • feat(diffusion): add LTX-2.3 video+audio finetuning (training + prepr… by @linnanwang :: PR: #3165
  • fix: harden checkpoint torch loads by @HuiyingLi :: PR: #3240
  • build: add MagiAttention optional dependency by @HuiyingLi :: PR: #3070
  • fix(checkpoint): harden PP safetensors consolidation by @yuhezhang-ai :: PR: #3248
  • fix(checkpoint): save all PEFT EP optimizer parts by @yuhezhang-ai :: PR: #3250
  • fix(ci): install CUDA torchvision with CUDA torch by @akoumpa :: PR: #3279
  • fix(docker): cap uv install concurrency to avoid ARM FD exhaustion by @thomasdhc :: PR: #3275
  • feat: support pre-extracted frame sequence of video sample by @GITsologun :: PR: #3211
  • fix(inkling): support HF 5.14 attention fields and embed norm FSDP by @HuiyingLi :: PR: #3281
  • fix(KD): Fix variable-length teacher PP batches on separate meshes by @HuiyingLi :: PR: #3280
  • feat: add Kimi K3 model support by @HuiyingLi :: PR: #3259
  • docs(models): add Kimi K3 coverage by @HuiyingLi :: PR: #3283
  • test(vlm): stabilize Qwen3.5 checkpoint robustness by @HuiyingLi :: PR: #3233
  • feat(vlm): integrate packed CP and vision sharding for Qwen3.5-MoE by @yuhezhang-ai :: PR: #3186
  • feat(checkpoint): defer distributed async consolidation by @khazic :: PR: #3125
  • test(retrieval): add Nemotron VL checkpoint coverage by @yuhezhang-ai :: PR: #3276
  • chore(ci): AUT-1135 pin GitHub Actions to commit SHAs by @svcnemo-autobot :: PR: #3277
  • feat: add ViSpec VLM draft training by @khazic :: PR: #3173
  • fix(docs): hide Kimi tokenizer regex from autodoc (3288) by @svcnvidia-nemo-ci :: PR: #3304
  • docs(distributed): review frozen multimodal FSDP guidance (3272) by @svcnvidia-nemo-ci :: PR: #3306
  • fix(vlm): expandable_segments:True for gemma4 31B FFPA 8k recipe (AMINT-203) (3273) by @svcnvidia-nemo-ci :: PR: #3307
  • fix(pp): use static metadata with PyTorch 2.13 (3290) by @svcnvidia-nemo-ci :: PR: #3305
  • ci: add Tulu-3 convergence + eval flow (3171) by @svcnvidia-nemo-ci :: PR: #3312
  • fix(distributed): preserve ERNIE router and dense Qwen3.5 SSM precision (3255) by @svcnvidia-nemo-ci :: PR: #3310
  • perf(moe): add scoped partial CUDA graphs (2917) by @svcnvidia-nemo-ci :: PR: #3314
  • fix(docs): avoid literal ampersand in Kimi autodoc (3317) by @svcnvidia-nemo-ci :: PR: #3320
  • fix(config): repair GLM-5.2 LoRA recipe (3313) by @svcnvidia-nemo-ci :: PR: #3323
  • fix(kd): preserve tensor-valued hidden states (3324) by @svcnvidia-nemo-ci :: PR: #3332
  • fix(kd): use mesh-safe gradient clipping (3302) by @svcnvidia-nemo-ci :: PR: #3334
  • refactor(docker): build torchao & FlashAttention as isolated wheel stages (3329) by @svcnvidia-nemo-ci :: PR: #3343
  • fix(distributed): restore Nemotron Flash TP2 training (3345) by @svcnvidia-nemo-ci :: PR: #3349
  • fix(nemotron-parse): sync RADIO preprocessing (3331) by @svcnvidia-nemo-ci :: PR: #3341
  • fix(distributed): reuse default group for world-sized meshes (3319) by @svcnvidia-nemo-ci :: PR: #3347
  • fix(training): prewarm Mamba SSD autotune kernels (3296) by @svcnvidia-nemo-ci :: PR: #3338
  • fix(docs): make Fern autodoc metadata MDX-safe (3355) by @akoumpa :: PR: #3357
  • fix(pp): preserve VLM media cursor with static metadata (3344) by @svcnvidia-nemo-ci :: PR: #3361
  • fix(deps): resolve 26.08 rc2 container CVEs (3346) by @svcnvidia-nemo-ci :: PR: #3363
  • fix(pp): route Qwen3.5 MoE pre-embedded inputs (3294) by @svcnvidia-nemo-ci :: PR: #3364
  • fix(kimi_k25_vl): don't int4 quantize LoRA adapter keys on save (3295) by @svcnvidia-nemo-ci :: PR: #3362
  • fix(diffusion): use spawn start method for GPU preprocessing pools (3339) by @svcnvidia-nemo-ci :: PR: #3366
  • fix(tokenizer): make special token insertion opt-in (3337) by @svcnvidia-nemo-ci :: PR: #3375
  • fix(distributed): stabilize SAC replay with TE and FSDP (3330) by @svcnvidia-nemo-ci :: PR: #3376
  • fix(ci): use a single container cache donor (3377) by @svcnvidia-nemo-ci :: PR: #3381
  • fix(fsdp): resolve fp32 master-weight compute dtype per parameter (3328) by @svcnvidia-nemo-ci :: PR: #3384
  • perf(checkpoint): reduce distributed save overhead (3369) by @svcnvidia-nemo-ci :: PR: #3389
  • fix(ci): enable LTX-2.3 diffusion finetuning (3372) by @svcnvidia-nemo-ci :: PR: #3391
  • ci: tune Nemotron single-GPU model load threads (3370) by @svcnvidia-nemo-ci :: PR: #3395
  • fix(model): correct Mistral4 attention and distributed MoE routing (3348) by @svcnvidia-nemo-ci :: PR: #3393
  • fix(checkpoint): survive an interrupted save instead of hanging all ranks (3261) by @svcnvidia-nemo-ci :: PR: #3394
  • perf: remove Python overhead from model hot paths (3374) by @svcnvidia-nemo-ci :: PR: #3401
  • fix(deps): resolve OSS CVE findings (3398) by @svcnvidia-nemo-ci :: PR: #3412
  • fix(moe): support EP-free DTensor state-dict conversion (3397) by @svcnvidia-nemo-ci :: PR: #3409
  • fix(checkpoint): harden PEFT PP checkpointing and Step-3.7 coverage (3316) by @svcnvidia-nemo-ci :: PR: #3403
  • docs(training): mark Mamba prewarm sections for review (3340) by @svcnvidia-nemo-ci :: PR: #3399
  • fix(retrieval): guard optional W&B imports (3380) by @svcnvidia-nemo-ci :: PR: #3402
  • fix(gemma4): enable expandable_segments on the CP tulu3 recipes (3408) by @svcnvidia-nemo-ci :: PR: #3410
  • feat: Add {% generation %} chat template for DiffusionGemma SFT/LoRA examples (3353) by @svcnvidia-nemo-ci :: PR: #3417
  • fix(docs): AUT-1324 qualify LTX model coverage slug (3418) by @svcnvidia-nemo-ci :: PR: #3419
  • fix(ci): run Kimi K3 HellaSwag on GB200 (3420) by @svcnvidia-nemo-ci :: PR: #3421
  • fix(transformers): default missing THD capability (3406) by @svcnvidia-nemo-ci :: PR: #3429
  • ci(convergence): (3416) by @svcnvidia-nemo-ci :: PR: #3432
  • cp: backport MoE checkpoint reload parity fixes to r0.6.0 by @yuhezhang-ai :: PR: #3438
  • fix(retrieval): use stock Ministral embedding backbone (3103) by @svcnvidia-nemo-ci :: PR: #3433
  • ci(minimax): extend M2.7 LoRA timeout (3428) by @svcnvidia-nemo-ci :: PR: #3430
  • fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (3411) by @svcnvidia-nemo-ci :: PR: #3445
  • fix(kimi_k25_vl): keep PEFT expert LoRA keys under the base_model.model. prefix (3431) by @svcnvidia-nemo-ci :: PR: #3451
  • ci(vlm): raise Slurm wall time for MiniMax-M3 and Gemma4 recipes (3448) by @svcnvidia-nemo-ci :: PR: #3453
  • fix: Gemma4-31B CoderForge CP8/64K data processing and recipe (3446) by @svcnvidia-nemo-ci :: PR: #3452
  • cp: backport PEFT v5 adapter output fixes to r0.6.0 by @yuhezhang-ai :: PR: #3458
  • fix(retrieval): repair Ministral3 training recipe (3441) by @svcnvidia-nemo-ci :: PR: #3454
  • test(checkpoint): redesign resume robustness around a shared trajectory (3427) by @svcnvidia-nemo-ci :: PR: #3457
  • test(checkpoint): stabilize MoE checkpoint parity gates (3440) by @svcnvidia-nemo-ci :: PR: #3460
  • fix(gemma4): exact router scalar and fp32 reference routing for MoE parity (3456) by @svcnvidia-nemo-ci :: PR: #3466
  • cp: backport Sentence Transformers metadata export to r0.6.0 by @yuhezhang-ai :: PR: #3464
  • fix: normalize MoE auxiliary loss during gradient accumulation (3359) by @svcnvidia-nemo-ci :: PR: #3459
  • fix(bagel): type SFT backend before AutoModel init (3449) by @svcnvidia-nemo-ci :: PR: #3469
  • fix(peft): enable v5 expert adapters for Nemotron and MiniMax (3439) by @svcnvidia-nemo-ci :: PR: #3468
  • fix(checkpoint): save ranked RNG state per global rank (3437) by @svcnvidia-nemo-ci :: PR: #3471
  • fix(peft): align frozen LoRA tensors with compute dtype (3470) by @svcnvidia-nemo-ci :: PR: #3480
  • cp: backport VLM CP gradient and vision sharding fixes to r0.6.0 by @yuhezhang-ai :: PR: #3479
  • fix(training): all-reduce grad-norm scalars on the mesh device (3461) by @svcnvidia-nemo-ci :: PR: #3481
  • fix(recipe): shard GPT-OSS 120B across 64 experts (3483) by @svcnvidia-nemo-ci :: PR: #3489
  • fix(ci): register flux2/wan2.2/qwen-image-edit diffusion recipes in CI (3485) by @svcnvidia-nemo-ci :: PR: #3487
  • perf(benchmarks): recompute deterministic MoE routers under AC (3474) by @svcnvidia-nemo-ci :: PR: #3490
  • chore: Bump gitpython to >= 3.1.59 (3482) by @svcnvidia-nemo-ci :: PR: #3488
  • perf(recipes): run GPT-OSS 120B with EP64 and no activation checkpointing by @akoumpa :: PR: #3492
  • cp: feat(models): make Inkling standalone (#3358) into r0.6.0 by @HuiyingLi :: PR: #3498
  • fix(models): preserve lm-head dtype boundaries (3491) by @svcnvidia-nemo-ci :: PR: #3502
  • fix(checkpoint): validate PEFT adapter-only state (3501) by @svcnvidia-nemo-ci :: PR: #3503
  • ci(convergence): fix gemma4 eval setup and re-baseline Qwen3-MoE (3493) by @svcnvidia-nemo-ci :: PR: #3505
  • fix(deepseek_v4): avoid TileLang boolx8 backward codegen (3467) by @svcnvidia-nemo-ci :: PR: #3506
  • fix(checkpoint): preserve FSDP2 mixed precision during recompute (3513) by @svcnvidia-nemo-ci :: PR: #3515
  • fix(deps): resolve Starlette and GitPython CVEs (3523) by @svcnvidia-nemo-ci :: PR: #3524
  • fix(minimax): declare packed-sequence and CP attention ownership by @athitten :: PR: #3548
  • feat(gemma4): add E-series tensor parallelism (3512) by @svcnvidia-nemo-ci :: PR: #3552
  • fix(retrieval): support canonical Sentence Transformers metadata (3546) by @svcnvidia-nemo-ci :: PR: #3551
  • ci: route DGX Spark recipes to GB10 (3538) by @svcnvidia-nemo-ci :: PR: #3556
  • fix(docker): build bitsandbytes for SM121 (3553) by @svcnvidia-nemo-ci :: PR: #3557
  • cp: backport packed THD VLM context parallelism to r0.6.0 by @qiaochuz-nv :: PR: #3558
  • fix(fsdp): uniform reduce dtype and EP-local expert gradients (3540) by @svcnvidia-nemo-ci :: PR: #3561
  • fix(deps): resolve msgpack, wandb, and mistune container CVEs (3563) by @svcnvidia-nemo-ci :: PR: #3566
  • fix(deps): resolve 26.08 rc9 container CVEs (3607) by @thomasdhc :: PR: #3611
  • fix(deps): resolve 26.08 rc10 container CVEs (3647) by @thomasdhc :: PR: #3658
  • fix(peft): support mixed-dtype memory-efficient LoRA backward (3675) by @svcnvidia-nemo-ci :: PR: #3682
  • beep boop 🤖: Bumping NeMo-Automodel to v0.5.1 by @nemo-automation-bot[bot] :: PR: #3687
  • fix(release): set version to 0.6.0 by @thomasdhc :: PR: #3692
  • ci: update package version to 0.5.0 (#2472) by @thomasdhc
  • feat: make mesh accept meshcontext (#2266) by @adil-a
  • docs(contributing): fix perk typo (#2471) by @grgkovac
  • feat(vlm): enable Qwen3.5 MoE VLM CP (#2432) by @HuiyingLi
  • ci: ask claude-review to flag stale ModelCapabilities (#2480) by @athitten
  • feat(speculative): add activation checkpointing for P-EAGLE draft layers (#2458) by @khazic
  • fix(speculative): release the EAGLE-3 target/W&B on any training exit (#2460) by @khazic
  • feat(model): flux2 (#2145) by @linnanwang
  • refactor(speculative): share grad-accum + LR utils across EAGLE/DFlash recipes (#2463) by @khazic
  • fix(speculative): require an explicit DFlash mask_token_id (#2461) by @khazic
  • fix(speculative): re-sync the target vocab mapping on EAGLE-3 resume (#2467) by @khazic
  • fix(speculative): drop position 0 from P-EAGLE depth-1 candidate pool (#2464) by @khazic
  • fix(speculative): finalize the last async checkpoint on training exit (#2469) by @khazic
  • fix(speculative): make regenerate clobber guard see existing shards (#2477) by @khazic
  • fix(speculative): validate config on precompute_eagle3 --resume (#2478) by @khazic
  • fix(speculative): keep DFlash DDP ranks in lockstep on no-valid-anchor skips (#2468) by @khazic
  • fix(speculative): serialize remote EAGLE-3 /generate end to end (#2479) by @khazic
  • feat(speculative): add EAGLE feature-noise augmentation to EAGLE-1/2 training (#2470) by @khazic
  • fix(speculative): gate unsupported P-EAGLE backend / cache / packing combos (#2466) by @khazic
  • fix(dflash): skip short validation micro-batches in eval loop (#2476) by @khazic
  • fix(speculative): retry the remote EAGLE-3 /generate wire path (#2459) by @khazic
  • docs(dllm): add DiffusionGemma 26B-A4B SFT/LoRA guide (#2504) by @zyzhou5
  • feat(dllm): add DiffusionGemma 26B-A4B block-diffusion SFT (full + LoRA) (#2506) by @zyzhou5
  • feat: add MSC cloud storage support for dcp checkpoints (#1709) by @edjson
  • feat(checkpoint): export torch_save checkpoints to hf (#2487) by @khazic
  • test(checkpoint): fix TestFormatLoad directory-based reader selection (#2520) by @khazic
  • fix(gemma4): cast dense params without casting buffers (#2359) by @DOGEUNNKIM
  • fix(config): glm4.7 yaml (#2527) by @akoumpa
  • docs(fern): relocate Fern under docs/ and remove legacy Sphinx tree (#2391) by @lbliii
  • fix: unwrap ModelOutput to extract logits (#2523) by @akoumpa
  • feat(speculative): add tqdm training progress bar to EAGLE and DFlash recipes (#2522) by @khazic
  • feat(speculative): add context parallelism for the EAGLE-3 target model (#2465) by @khazic
  • fix(speculative): compile variable-shape flex callables with dynamic shapes (#2532) by @khazic
  • fix(checkpoint): map BOOL in the safetensors backport DTYPE_MAP (#2533) by @khazic
  • docs: add MiniMax M3 VL user guide (#2536) by @athitten
  • feat: add MiniMax M3 VL (#2538) by @athitten
  • fix(docs): restore old .md (#2540) by @akoumpa
  • docs(fern): point restored guide pages to .md in nightly nav (#2541) by @lbliii
  • docs(dllm): use absolute image URLs in the DiffusionGemma guide (#2543) by @zyzhou5
  • fix(fern): generate autodoc library reference before publish (#2537) by @lbliii
  • fix(docs): restore previous md (#2544) by @akoumpa
  • feat(examples): add Nemotron-3-Ultra-550B benchmark and full-SFT recipes (#2539) by @adil-a
  • feat: Add MagiAttention (FFA / context-parallel) attention backend (#2384) by @HuiyingLi
  • fix(qwen3_5): make dense VLM pipeline-parallel safe (#2524) by @HuiyingLi
  • feat(vlm): enable Qwen3.5 dense VLM CP (#2505) by @HuiyingLi
  • fix(models): keep RoPE frequency buffers fp32 under bf16 model cast (#2549) by @akoumpa
  • ci: schedule ep-parallel finetune recipes at documented node counts (#2546) by @akoumpa
  • feat: add gemma4 moe ring CP (#1914) by @khazic
  • ci(fern): run docs check on every /ok to test (drop paths filter) (#2547) by @akoumpa
  • feat(qwen3_5): port dense Qwen3.5 to a native custom-model implementation (#2557) by @HuiyingLi
  • ci: flag RoPE/precision-buffer dtype hazards in automated PR review (#2552) by @akoumpa
  • fix(diffusion): resolve flux nightly CI failures (#2529) by @pthombre
  • docs: announce Gemma diffusion support (#2568) by @pthombre
  • docs: Update diffusiongemma.mdx (#2542) by @zyzhou5
  • feat(moe): MTP FLOPs accounting, inline shared experts, checkpoint warning fix (#2486) by @adil-a
  • fix(checkpoint): preserve tied lm_head on resume (#2511) by @yuhezhang-ai
  • test: fix all 5 vllm_deploy tests (token drift, nemotron OOM + mamba merge) (#2559) by @adil-a
  • fix(ci): bump ling_1t_lora_pp local_batch_size to satisfy PP assert (#2575) by @akoumpa
  • fix(ci): set node counts for multi-node VLM finetune recipes (#2574) by @akoumpa
  • fix(recipe): reshard MoE experts after forward in nemotron_nano_v3_cp_test (#2577) by @akoumpa
  • ci: use digits for spark recipes (#2581) by @akoumpa
  • ci: Enable activation checkpointing for gemma_2_9b_it_squad (AM-464) (#2585) by @akoumpa
  • fix(test): load checkpoint-robustness HF reference via device_map (#2582) by @akoumpa
  • fix(distributed): register Falcon-H1 TP plan to fix 34B PEFT OOM (#2589) by @akoumpa
  • feat(mtp): enable MTP to accept pre-fused input embeddings for multimodal models (#2510) by @Slyne
  • fix(peft): LoRA MLP QLoRA/PP/gemma3n fixes (AM-435, AM-447, AM-453) (#2584) by @akoumpa
  • fix(gemma4): FSDP2-safe kv-sharing + skip frozen audio tower on grad-accum (#2566) by @athitten
  • fix(vlm): enable activation checkpointing for 35B Qwen3.5/3.6 VLM recipes (#2600) by @akoumpa
  • fix(vlm): use FusedLinearCrossEntropy for qwen3_5_9b to avoid logits OOM (#2603) by @akoumpa
  • perf(distributed): add retrieval tuning knobs (#2452) by @yuhezhang-ai
  • fix(moe): preserve fp32 A_log in Qwen3.5-MoE and Qwen3-Next GatedDeltaNet (#2484) by @yuhezhang-ai
  • fix(moe): weight GroupedExpertsTE down-projection bias by routing probability (#2591) by @akoumpa
  • fix: use TE attention for gpt_oss packed-sequence recipe (AM-438) (#2587) by @akoumpa
  • fix(oom): use FusedLinearCrossEntropy in qwen3 tulu3 configs to avoid OOM (#2609) by @akoumpa
  • fix(bagel): distributed setup init (#2608) by @zyzhou5
  • feat(deepseek-v4): support context parallel training (#2590) by @HuiyingLi
  • fix(qwen3_5_moe): convert MTP experts as grouped tensors (AM-442) (#2595) by @HuiyingLi
  • fix(transformers): keep gemma3n KV sharing working under FSDP2 (AM-454) (#2594) by @HuiyingLi
  • feat(vlm): add Gemma 4 31B joint drafter example config (#2620) by @khazic
  • feat(datasets): add cp=1 shared-prefix prefix-tree attention for rollouts (#2564) by @khazic
  • feat(gemma4): context parallelism for dense 31B (#2592) by @HuiyingLi
  • fix(qwen3_moe): keep native forward under PP so CP+THD works (#2625) by @akoumpa
  • fix(docker): build DeepEP against the NVSHMEM wheel matching the apt runtime (#2614) by @akoumpa
  • fix(config): require pp_size and distributed.pipeline to agree (#2616) by @akoumpa
  • feat(gemma4): context parallelism for dense E2B/E4B (#2621) by @HuiyingLi
  • fix(moe): default ignore_router_for_ac=True for activation checkpointing (#2635) by @akoumpa
  • ci: raise ci.time for slow finetune recipes hitting 10-min default (#2637) by @akoumpa
  • ci: cap MAX_STEPS to 10 for slow vlm_finetune recipes (#2639) by @akoumpa
  • fix(examples): enable activation checkpointing for phi_4_squad (#2634) by @akoumpa
  • fix(parallelizer): resolve NemotronH decoder blocks for Nemotron-V3 (#2638) by @akoumpa
  • refactor(moe): remove enable_deepep, switch failing ep recipes to hybridep (#2630) by @akoumpa
  • test: run slow CP unit-test files GPU-only (gemma4, deepseek_v4) (#2647) by @akoumpa
  • fix(datasets): decode MedPix images on demand instead of up front (#2645) by @akoumpa
  • feat(config): add wandb.enable flag; disable W&B in example configs (#2643) by @akoumpa
  • fix(datasets): nest processing kwargs under processor_kwargs in default_collate_fn (#2649) by @akoumpa
  • fix(devstral2,ministral3): load FP8 checkpoints via custom mistral3_vlm path (drop HF FineGrainedFP8) (#2654) by @akoumpa
  • fix(merge_lora): save tokenizer faithfully via NeMoAutoTokenizer (#2653) by @akoumpa
  • fix(distributed): make fully_shard_by_dtype produce storage-uniform FSDP2 groups (#2655) by @akoumpa
  • fix(models): default yarn original_max_position_embeddings in Step3p7Config (#2652) by @akoumpa
  • fix(model): apply dict config overrides with custom dispatch (#2657) by @akoumpa
  • fix(recipe): disable fused RoPE for MLA packed-sequence MoE recipes (#2675) by @akoumpa
  • fix(training): flag final step when epochs exhaust before max_steps (#2672) by @akoumpa
  • fix(docker): bump DeepEP to 42144303 to pad HybridEP token capacity (#2678) by @akoumpa
  • fix(llama3_3): reduce pp_microbatch_size to fit 70b squad on h100 (#2673) by @akoumpa
  • fix(checkpoint): fp8 MoE expert weights silently not loaded at ep_shard=1 (random init → garbage loss) (#2682) by @akoumpa
  • fix(qwen3_moe): disable fused RoPE for the packed-sequence LoRA recipe (#2687) by @akoumpa
  • feat(glm_moe_dsa): GLM-5.2 IndexShare DSA support (#2633) by @HuiyingLi
  • fix(vlm): bump mistral3p5_128b_medpix max_length 1024->2048 (#2689) by @akoumpa
  • fix(moe): free DeepEP buffer before process-group destroy to avoid hang (#2686) by @akoumpa
  • test(models): speed up qwen3.5 moe/vl-moe from_pretrained unit tests (#2698) by @akoumpa
  • perf(checkpoint): mmap HF DCP read_data to avoid host-RAM OOM on large loads (#2690) by @akoumpa
  • fix(mistral3): remap FP8 VLM checkpoint prefixes (#2692) by @akoumpa
  • fix(loss): reuse LM head gather for MTP loss (#2694) by @akoumpa
  • fix(gemma4_moe): re-tie lm_head to active embed_tokens on MoE path (#2601) by @Achyuthan-S
  • ci: fix 26.06 release cves (#2705) by @thomasdhc
  • fix(moe): handle non-EP expert weight DTensors (#2697) by @akoumpa
  • ci: add cluster_tag to gb200 benchmarks (#2714) by @thomasdhc
  • build: install tilelang + tile_kernels for DeepSeek-V4 recipes (#2683) by @akoumpa
  • fix(checkpoint): super-49B consolidated reload and vllm_deploy (#2626) by @adil-a
  • fix(models): use bool sparse masks for sdpa (#2624) by @yuhezhang-ai
  • fix(distributed): register DeciLM Nemotron TP plan (#2703) by @akoumpa
  • ci: add time budgets for 12 new timeout failures (#2707) by @thomasdhc
  • fix(qwen3_5): handle packed MTP attention (#2727) by @akoumpa
  • docs: update container references to 26.06 and fix mount instructions (#2716) by @adil-a
  • perf(DSV4): use generic checkpoint wrapper for activation checkpointing (#2704) by @HuiyingLi
  • feat(glm_moe_dsa): add TileLang DSA kernels (#2691) by @HuiyingLi
  • ci: add gb200 cluster specification for nemotron_ultra recipe (#2733) by @thomasdhc
  • fix(ci): pin qwen3_moe_30b mxfp8 finetune to gb200 (#2735) by @thomasdhc
  • ci: address diffusers cve (#2706) by @thomasdhc
  • feat(glm_moe_dsa): add GLM5.2 context parallel support (#2695) by @HuiyingLi
  • ci: address thrift cve bump to 0.23.0 (#2736) by @thomasdhc
  • fix(deepseek-v4): restore batch axis for packed-sequence (THD) forward (#2651) by @akoumpa
  • fix(deepseek-v4): avoid bf16 -inf overflow in additive attention mask (#2658) by @akoumpa
  • build: install TileKernels for DeepSeek V4 (#2740) by @akoumpa
  • fix(ci): reduce mixtral release smoke batch (#2728) by @akoumpa
  • ci: Add gemma4 e4b to nightly test (#2749) by @athitten
  • fix(models): audit fp32 protected tensors (#2598) by @yuhezhang-ai
  • fix(fsdp2): guard uninitialized accumulated grads (#2744) by @akoumpa
  • fix(qwen3_moe): step-0 NaN in MXFP8 packed finetune — expert unload + fused RoPE (#2722) by @hemildesai
  • fix(diffusion): reuse warm HF cache instead of re-downloading models (#2747) by @pthombre
  • fix(diffusion): raise qwen-image dist timeout for checkpoint consolidation (#2748) by @pthombre
  • ci: bump benchmark glm_4.7_flash_te_deepep time (#2757) by @thomasdhc
  • fix(mistral3): preserve medium VLM checkpoint layout (#2758) by @akoumpa
  • ci: avoid remote config load for Nemotron Nano test (#2764) by @akoumpa
  • fix(sdpa): apply resolved backend constraints to custom models (#2761) by @akoumpa
  • fix(wandb): log different val datasets separately in wandb (#2526) by @grgkovac
  • docs(fern): document DistributedSetup Python API for NeMoAutoModel loaders (#2766) by @akoumpa
  • fix(distributed): use flattened CP FSDP mesh (#2768) by @HuiyingLi
  • fix: Remove dali from container (#2770) by @chtruong814
  • ci: run dsv32_lora, kimi_k2 and qwen3_moe_235b deepep benchmarks online (#2773) by @thomasdhc
  • fix: skip fused LoRA MLP install for meta weights (#2775) by @akoumpa
  • fix(checkpoint): resolve tie_word_embeddings top-level-first to match HF tying (#2732) by @Achyuthan-S
  • fix(benchmark): skip unsupported MTP flops (#2767) by @akoumpa
  • fix: Qwen3.5 MedPix EP32 NCCL timeout (#2777) by @akoumpa
  • fix(ci): address go-git/go-billy and rustls-webpki CVEs (#2780) by @thomasdhc
  • fix(ci): stabilize diffusion finetune smoke tests (#2788) by @pthombre
  • build(deps): move ffmpeg/opencv deps to opt-in media extra (#2743) by @thomasdhc
  • fix(vlm): keep Qwen3.5 media tokens aligned (#2772) by @yuhezhang-ai
  • fix(ci): HybridEP for multi-node MoE benchmarks + LoRA OOM fixes (#2789) by @hemildesai
  • fix: qwen3.5 and 3.6 mtp expert checkpoint layout (#2778) by @HuiyingLi
  • feat: CP support for MiniMax M3 (#2551) by @athitten
  • docs(speculative): fix EAGLE drafter layer count and EAGLE-1 loss (#2796) by @khazic
  • fix(ci): drop base-image uv/wandb copies flagged for CVEs (#2800) by @thomasdhc
  • chore(skills): refresh automodel skill signatures (#2804) by @akoumpa
  • feat(distributed): enable selective checkpointing for DDP (#2786) by @yuhezhang-ai
  • docs: align nightly navigation routes (#2801) by @akoumpa
  • feat: vision biencoder finetuning + Nemotron VL 1B finetuning support (#2354) by @gabrielspmoreira
  • fix(docs): catch Fern MDX syntax errors (#2806) by @akoumpa
  • feat(moe): add Rollout Routing Replay (R3) for MoE RL training (#2797) by @khazic
  • fix(optim): align Dion mesh with FSDP sharding (#2808) by @akoumpa
  • fix(ci): stabilize failed benchmark recipes (#2817) by @akoumpa
  • docs: document opt-in media extras (vlm-media/diffusion-media) (#2799) by @thomasdhc
  • feat(speculative): add SGLang target backend for EAGLE-3 training (#2449) by @khazic
  • feat(speculative): add Domino online training path on top of DFlash (#2819) by @khazic
  • feat(speculative): add dspark draft model and training objective (#2810) by @kashif
  • docs: add GLM-5.2 and speculative decoding (DSpark) updates (#2828) by @khazic
  • feat(diffusion): support Hugging Face datasets (#2816) by @pthombre
  • feat(speculative): add vLLM target backend for EAGLE-3 training (#2798) by @khazic
  • fix(speculative): gather sharded target lm_head in EAGLE-1/2 token loss (#2823) by @khazic
  • fix(model): honor yarn/linear/dynamic RoPE in LlamaRotaryEmbedding (#2825) by @khazic
  • fix(speculative): keep DFlash padding blocks self-attending to avoid NaN (#2826) by @khazic
  • feat(speculative): support target tensor parallelism in EAGLE-3 colocated path (#2827) by @khazic
  • feat(speculative): support target tensor parallelism in EAGLE-1/2 (#2829) by @khazic
  • feat(speculative): support target tensor parallelism in DFlash (#2830) by @khazic
  • fix(speculative): validate DFlash mask_token_id on checkpoint resume (#2824) by @khazic
  • docs: enable Fern multi-source (#2845) by @lbliii
  • fix(distributed): propagate NCCL timeout to derived device meshes (#2846) by @HuiyingLi
  • feat(speculative): add serve_vllm for EAGLE-3 / P-EAGLE drafts (#2841) by @khazic
  • fix(speculative): keep DSpark draft RoPE inv_freq in fp32 under bf16 (#2859) by @khazic
  • feat(speculative): compress EAGLE-3 offline cache target_probs via top-k (#2847) by @khazic
  • perf: FFPA D=512 attention backend for Gemma4 (3× fwd / 6× bwd vs SDPA) (#2436) by @Butterfingrz
  • feat(speculative): add JetSpec causal parallel drafting training (#2867) by @khazic
  • feat(models): reject tie_word_embeddings=True on separate-head model families (#2805) by @Achyuthan-S
  • ci: Ensure diffusion-media extra installed for diffusion tests (#2858) by @chtruong814
  • feat(speculative): add DeepSeek-V3 (MLA) EAGLE-3 draft model (#2849) by @khazic
  • feat(speculative): add DeepSeek V4 DSpark drafter and V4-Flash training (#2866) by @khazic
  • chore(models): add 26.10 deprecation warnings for custom model classes (#2807) by @athitten
  • feat(speculative): add DSpark draft for MiniMax M3 VL (text + multimodal) (#2877) by @khazic
  • ci: use NVIDIA inference for Claude review (#2882) by @chtruong814
  • feat(speculative): add GLM-5.2 DSpark draft model and training (#2885) by @khazic
  • fix(checkpoint): write consolidated safetensors without append (#2627) by @huahuajhu
  • ci(automodel): set AM-576 release timeouts (#2871) by @yuhezhang-ai
  • fix(diffusion_gemma): make parity test work on transformers >= 5.11 (#2883) by @zyzhou5
  • fix(bagel): torch2.12 DTensor sharding-prop leak (#2888) by @zyzhou5
  • fix(speculative): plumb EAGLE-3.1 fc_norm/norm_output into the draft config (#2897) by @khazic
  • fix(speculative): make save_consolidated final export the last EAGLE checkpoint (#2898) by @khazic
  • fix(speculative): gather TP target outputs in the Domino trainer (#2899) by @khazic
  • fix(speculative): default DFlash/Domino loss_decay_gamma to the paper value (#2908) by @khazic
  • refactor(speculative): dispatch all DSpark drafts through the registry (#2909) by @khazic
  • feat(speculative): serve DFlash and JetSpec drafts on vLLM (#2913) by @khazic
  • docs(speculative): document DSpark, Domino, JetSpec, and the vLLM serve path (#2914) by @khazic
  • feat(speculative): add bench_vllm acceptance/speedup benchmark (#2912) by @khazic
  • fix(speculative): read DeepSeek EAGLE-3 draft rope_theta from rope_parameters (#2920) by @khazic
  • fix(speculative): send chat_template_kwargs as a top-level request field (#2921) by @khazic
  • ci(deps): add albumentations to vlm-media for nemotron-parse processor (#2933) by @thomasdhc
  • fix(glm): select HybridEP for GLM-5.2 recipes (#2934) by @HuiyingLi
  • fix(distributed): omit unset FSDP reshard argument (#2926) by @zhiqi-li
  • feat(speculative): add DSpark offline cache path (#2924) by @fkuner
  • chore: update readme for 26.06-26.08 (#2886) by @akoumpa
  • ci: validate ci section on newly added recipes (#2936) by @thomasdhc
  • fix(bagel): make non-backbone init topology independent (#2939) by @zyzhou5
  • feat(speculative): add fp8 draft training and LoRA draft adaptation (#2963) by @khazic
  • feat(speculative): report simulated accept length during EAGLE-3 training (#2959) by @khazic
  • fix(distributed): fix and scope vlm activation checkpointing (#2840) by @yuhezhang-ai
  • feat(models): complete tie_word_embeddings guards for #2512 families (#2896) by @Achyuthan-S
  • feat(speculative): add distributed DSpark offline precompute (#2953) by @khazic
  • ci: route gb200 L2_HF_DCP to the IMEX runner for DeepEP NVLink P2P (#2968) by @ko3n1g
  • docs: add 0.5.0 release notes (#2881) by @lbliii
  • fix(vlm): gemma4 FFPA mock + e4b cp16 CI recipe fixes (AM-626, AM-627) (#2949) by @athitten
  • build(deps): upgrade transformers to 5.12.1 (#2873) by @athitten
  • ci: update claude review guidelines (#2739) by @akoumpa
  • fix(moe): preserve MLP dispatch through checkpoint wrappers (#2955) by @akoumpa
  • feat(speculative): add multi-dataset acceptance-length benchmark sweep (#2966) by @khazic
  • fix(transformers): gate _tie_weights_nemo on tie_word_embeddings flag (#2942) by @yuhezhang-ai
  • fix(loss): avoid pkg_resources in linear CE (#2700) by @akoumpa
  • fix(deepseek_v4): support packed THD with context parallel (#2731) by @HuiyingLi
  • feat(speculative): FSDP2-shard dense DSpark target + add Gemma4-31B config (#2976) by @khazic
  • fix(attention): initialize FlexAttention masks (#2982) by @akoumpa
  • fix(deepseek-v4): initialize random training state (#2991) by @akoumpa
  • ci(review): strengthen model review invariants (#2994) by @akoumpa
  • ci: Bump claude_review to v1.8.2 (#3012) by @chtruong814
  • ci: AUT-779 repin _claude_review.yml to v1.8.3 (#3013) by @svcnemo-autobot
  • test(ci): accept SHA-pinned claude-review template refs in the policy test (#3022) by @khazic
  • feat(speculative): context parallelism for draft-model training (#2918) by @kashif
  • feat(pp): support FusedLinearCrossEntropy under pipeline parallelism (#2927) by @beccohov
  • feat(speculative): support sequence packing in the DeepSeek MLA EAGLE-3 draft (#3002) by @khazic
  • fix(datasets): reject non-prefix multiturn mask templates (#3024) by @khazic
  • feat(speculative): support sequence packing in the EAGLE-1/2 draft (#3003) by @khazic
  • feat(speculative): add variable-prefix DFlash and LK EAGLE-3 training losses (#3032) by @khazic
  • feat(dspark): log acceptance and confidence metrics during training (#2957) by @kashif
  • feat(speculative): support sequence packing in the DFlash draft (#3004) by @khazic
  • feat(speculative): add periodic real accept-length eval (#3037) by @khazic
  • feat(speculative): support sequence packing in the Domino and JetSpec drafts (#3023) by @khazic
  • feat(speculative): support sequence packing in the DSpark Qwen3 draft (#3005) by @khazic
  • fix(models): temporarily disable TE fused RoPE globally (#3028) by @HuiyingLi
  • fix(glm_moe_dsa): clamp DSA indexer top-k for short sequences (#3035) by @HuiyingLi
  • feat(eagle3): on-policy regeneration loop with step-cadence dataloader swap (#3042) by @khazic
  • fix(docs): update technical accuracy (#3031) by @akoumpa
  • fix(checkpoint): compare torch versions semantically (#3043) by @HuiyingLi
  • docs: add uv run prefix (#3057) by @akoumpa
  • fix(recipe): add generation-marked chat templates (#3051) by @khazic
  • ci: include cutlass deps (#3056) by @akoumpa
  • fix(dspark): omit unmeasured positional acceptance rates (#3050) by @khazic
  • test(transformers): add source-load parity coverage (#2960) by @yuhezhang-ai
  • fix(distributed): resolve VLM AC layer-group paths structurally, not by version (#3025) by @yuhezhang-ai
  • fix(vlm): load Tulu-3 directly from HF Hub instead of local meta JSON (#3001) by @athitten
  • fix(optim): drop zero-numel local DTensor shards before TE FusedAdam (#2997) by @yuhezhang-ai
  • fix(moe): checkpoint trainable vision towers on the expert-parallel path (#2993) by @yuhezhang-ai
  • feat(datasets): add optional parallel pre-tokenization before packing (#3045) by @khazic
  • fix(models): handle checkpoint-wrapped fp32 buffers (#3059) by @akoumpa
  • feat(optim): support per-parameter-group learning rate (#3046) by @khazic
  • fix(cp): preserve gradients when sharding VLM inputs (#2931) by @HuiyingLi
  • refactor(vlm): gate build_model via recipe-side model-target allowlist (#2357) by @athitten
  • fix(ci): use HybridEP for Step 3.5 benchmark (#3069) by @HuiyingLi
  • fix: Kimi K2 config loading without remote code (#3065) by @HuiyingLi
  • feat(speculative): add Gemma4 EAGLE-3 target support (#3071) by @khazic
  • fix(dllm): remove obsolete DiffusionGemma HF compatibility (#3067) by @zyzhou5
  • feat(speculative): add Gemma4-E4B EAGLE-3 example config (#3073) by @khazic
  • fix(test): use SDPA for Nemotron-H HF reload (#3060) by @yuhezhang-ai
  • feat(speculative): add Gemma4-31B EAGLE-3 example config (#3077) by @khazic
  • feat(speculative): add Gemma4-26B-A4B MoE EAGLE-3 example config (#3079) by @khazic
  • chore(ci): AUT-852 bump claude review template to v1.8.4 (#3087) by @svcnemo-autobot
  • feat: add Ministral3 embedding distillation recipe (#3058) by @vinay-raman
  • fix(perf): use GPT-OSS head dimension in FLOPs accounting (#3091) by @yaoyu-33
  • test(speculative): add EAGLE-3 fp8 draft convergence smoke for SM89+ (#3081) by @khazic
  • fix(speculative): keep DSpark target_layer_ids in [0, N-2] for SGLang (#3083) by @khazic
  • chore(skills): add Regent Open Plugin manifest (#3097) by @ko3n1g
  • chore(skills): remove Open Plugin manifest (superseded) (#3099) by @ko3n1g
  • test: clean up KD test process group (#3090) by @yuhezhang-ai
  • refactor(models): single tie_word_embeddings guard via TieSupport (#2998) by @Achyuthan-S
  • feat(speculative): add DFlash validation metrics (#3072) by @khazic
  • refactor(datasets): typed Config + build per dataset, config-driven dataloader (#2390) by @akoumpa
  • fix(bagel): enable periodic garbage collection (#3105) by @zyzhou5
  • fix(ci): pass PR number and commit to Codecov for fork PRs (#3101) by @svcnemo-autobot
  • fix(test): use eager attention for Nemotron-H HF loads (#3100) by @yuhezhang-ai
  • feat(kd): support separate student and teacher meshes (#2954) by @akoumpa
  • fix(moe): safe TP and EP/TP gradient correctness for custom MoE models (#2995) by @yuhezhang-ai
  • test(ci): remove deprecated 26.10 models from nightly and release CI (#2884) by @athitten
  • feat(training): add opt-in setup-time prewarms (#2992) by @yuhezhang-ai
  • fix: bound chunked CE memory and preserve packing masks (#2996) by @yuhezhang-ai
  • fix(ci): restore Ministral3 checkpoint robustness (#3111) by @yuhezhang-ai
  • ci: add HF hub cache preflight check (#3119) by @thomasdhc
  • fix(vlm): resolve get_rope_index from the base model for packed mRoPE (#3113) by @khazic
  • feat(models): add Inkling VLM MoE support (#3095) by @hemildesai
  • ci(dllm): add dLLM SFT nightly train-to-generate launcher and recipes (#2783) by @zyzhou5
  • feat(vlm): add THD packed-sequence support for Qwen3-VL-MoE (#3052) by @khazic
  • docs: Add Inkling README news (#3129) by @HuiyingLi
  • fix(checkpoint): gather PEFT adapter across PP stages (#3096) by @hyfine
  • build(deps): bump base container to 26.06-cuda13.3 (#2983) by @thomasdhc
  • fix(transformers): keep custom M3 config when transformers ships its own (#3134) by @HuiyingLi
  • ci: AUT-911 shard GPU unit tests into 5 pytest-shard chunks (#3138) by @svcnemo-autobot
  • fix(checkpoint): isolate consolidation timeout from NCCL (#3108) by @yuhezhang-ai
  • test(vlm): add checkpoint robustness coverage (#3112) by @yuhezhang-ai
  • fix(ci): preserve retrieval evaluation schedule (#3139) by @yuhezhang-ai
  • feat(bagel): add TE support (#2895) by @zyzhou5
  • fix(moe): preserve HSDP replica gradient synchronization (#3135) by @wangzhxg
  • docs: add Laguna SFT docs and recipe (#3146) by @HuiyingLi
  • ci: AUT-895 gate GB200 tests on DISABLE_GB200_TESTS variable (#3118) by @svcnemo-autobot
  • feat(dllm): add LLaDA2 generation support (#3092) by @zyzhou5
  • feat(models): add Laguna model implementation (#3148) by @HuiyingLi
  • fix: prevent damaged token embeddings from dominating grad clipping (#3136) by @akoumpa
  • fix(checkpoint): add bounded retention window (#2416) by @oliverholworthy
  • fix(kd): use fp32 master weight copy (#3019) by @akoumpa
  • fix(checkpoint): restore adapters through DDP wrappers (#3150) by @yuhezhang-ai
  • feat(distributed): block-diagonal varlen CP for packed sequences (#2989) by @yuhezhang-ai
  • fix(distributed): preserve canonical activation checkpoint keys (#3152) by @yuhezhang-ai
  • feat(examples): Gemma4-31B CoderForge data pipeline + CP SFT recipes (#3151) by @athitten
  • feat(attention): support flash_attention_3 and flash_attention_4 (#2929) by @HuiyingLi
  • feat(checkpoint): add DCP CPU offload option (#3130) by @khazic
  • test(packing): fix stale unsupported-backend test after fa3/fa4 support (#3199) by @khazic
  • test(ci): add DeepSeek V4 Flash pretrain coverage (#3128) by @akoumpa
  • refactor(distributed): unify CP input prep and dispatch across models (#2937) by @HuiyingLi
  • docs(fern): add legacy URL redirects (#3203) by @akoumpa
  • ci: reduce L0 GPU tests to two shards (#3201) by @akoumpa
  • feat: add checkpoint staging wait option (#3131) by @khazic
  • fix(datasets): keep system turns in the sharegpt conversation converter (#3144) by @khazic
  • fix(peft): honor memory-efficient LoRA opt-out (#3126) by @akoumpa
  • feat: enable text inclusion alongside images in retrieval training (#3189) by @rnyak
  • feat: Support preemption checkpointing (#3007) by @edjson
  • fix(ci): AUT-964 restore Codecov after skipped GB200 jobs (#3205) by @svcnemo-autobot
  • fix(ci): preserve HF meta init for device-mapped loads (#3188) by @yuhezhang-ai
  • feat(dllm): add lora to dllm (#3163) by @zyzhou5
  • fix(gemma4): run SDPA in fp32 to avoid #2208 NaN on Hopper (#3141) by @aminehd
  • docs(retrieval): complete fine-tuning guide integration (#2306) by @oliverholworthy
  • fix(glm): apply LoRA to TileLang MLA KV projection (#3176) by @shahafwa
  • fix: prefer Automodel config registry lookup (#3202) by @akoumpa
  • feat(distributed): shape TPLinear/LinearLoRA graphs for async-TP fusion (#2987) by @yuhezhang-ai
  • fix(distributed): Megatron-FSDP 0.5.0 compatibility (#2986) by @yuhezhang-ai
  • feat(checkpoint): preserve intrinsic fp32 during offline consolidation (#3193) by @yuhezhang-ai
  • feat(kernels): add QuACK backend (#3115) by @akoumpa
  • feat(dllm): add DiffusionGemma generation via the built-in HF sampler (#3161) by @zyzhou5
  • refactor(diffusion): migrate recipe onto typed RecipeConfig build() path (#3122) by @pthombre
  • fix(dist): keep profiler record-function ops out of SAC replay accounting (#3133) by @HuiyingLi
  • fix(vlm): size qwen3.6 medpix CI configs (#3209) by @HuiyingLi
  • feat(glm_moe_dsa): expose update_moe_gate_bias on GLM MoE DSA models (#3207) by @jQizhang
  • fix(ci): shard Nemotron Super vLLM deploy across 8 GPUs (#3061) by @yuhezhang-ai
  • fix(glm): size GLM5.2 release CI jobs (#3210) by @HuiyingLi
  • fix(distributed): activation-checkpoint Qwen3-Next linear_attn layers (#3192) by @amolkhanna
  • fix(distributed): support packed CP for Llama, Qwen2, and Qwen3 (#2999) by @akoumpa
  • fix(dllm): recipe defaults, corruption seeding, sampler and docs fixes (#3162) by @zyzhou5
  • perf(bagel): grouped MoT routing + fused SwiGLU/RoPE (#3214) by @zyzhou5
  • feat(retrieval): add normalized Arrow dataset tooling (#2596) by @yuhezhang-ai
  • feat(diffusion): context parallelism via diffusers ContextParallelConfig (#3157) by @pthombre
  • docs: explain embedding row repair (#3175) by @akoumpa
  • fix(transformers): custom configs override builtin ones by default (#3221) by @HuiyingLi
  • build(deps): bump ffpa-attn to 0.2.2 (#3225) by @akoumpa
  • fix(distributed): checkpoint linear attention with compile (#3213) by @yuhezhang-ai
  • fix(glm): own the GLM MoE DSA config so qk_rope_head_dim survives (#3222) by @HuiyingLi
  • feat(speculative): add the multimodal speculative decoding (MSD) core (#3166) by @khazic
  • fix(speculative): average DFlash train metrics over the log window (#3234) by @khazic
  • fix(speculative): honor --dflash-causal on the vLLM remapping export (#3235) by @khazic
  • fix(speculative): align the decode-eval cadence on resume (#3237) by @khazic
  • fix(speculative): validate mask_token_id when resuming DSpark (#3238) by @khazic
  • build: cache CUDA extension source builds (#3204) by @akoumpa
  • docs(fern): cut v0.5.0 version train (#2970) by @lbliii
  • fix(checkpoint): infer Qwen MTP layout from checkpoint keys (#3229) by @HuiyingLi
  • docs: fix remaining broken link sources (#3249) by @akoumpa
  • fix(distributed): control frozen multimodal FSDP sharding (#2763) by @yuhezhang-ai
  • fix(ci): preserve main wheelhouse cache builds (#3253) by @akoumpa
  • feat(diffusion): add Qwen-Image-Edit-2511 training support (#3217) by @pthombre
  • fix(ci): omit token type IDs from vLLM parity prompts (#3198) by @yuhezhang-ai
  • fix(deepseek-v4): accept cu_seqlens for THD packing (#3231) by @HuiyingLi
  • fix(deepseek-v4): preserve fp32 model dtype (#3227) by @HuiyingLi
  • fix(thd): derive packed padding mask from the pack layout, not token ids (#3223) by @HuiyingLi
  • fix(nemotron-v3): recompute deterministic LoRA router (#3258) by @HuiyingLi
  • fix(moe): preserve top-k routing under full AC (#3140) by @yuhezhang-ai
  • fix(gpt-oss): route packed attention through THD (#3226) by @akoumpa
  • perf(benchmark): tune Qwen3-VL LoRA pipeline batches (#3230) by @akoumpa
  • fix(moe): preserve gate load across activation recompute (#3247) by @akoumpa
  • fix(ci): narrow CUDA wheelhouse cache keys (#3269) by @akoumpa
  • feat(distributed): frame-level context-parallel vision-tower sharding (#2990) by @yuhezhang-ai
  • fix(vlm): preserve mRoPE axes in PP chunking (#3208) by @HuiyingLi
  • feat(diffusion): add LTX-2.3 video+audio finetuning (training + prepr… (#3165) by @linnanwang
  • fix: harden checkpoint torch loads (#3240) by @HuiyingLi
  • build: add MagiAttention optional dependency (#3070) by @HuiyingLi
  • fix(checkpoint): harden PP safetensors consolidation (#3248) by @yuhezhang-ai
  • fix(checkpoint): save all PEFT EP optimizer parts (#3250) by @yuhezhang-ai
  • fix(ci): install CUDA torchvision with CUDA torch (#3279) by @akoumpa
  • fix(docker): cap uv install concurrency to avoid ARM FD exhaustion (#3275) by @thomasdhc
  • feat: support pre-extracted frame sequence of video sample (#3211) by @GITsologun
  • fix(inkling): support HF 5.14 attention fields and embed norm FSDP (#3281) by @HuiyingLi
  • fix(KD): Fix variable-length teacher PP batches on separate meshes (#3280) by @HuiyingLi
  • feat: add Kimi K3 model support (#3259) by @HuiyingLi
  • docs(models): add Kimi K3 coverage (#3283) by @HuiyingLi
  • test(vlm): stabilize Qwen3.5 checkpoint robustness (#3233) by @HuiyingLi
  • feat(vlm): integrate packed CP and vision sharding for Qwen3.5-MoE (#3186) by @yuhezhang-ai
  • feat(checkpoint): defer distributed async consolidation (#3125) by @khazic
  • test(retrieval): add Nemotron VL checkpoint coverage (#3276) by @yuhezhang-ai
  • chore(ci): AUT-1135 pin GitHub Actions to commit SHAs (#3277) by @svcnemo-autobot
  • feat: add ViSpec VLM draft training (#3173) by @khazic
  • fix(docs): hide Kimi tokenizer regex from autodoc (3288) (#3304) by @svcnvidia-nemo-ci
  • docs(distributed): review frozen multimodal FSDP guidance (3272) (#3306) by @svcnvidia-nemo-ci
  • fix(vlm): fused linear CE in gemma4 31B FFPA 8k recipe (AMINT-203) (3273) (#3307) by @svcnvidia-nemo-ci
  • fix(pp): use static metadata with PyTorch 2.13 (3290) (#3305) by @svcnvidia-nemo-ci
  • ci: add Tulu-3 convergence + eval flow (3171) (#3312) by @svcnvidia-nemo-ci
  • fix(distributed): preserve ERNIE router and dense Qwen3.5 SSM precision (3255) (#3310) by @svcnvidia-nemo-ci
  • perf(moe): add scoped partial CUDA graphs (2917) (#3314) by @svcnvidia-nemo-ci
  • fix(docs): avoid literal ampersand in Kimi autodoc (3317) (#3320) by @svcnvidia-nemo-ci
  • fix(config): repair GLM-5.2 LoRA recipe (3313) (#3323) by @svcnvidia-nemo-ci
  • fix(kd): preserve tensor-valued hidden states (3324) (#3332) by @svcnvidia-nemo-ci
  • fix(kd): use mesh-safe gradient clipping (3302) (#3334) by @svcnvidia-nemo-ci
  • refactor(docker): build torchao & FlashAttention as isolated wheel stages (3329) (#3343) by @svcnvidia-nemo-ci
  • fix(distributed): restore Nemotron Flash TP2 training (3345) (#3349) by @svcnvidia-nemo-ci
  • fix(nemotron-parse): sync RADIO preprocessing (3331) (#3341) by @svcnvidia-nemo-ci
  • fix(distributed): reuse default group for world-sized meshes (3319) (#3347) by @svcnvidia-nemo-ci
  • fix(training): prewarm Mamba SSD autotune kernels (3296) (#3338) by @svcnvidia-nemo-ci
  • fix(docs): make Fern autodoc metadata MDX-safe (3355) (#3357) by @akoumpa
  • fix(pp): preserve VLM media cursor with static metadata (3344) (#3361) by @svcnvidia-nemo-ci
  • fix(deps): resolve 26.08 rc2 container CVEs (3346) (#3363) by @svcnvidia-nemo-ci
  • fix(pp): route Qwen3.5 MoE pre-embedded inputs (3294) (#3364) by @svcnvidia-nemo-ci
  • fix(kimi_k25_vl): don't int4 quantize LoRA adapter keys on save (3295) (#3362) by @svcnvidia-nemo-ci
  • fix(diffusion): use spawn start method for GPU preprocessing pools (3339) (#3366) by @svcnvidia-nemo-ci
  • fix(tokenizer): make special token insertion opt-in (3337) (#3375) by @svcnvidia-nemo-ci
  • fix(distributed): stabilize SAC replay with TE and FSDP (3330) (#3376) by @svcnvidia-nemo-ci
  • fix(ci): use a single container cache donor (3377) (#3381) by @svcnvidia-nemo-ci
  • fix(fsdp): resolve fp32 master-weight compute dtype per parameter (3328) (#3384) by @svcnvidia-nemo-ci
  • perf(checkpoint): reduce distributed save overhead (3369) (#3389) by @svcnvidia-nemo-ci
  • fix(ci): enable LTX-2.3 diffusion finetuning (3372) (#3391) by @svcnvidia-nemo-ci
  • ci: tune Nemotron single-GPU model load threads (3370) (#3395) by @svcnvidia-nemo-ci
  • fix(model): correct Mistral4 attention and distributed MoE routing (3348) (#3393) by @svcnvidia-nemo-ci
  • fix(checkpoint): survive an interrupted save instead of hanging all ranks (3261) (#3394) by @svcnvidia-nemo-ci
  • perf: remove Python overhead from model hot paths (3374) (#3401) by @svcnvidia-nemo-ci
  • fix(deps): resolve OSS CVE findings (3398) (#3412) by @svcnvidia-nemo-ci
  • fix(moe): support EP-free DTensor state-dict conversion (3397) (#3409) by @svcnvidia-nemo-ci
  • fix(checkpoint): harden PEFT PP checkpointing and Step-3.7 coverage (3316) (#3403) by @svcnvidia-nemo-ci
  • docs(training): mark Mamba prewarm sections for review (3340) (#3399) by @svcnvidia-nemo-ci
  • fix(retrieval): guard optional W&B imports (3380) (#3402) by @svcnvidia-nemo-ci
  • fix(gemma4): enable expandable_segments on the CP tulu3 recipes (3408) (#3410) by @svcnvidia-nemo-ci
  • feat: Add {% generation %} chat template for DiffusionGemma SFT/LoRA examples (3353) (#3417) by @svcnvidia-nemo-ci
  • fix(docs): AUT-1324 qualify LTX model coverage slug (3418) (#3419) by @svcnvidia-nemo-ci
  • fix(ci): run Kimi K3 HellaSwag on GB200 (3420) (#3421) by @svcnvidia-nemo-ci
  • fix(transformers): default missing THD capability (3406) (#3429) by @svcnvidia-nemo-ci
  • ci(convergence): (3416) (#3432) by @svcnvidia-nemo-ci
  • cp: backport MoE checkpoint reload parity fixes to r0.6.0 (#3438) by @yuhezhang-ai
  • fix(retrieval): use stock Ministral embedding backbone (3103) (#3433) by @svcnvidia-nemo-ci
  • ci(minimax): extend M2.7 LoRA timeout (3428) (#3430) by @svcnvidia-nemo-ci
  • fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (3411) (#3445) by @svcnvidia-nemo-ci
  • fix(kimi_k25_vl): keep PEFT expert LoRA keys under the base_model.model. prefix (3431) (#3451) by @svcnvidia-nemo-ci
  • ci(vlm): raise Slurm wall time for MiniMax-M3 and Gemma4 recipes (3448) (#3453) by @svcnvidia-nemo-ci
  • fix: Gemma4-31B CoderForge CP8/64K data processing and recipe (3446) (#3452) by @svcnvidia-nemo-ci
  • cp: backport PEFT v5 adapter output fixes to r0.6.0 (#3458) by @yuhezhang-ai
  • fix(retrieval): repair Ministral3 training recipe (3441) (#3454) by @svcnvidia-nemo-ci
  • test(checkpoint): redesign resume robustness around a shared trajectory (3427) (#3457) by @svcnvidia-nemo-ci
  • test(checkpoint): stabilize MoE checkpoint parity gates (3440) (#3460) by @svcnvidia-nemo-ci
  • fix(gemma4): exact router scalar and fp32 reference routing for MoE parity (3456) (#3466) by @svcnvidia-nemo-ci
  • cp: backport Sentence Transformers metadata export to r0.6.0 (#3464) by @yuhezhang-ai
  • fix: normalize MoE auxiliary loss during gradient accumulation (3359) (#3459) by @svcnvidia-nemo-ci
  • fix(bagel): type SFT backend before AutoModel init (3449) (#3469) by @svcnvidia-nemo-ci
  • fix(peft): enable v5 expert adapters for Nemotron and MiniMax (3439) (#3468) by @svcnvidia-nemo-ci
  • fix(checkpoint): save ranked RNG state per global rank (3437) (#3471) by @svcnvidia-nemo-ci
  • fix(peft): align frozen LoRA tensors with compute dtype (3470) (#3480) by @svcnvidia-nemo-ci
  • cp: backport VLM CP gradient and vision sharding fixes to r0.6.0 (#3479) by @yuhezhang-ai
  • fix(training): all-reduce grad-norm scalars on the mesh device (3461) (#3481) by @svcnvidia-nemo-ci
  • fix(recipe): shard GPT-OSS 120B across 64 experts (3483) (#3489) by @svcnvidia-nemo-ci
  • fix(ci): register flux2/wan2.2/qwen-image-edit diffusion recipes in CI (3485) (#3487) by @svcnvidia-nemo-ci
  • perf(benchmarks): recompute deterministic MoE routers under AC (3474) (#3490) by @svcnvidia-nemo-ci
  • chore: Bump gitpython to >= 3.1.59 (3482) (#3488) by @svcnvidia-nemo-ci
  • perf(recipes): run GPT-OSS 120B with EP64 and no activation checkpointing (#3492) by @akoumpa
  • cp: feat(models): make Inkling standalone (#3358) into r0.6.0 (#3498) by @HuiyingLi
  • fix(models): preserve lm-head dtype boundaries (3491) (#3502) by @svcnvidia-nemo-ci
  • fix(checkpoint): validate PEFT adapter-only state (3501) (#3503) by @svcnvidia-nemo-ci
  • ci(convergence): fix gemma4 eval setup and re-baseline Qwen3-MoE (3493) (#3505) by @svcnvidia-nemo-ci
  • fix(deepseek_v4): avoid TileLang boolx8 backward codegen (3467) (#3506) by @svcnvidia-nemo-ci
  • fix(checkpoint): preserve FSDP2 mixed precision during recompute (3513) (#3515) by @svcnvidia-nemo-ci
  • fix(deps): resolve Starlette and GitPython CVEs (3523) (#3524) by @svcnvidia-nemo-ci
  • fix(minimax): declare packed-sequence and CP attention ownership (#3548) by @athitten
  • feat(gemma4): add E-series tensor parallelism (3512) (#3552) by @svcnvidia-nemo-ci
  • fix(retrieval): support canonical Sentence Transformers metadata (3546) (#3551) by @svcnvidia-nemo-ci
  • ci: route DGX Spark recipes to GB10 (3538) (#3556) by @svcnvidia-nemo-ci
  • fix(docker): build bitsandbytes for SM121 (3553) (#3557) by @svcnvidia-nemo-ci
  • cp: backport packed THD VLM context parallelism to r0.6.0 (#3558) by @qiaochuz-nv
  • fix(fsdp): uniform reduce dtype and EP-local expert gradients (3540) (#3561) by @svcnvidia-nemo-ci
  • fix(deps): resolve msgpack, wandb, and mistune container CVEs (3563) (#3566) by @svcnvidia-nemo-ci
  • fix(deps): resolve 26.08 rc9 container CVEs (3607) (#3611) by @thomasdhc
  • fix(deps): resolve 26.08 rc10 container CVEs (3647) (#3658) by @thomasdhc
  • fix(peft): support mixed-dtype memory-efficient LoRA backward (3675) (#3682) by @svcnvidia-nemo-ci
  • beep boop 🤖: Bumping NeMo-Automodel to v0.5.1 [skip ci] by @github-actions[bot]
  • fix(release): set version to 0.6.0 (#3692) by @thomasdhc