Repository navigation
v0.1.1
TL;DR
More of the latest frontier models land. Qwen3.8-Flash-Next and GLM-5.3-Flash arrive with colocated RL, Kimi-K3 moves from a separate image into the main release, and DeepSeek-V4 and Nemotron-H gain a native Megatron path, packed sequences, and fixes.
Multi-LoRA with Tinker compatibility. A Tinker-compatible gateway lets isolated training jobs share one base model, each adapter with its own optimizer state and checkpoints.
Enhanced agentic training. Agentic training works with Anthropic-API harnesses such as Claude Code and runs Harbor tasks directly in E2B and Daytona sandboxes.
Optimizations across the stack. On memory, colocated training has a lower peak host memory, policy logits stay in BF16 without changing gradients, and optimizer state for both Adam and Muon can stream to NVMe. Beyond memory, fused NVFP4 fake-QAT kernels and leaner NVFP4 weight sync cut low-precision overhead, and checkpoint saves overlap the HF export with the native async write.
Highlights
472 PRs from 31 contributors.
Model Support
| Model | Training | PRs | Docs |
|---|---|---|---|
| Kimi-K3 | Full, LoRA | #1825 | link |
| Qwen3.8-Flash-Next | Full | #2777 | link |
| GLM-5.3-Flash | Full | #2786 | link |
Qwen3.8-Flash-Next. A 48-layer hybrid MoE with QSA sparse attention, hyper-connections, and a host-resident PLE n-gram embedding. Colocated on 8x4 GB300 with R3, weight-update equality passes and train_rollout_logprob_abs_diff stays at 0.0083 to 0.0131 over 5 steps as raw reward rises from 0.47 to 0.75 (#2777, sgl-project/sglang#40821, radixark/Megatron-LM#89).
GLM-5.3-Flash. 34 KDA linear-attention and 11 DSA sparse-attention layers over a 288-expert MoE. Colocated in BF16 on 8x4 GB300 with R3, weight sync checks equal and train_rollout_logprob_abs_diff stays at 0.0085 to 0.0098 over 3 steps as raw reward rises from 0.50 to 0.72; R3 needs --sglang-moe-runner-backend triton, which the launcher sets (#2786, sgl-project/sglang#40821).
Kimi-K3 in mainline. Full and LoRA RL now run on the stock radixark/miles image, with no separate branch. A 236-step full run on 16x4 GPUs lifts AIME from 0.467 to a 0.867 peak; on the four-layer prune, train/rollout KL with a trained adapter (0.0057) is no higher than the MXFP4-vs-BF16 floor without one (0.0063) (#1825, #3194).
DeepSeek-V4. --dsv4-impl megatron adds Megatron's native implementation (TP1) beside the default miles one, loading the same checkpoint; on the same rollout at TP1 on 8x H200 they reach train_rollout_logprob_abs_diff 0.0885 and 0.0837. DeepSeek-V4-Flash also gains THD packing with CP (--qkv-format thd, validated on 4x8 MI355X), and DeepSeek-V4-Flash-0731 loads through a bit-exact MXFP4 to FP8 converter (#2706, #2038, #2717).
Nemotron-H. Mamba convolution weights now load and sync (an HF to Megatron to HF audit of Nemotron-3-Nano-30B-A3B goes from 46 missing tensors to all 6,243 matching), and MTP-disabled runs no longer fail on the first forward (#3163, #3164).
Features
Sampling-support replay. For top-p/top-k rollouts, the trainer scores the actor on the exact token support SGLang sampled from instead of the full vocabulary; the Qwen3-0.6B FSDP E2E (2x H200, 5 GRPO steps) checks rollout/trainer log-prob equality on every rollout. It turns on when top-p < 1 or top-k > 0, needs a positive top-k, and for now excludes KL and on-policy distillation (#2200, #2595, #3354).
Multi-LoRA v2 with a Tinker-compatible gateway. Adapter slots with their own optimizer state, checkpoints, and losses serve isolated Tinker tenants on one base model; a saved slot restores into a fresh one with weights, FP32 masters, and Adam moments equal at zero tolerance. LoRA targets are chosen by HF module group with --target-modules attn,mlp,unembed (#2839, #2840, #2841, #2842, #2843, #2844, #2845, #2849, #3318).
Agentic training. The session server serves the Anthropic Messages API (/v1/messages), so harnesses like Claude Code train through TITO unchanged. Harbor trials run in-process on cloud sandboxes without an agent server, passing end to end with Harbor's oracle agent on E2B and Daytona. New recipes cover Terminus 2 compaction on all 89 Terminal-Bench 2 tasks and a NeMo Gym Workplace Assistant (#2358, #2806, #2836, #2741, #3658).
Launchers and eval. Launchers detect the node with --hardware auto and size GPUs per node from it, and --debug-train-only runs evaluate checkpoints through the same snapshot-eval dispatcher as RL (#2738, #3172).
AMD. ROCm 10 MI35X and ROCm 7.2.4 images, and a Qwen3-4B LoRA launcher that completed 100 GRPO steps on 8x MI355X (#2854, #3326, #3183).
Optimizations
Host memory for colocated frontier models. --colocate-memory-peak-device gpu (default cpu) keeps the engine weight mirror and the trainer backup from coexisting in host RAM, cutting the handoff peak from 324 to 257 GB (Qwen3.5-35B-A3B LoRA, 8x H200). NVMe offload now covers Muon optimizer state and Adam's main-parameter init: at full offload on 4x GB300 Muon's pinned-CPU path was killed at 908 of 919 GB, while the disk backend finished with grad_norm inside the no-offload band. Clearing TE's quantized-weight cache before offload (on by default) moves 45% less host memory per offload (4-layer DeepSeek-V4-Flash, MXFP8, 4x B200) (#2781, #3100, #2739, #2653, #2142).
NVFP4 RL. Fused CuTe DSL fake-QAT kernels remove the quantize/dequantize memory round trips and back a GLM-5.2 W4A16 recipe; NVFP4 now covers routed experts only, keeping shared experts at source precision end to end (#2864, #2855, #3217).
Policy logits in BF16. Log-prob and loss forwards no longer upcast the full logits to FP32, cutting forward plus backward memory from 768 to 304 MiB at [2048, 32768] with loss and gradients unchanged (max abs diff 0, 1x H200) (#2764, #2818).
Refactors
Rollout and trainer orchestration. Rollout splits into an InferenceController that manages the engines and a RolloutExecutor that generates; engines, routers, and session servers launch as command subprocesses under one RayWorkerManager, and trainers and the rollout process call SGLang over HTTP instead of through Ray actors. Training runs through a single cell-based TrainerController, which the new fault-tolerance machinery (still being merged, #1837) builds on; user-visible changes are under Breaking Changes (#1842, #1843, #1861, #1985, #2051, #2054, #2063, #2154, #2161, #2176).
Weight update on one driver. Every transfer protocol (broadcast, P2P, RDT, disk-delta, colocated) streams size-bounded buckets from a backend-neutral HfWeightIterator through one WeightUpdater that owns the pause, version, and resume frame. Distributed syncs now read the weights-backuper snapshot instead of live parameters, which differ under --rematerialize-param-from-master-weight, and LoRA on distributed engines no longer requires PP=1 (#2740, #2745, #2746, #2747, #2751, #2752, #2753, #2754, #2756).
Session request pipeline. Each request is resolved once, server constraints first and then model rules over the full request, so local rendering and the backend see the same arguments. Client fields that conflict with the server (e.g. input_ids) get HTTP 400, rejected or retried requests no longer move the served path or drop committed turns, and each turn records its resolved request for continuations (#3258, #3301, #3264, #3302, #3311, #3249).
Checkpoint writing. Single-LoRA saves and full-model HF exports share one SnapshotPublisher and directory writer, and readers that overlap a later save get distinct snapshot paths. Force-sync saves (final step, external triggers) now overlap the HF export with the native async write, as periodic saves already did (#3198, #3342).
Full release notes by category below; breaking changes are near the end.
RL Algorithms & Training
- [Sampling Replay] Bounded sampling-support replay: #2200, #2595, #3354
- fix: chunk true-on-policy logprob computation: #2825
- perf(megatron): keep policy logits in model precision: #2764
- fix(megatron): keep SFT logits in model precision: #2818
- [PPO] Support the critic under --megatron-to-hf-mode bridge: #2878
- Reject indep_dp when shared actor-critic PPO is enabled: #2891
- Decide the pipeline stage from the parallel state in compute_advantages_and_returns: #3125
- feat(megatron): reset optimizer state by optimizer class: #3091
- fix(megatron): respect --no-save-optim when saving checkpoints: #3083
- fix: resume from the checkpoint step in bridge mode: #2682
- Fix MTP gradient detachment in AutoBridge mode: #3211
- fix: align MTP loss masks with input tokens before context-parallel slicing: #3219
- Fix trainer-rollout diagnostics with rollout logprobs enabled: #3655
- [Feat] snapshot eval during --debug-train-only by reusing the eval dispatcher: #3172
- Skip the last rollout's engine handoff unless a final eval follows: #3182
- fix(train): hand off to the engine before an off-cadence final eval: #3615
Rollout & Session
- Session v2 request pipeline ([1/6] to [6/6]): #3258, #3301, #3264, #3302, #3311, #3249
- feat(session): serve Anthropic Messages through SessionCore: #2358
- Tolerate missing optional SGLang Anthropic helpers: #3114
- fix(inkling): parse raw completions in the session server: #2302
- feat(session): configure session server workers explicitly: #2774
- Let agents outside the cluster reach session servers on more than one host: #2851
- fix(session): use incremental R3 for all non-retract cases, and fix asserts: #2834
- fix(session): disable training replay outputs during evaluation: #3323
- fix(session): reject aborted backend generations: #3337
- perf(session): decode samples off the event loop: #3339
- feat(session): fill omitted sampling fields from defaults set at session creation: #3585
- fix(session): default picker trims only identical re-sends, incl. first turns: #3657
- feat(tito): support Qwen3.5 and Qwen3.6 templates: #2759
- feat(tito): support Qwen3.8 27B and Flash-Next templates: #2760
- feat: expand GLM 53 TITO model support: #3087
- fix(rollout): drop groups with missing rewards: #2814
- Skip non-numeric rewards in the episode average: #2743
- fix(rollout): isolate cancelled groups in fully async generation: #3319
- fix(async): publish weights before the next fully-async drain: #3343
- fix(args): treat the fully-async rollout path as the mode that selects it: #3587
- fix(data): support mixed text and multimodal prompts: #3600
Agentic & Environments
- rollout: shared miles-side layer for agent-function legs (session URL, sandbox credentials): #2805
- examples: in-process Harbor agent function: #2806
- examples: launcher for the in-process Harbor path: #2836
- feat(agentic): use Modal Sandbox v2 readiness: #3341
- feat(agentic): forward Harbor sandbox resource controls: #3345
- harbor daytona: reclaim orphaned sandboxes by default, and a storage override: #3215
- Add Terminus 2 compaction training example: #2741
- Add experimental Nemotron Workplace Assistant RL example: #3658
- feat(openenv): report episode end reason: #2677
- openenv e2b backend: declare the SDK as an extra, floor it at 2.12: #2809
- openenv: raise the tbench2_env contract floor to OpenEnv#1025: #2811
- openenv tb2 recipe: pin mcp below 2 in the server layer: #2863
- fix(verifiers): cap renderers before pool API removal: #3084
- fix(swe-agent): flush all session-server instances on abort: #2748
LoRA & Multi-LoRA
- [multi-lora v2 0/7] to [7/7] Multi-LoRA v2 and the Tinker gateway: #2839, #2840, #2841, #2842, #2843, #2844, #2845, #2849
- Unify LoRA target selection with HF model layouts: #3318
- fix: Multi-LoRA correctness fixes across models: #3584
- Fix Tinker gateway cleanup on launcher exit: #3656
- Native LoRA fixes: grad dtype order, gloo barriers, strict adapter load, trainer-owned adapter: #3194
- fix colocate lora duplicated cpu param backup by keeping TMS only: #3089
- Bound the trainer's GPU memory during the LoRA weight sync: #3147
- fix(lora): pass config to the bridge value head instead of sequence_parallel: #2869
- fix(megatron): respect --no-load-optim when resuming a LoRA adapter: #3245
Weight Update
- Publish the new weight version to the rollout executor: #2502
- Synchronize disk-delta baseline reloads across trainer ranks: #2500
- fix(disk-delta): reject noncanonical tensor layouts: #2692
- fix: use host weights when syncing an offloaded actor: #3128
- fix: guard connection allocations during offloaded weight sync: #3129
Low Precision
- Add fused NVFP4 fake-QAT QDQ kernels: #2864
- Restrict NVFP4 RL to routed experts: #2855
- Limit NVFP4 conversion to main language decoder experts: #3638
- Reduce NVFP4 weight synchronization overhead: #3217
Memory & Offload
- Add --colocate-memory-peak-device: overlap the trainer/rollout handoff on the GPU: #2781
- Support --colocate-memory-peak-device gpu under LoRA: #3100
- feat: dist_muon offloading in megatron: #2739
- Stream optimizer main initialization to NVMe: #2653
- Drop TE cached quantized weights before offloading the training actor: #2142
Model Support & Fixes
- DeepSeek-V4 support-v2: native megatron implementation as option: #2706
- [Feat] Enable THD packed-sequence training for DeepSeek-V4-Flash, with context parallelism: #2038
- add DeepSeek-V4-Flash-0731 support and mxfp4->fp8 converter: #2717
- Fix missing Nemotron-H Mamba convolution mappings: #3163
- Fix Nemotron-H forward compatibility when MTP is disabled: #3164
- fix: align MLA RoPE types with model configs: #2716
- fix(convert): restore auto pipeline parallelism when ETP is unset: #3653
Launchers & Configuration
- feat: resolve --hardware from the node in every launcher: #2738
- Refuse a static port that something else is already serving: #2885
- fix: support dataclass SGLang ServerArgs: #3581
- Dump debug train data once per DP/CP shard instead of on every rank: #3141
Dashboard & Metrics
- fix(metrics): propagate and correctly aggregate speculative decoding counters: #2598
- Log compaction-aware rollout metrics: #2710
- feat(metrics): log per-rollout response lengths: #2765
- fix(metrics): scope episode rewards by adapter: #2778
- Scan full responses for repetition metrics: #3115
- metrics(async): report selected policy provenance: #3334
- [Dashboard for AMD GPU]: add AMD SMI GPU telemetry: #2696
- feat(dashboard): render conversation images: #3239
- Dashboard: handle multiple TITO leaves per sample index: #2881
AMD / ROCm
- [AMD] docker: retire the rocm700 variants, add a rocm10-mi35x variant: #2854
- [AMD] replace the ROCm 7.2.0 image with ROCm 7.2.4: #3326
- [AMD][LoRA] Add a Qwen3-4B LoRA RL launcher: #3183
- [AMD] Point the ROCm image at the v0.5.16 wheels release with TE 2.17.0: #2660
Refactors
- Rollout and trainer orchestration rewrite (InferenceController, RayWorkerManager, cell-based TrainerController), the base for the new fault tolerance; 205 PRs tracked in #1837, including: #1842, #1843, #1861, #1985, #2051, #2054, #2063, #2154, #2161, #2176
- Worker RPC framework: #1942, #1943, #1944, #1945, #1946, #1947, #1948, #1949, #1950, #1951, #1952
- Weight update on one backend-neutral driver: #2740, #2745, #2746, #2747, #2751, #2752, #2753, #2754, #2756
- Unify HF snapshot and checkpoint writing: #3198
- perf(checkpoint): overlap force-sync HF export with async save: #3342
Docs Updates
- Qwen3.8-2.4T-A95B: fuller RL recipe on the Qwen3.8 page: #2650
- Mooncake rollout data transfer guide: #2535
- Sandbox providers: setup and new-provider pages, Harbor on Modal, and the in-process Harbor example: #3213, #3214, #2837
- AMD ROCm page and Hardware Platforms navigation: #3209, #3313
- Stale paths, flags, env vars, and metric names fixed across the docs: #2651
- README and docs landing refresh: #2606, #2655, #2657
Dependencies
- SGLang: v0.5.16 to v0.5.20, with
sglang-milesrebased onto each release (locked at880e3d24, re-pinned after the cut to pick up the KV offload synchronization fix): #2714, #3124, #3321 - Megatron-LM: rebased onto NVIDIA's latest dev as
miles-main-20260819, with the image trackingmiles-mainagain (locked atf148a32b): #2673, #2734 - Megatron-Bridge: pinned to
8cd3466d: #2872, #3336, #3584 - flash-linear-attention 0.5.2: #2790
- TileLang 0.1.14: #3309
- nvidia-cudnn-frontend 1.28.0: #3296
- torch_memory_saver
b5588e83: #2670, #3192
Docker
docker pull radixark/miles:v0.1.1 # CUDA 13 (amd64 + arm64)Breaking Changes & Upgrade Notes
- Versioned releases ship CUDA 13 only. Starting with v0.1.1 there is no
radixark/miles:v0.1.1-cu12; the release image is CUDA 13 on amd64 and arm64. CUDA 12 users stay on thedev-cu12development builds or the existingv0.1.0-cu12tag: #3703 - External router mode is removed. Miles always starts its own router per model, and startup asserts
--sglang-router-ipis unset. Drop--sglang-router-ipand--sglang-router-portfrom launchers that pointed at an external router; there is no replacement: #1996 - Session server ports and count are set separately.
--session-server-portnow takes one base port (it took one port or aSTART ENDrange), instance i listens on base + i, and the instance count comes from--session-server-workers, which defaults to 32 (v0.1.0 started one server when no port was given).--session-server-ipis now the bind address and defaults to the worker's placed address instead of--sglang-router-ip. Replace--session-server-port START ENDwith--session-server-port START --session-server-workers N: #2053, #2774 --control-server-portis renamed--api-server-port. Left unset it is 0 (disabled) unless--use-fault-toleranceis on, and the new--api-server-hostbinds to127.0.0.1by default; set0.0.0.0to accept remote controllers: #1971, #2897- Multi-LoRA v1 is removed. The dataset-driven v1 driver and its flags are gone (
--multi-lora-adapter,--multi-lora-api-port,--multi-lora-backend-path,--multi-lora-disable-service-mode,--multi-lora-http-server-path,--multi-lora-idle-poll-s,--multi-lora-max-adapter-global-batch-size,--multi-lora-max-coalesce-wait-s,--multi-lora-max-empty-wait-s). Multi-adapter training moves to the multi-LoRA v2 Tinker gateway (serve_tinker.py); there is no v1 fallback: #2839, #2845 - Overwriting a checkpoint deletes the old one first. A save to an existing checkpoint directory removes it before writing, so a failed overwrite leaves no usable checkpoint there. Save to a new path when the previous checkpoint must survive a failed save: #3198
All Contributors
We would like to thank all community contributors of Miles and SGLang for their invaluable contributions and continued support.
All Miles contributors to this release (in alphabetical order):
@alisonshao @Arist12 @BaiqingL @behroozazarkhalili @bojiang-li @fzyzcjy @guapisolo @HJSang @ilyasher-harmonic @JessicaJiang-123 @kaixih @Michaelsqj @michaelzhang-ai @MinhaoLi0318 @nanjiangwill @nblintao @QQontheMoon @Rockdu @Shi-Dong @sreerohi @Xinyu-Kang @XinyuJiangCMU @xiuhu17 @xyuzh @YolandaLyj @yueming-yuan @yushengsu-thu @Zhichenzzz @zianglih @zxpdemonio @zyzshishui
Full Changelog: v0.1.0...v0.1.1