Repository navigation
OpenAI's Navier–Stokes research effort involved a group of roughly 10,000 concurrent agents and used approximately 130 billion output tokens during problem-solving. For us, it offered a concrete glimpse of how large agent workloads can become—and prompted us to think more ambitiously about the RL infrastructure that trains them.
Preparing for that scale means giving rollout orchestration room to grow across nodes, making data flow durable, and preserving useful work when training fails. It also means focusing our development effort so that algorithm exploration and training at scale share a foundation we can keep improving.
slime v0.4.0 is a first step in that direction. We are focusing on synchronous and fully asynchronous training, distributing rollout orchestration across the cluster, introducing persistent queues and tensor storage with straw, and enabling Megatron to recover independently of healthy SGLang serving resources. These changes lay the groundwork for higher concurrency and longer runs, with training correctness remaining central to slime's design.
One training lifecycle, two modes
We are focusing development on two training modes:
- Synchronous training for exploring algorithmic limits and validating correctness.
- Fully asynchronous training for keeping rollout and training resources busy at larger scale.
Both modes now use train.py, with asynchronous scheduling handled by the rollout implementation. We have removed the legacy one-step off-policy driver, train_async.py, giving both workflows a shared training lifecycle to maintain and improve. ([#2391](#2391))
Alongside this simplification, we are expanding the tools available for stable training. v0.4.0 adds score centering, using a REINFORCE objective with support for top-p replay and the built-in importance-sampling corrections. Enabled with --use-score-centering, it provides another tool for addressing training–inference mismatch as we scale rollout and training. ([#2405](#2405), [#2406](#2406))
A distributed, persistent rollout pipeline
Increasing rollout capacity also requires orchestration to grow with the cluster. With a new storing backend straw enabled, fully asynchronous rollout now runs one rollout worker per Ray node with CPU resources, each managing its own generation pool and asyncio event loop. This distributes orchestration across the CPU resources already available in the cluster and reduces the bottleneck of a single process. Training batches can collect completed groups without waiting for every worker. ([#2410](#2410))
This distributed execution is backed by [straw](https://github.com/zhuzilin/straw), a library providing durable queues and shared tensor storage on a filesystem. Rollout workers publish payloads to shared storage, and training ranks read their assigned data by reference. Bulk data between rollout and training moves off Ray’s object-store transport, while Ray continues to carry control messages and references.
The same data layer persists prompt tasks, partial rollouts, completed samples, and converted training batches. Coordinating this progress with model checkpoints establishes a consistent recovery boundary for asynchronous training and supports checkpoint rollback. Samples and tensors are packed into large append-only files to reduce small-file and shared-filesystem metadata overhead. ([#2410](#2410), [#2427](#2427))
Select straw with --rollout-data-transport straw. JuiceFS is the currently supported shared-filesystem target; other network filesystems require deployment-specific validation. See the [straw guide](https://github.com/THUDM/slime/blob/v0.4.0/docs/en/advanced/straw.md) for configuration and recovery details.
Recover training while preserving serving and rollout work
As RL jobs grow, failures and configuration mistakes become more expensive. v0.4.0 separates the lifetime of Megatron training from the SGLang serving cluster, allowing healthy engines and routers to remain running after a training failure.
You can adjust supported trainer parallelism or memory settings and resubmit training to the same live Ray cluster. With Straw or saved rollout debug data, the new training attempt restores the last committed checkpoint and replays retained batches, reusing accepted samples and conversion results.
This reduces the cost of training retries by preserving serving resources and completed rollout work. See the [recovery guide](https://github.com/THUDM/slime/blob/v0.4.0/docs/en/advanced/fault-tolerance.md) for the supported restart workflow. ([#2444](#2444))
Other Improvements
-
PipelineRL:
--flush-cache-intervalallows unfinished requests and their KV caches to survive selected weight updates, reducing repeated rollout work. Periodic full refreshes remain configurable. This mode requires separate training and rollout GPUs. ([#2437](#2437)) -
Updated training stack: The Docker and Conda environments now use Megatron Core 0.19.2, Transformer Engine 2.18.0, Flash Linear Attention 0.5.2, and NCCL 2.30.7. The Megatron patch is also smaller, with more integration logic maintained in slime. ([#2448](#2448))
-
Streaming and agent rollouts: Added support for cumulative and incremental generation streams, partial-prefix retention, and cancellation of individual requests. Agent timeout cleanup now targets the router’s SGLang workers. ([#2272](#2272), [#2340](#2340))
-
Accelerator backends: Added Biren SUPA and Ascend NPU integrations to slime’s accelerator abstraction, including runtime initialization, device mapping, and communication backend selection. ([#2393](#2393), [#2424](#2424))
-
Correctness fixes: Corrected linear-attention gradients under context parallelism, preserved reference log probabilities over the full vocabulary during actor top-p replay, and fixed checkpoint shard selection in disk-delta weight synchronization. ([#2443](#2443), [#2434](#2434), [#2442](#2442))
What's Changed
- feat: support streaming external rollouts by @Aphoh in [#2272](#2272)
- fix(agent): abort timed-out SGLang requests via router workers by @Arlo-mt in [#2340](#2340)
- sync from internal by @zhuzilin in [#2390](#2390)
- refactor: remove legacy async training entrypoint by @zhuzilin in [#2391](#2391)
- Remove r3 dtype bug by @zhuzilin in [#2392](#2392)
- feat: add backend-aware SUPA (Biren) support by @kakyo82 in [#2393](#2393)
- Don't use reloadable pg when not offloading train by @zhuzilin in [#2394](#2394)
- [slime] Support score centering by @zhuzilin in [#2405](#2405)
- Support score centering with top_p replay by @zhuzilin in [#2406](#2406)
- Remove obsolete rollout global dataset flag by @zhuzilin in [#2407](#2407)
- Support distributed fully async rollout and add straw to slime by @zhuzilin in [#2410](#2410)
- feat(accelerator): add Ascend NPU backend by @ke-working in [#2424](#2424)
- Remove buffer in QueueDataSource and some cleanup for readability by @zhuzilin in [#2426](#2426)
- Support straw checkpoint rollback and indexed rollout archives by @zhuzilin in [#2427](#2427)
- Remove obsolete Straw queue limits by @zhuzilin in [#2432](#2432)
- Support pipeline rl with --flush-cache-interval by @zhuzilin in [#2437](#2437)
- fix(ppo): avoid top-p replay for reference logprobs by @Beverly621 in [#2434](#2434)
- Support manual Megatron restarts with retained serving by @zhuzilin in [#2444](#2444)
- fix: linear-attention gradients are wrong when CP > 1 (Qwen3.5 / Qwen3-Next) by @sdc17 in [#2443](#2443)
- fix(weight-sync): honor checkpoint shard indexes in disk delta sync by @chaoworks in [#2442](#2442)
- [release] bump to v0.4.0 by @zhuzilin in [#2448](#2448)
New Contributors
- @Aphoh made their first contribution in [#2272](#2272)
- @Arlo-mt made their first contribution in [#2340](#2340)
- @kakyo82 made their first contribution in [#2393](#2393)
- @ke-working made their first contribution in [#2424](#2424)
- @Beverly621 made their first contribution in [#2434](#2434)
- @sdc17 made their first contribution in [#2443](#2443)
- @chaoworks made their first contribution in [#2442](#2442)
Full Changelog: [v0.3.2...v0.4.0](v0.3.2...v0.4.0)