Sync vs fully-async agentic RL on verl: multi-turn GRPO, long-tail rollout profiling, staleness ablations — quantifying when async pays off.
reinforcement-learning rollout slime post-training multi-turn megatron gspo llm-agent tool-calling verl dapo agentic-rl async-training
-
Updated
Aug 14, 2026 - Python