-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Legato
Venue: RSS 2026 (Manipulation session) Β· Authors: Yufeng Liu, Hang Yu, Juntu Zhao, Bocheng Li, β¦ Dequan Wang, Yang Gao β SJTU Γ Spirit AI Γ Tsinghua Γ Tongji Γ USTC (work done during an internship at Spirit AI) Β· arXiv: 2602.12978 Β· project Category: Real-time execution for chunked flow policies Trend tag: RSS 2026 thread 7 β action representation & inference mechanics
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Figure 1 of the paper. Top: smoothness (NSPARC) vs completion time across five real tasks (bowl, drawer, pickplace, towel, pour) β the Legato points (blue) sit left of and below their RTC counterparts (grey) on both axes: shorter execution and smoother trajectories. Bottom: an execution trace on the pour task β RTC's y-position trace shows multimodal switching with poor overlap alignment between the executed and discarded chunk continuations (photo inset: visible "hesitation"), while Legato's trace is smooth with well-aligned chunk overlaps ("moving"). Hesitation-induced slowdowns are the concrete cost that native continuation removes.
Action chunking gives VLAs real-time throughput but discontinuities at chunk boundaries. Real-Time Chunking patches this externally (inference-time guidance) β which causes spurious multimodal switching and trajectories that are never intrinsically smooth.
Legato makes continuation native to training:
- Denoising initializes from a schedule-shaped mixture of known (already-committed) actions and noise, exposing the model to partial action information during training;
- The learned flow dynamics are reshaped so denoising stays consistent between training and inference under per-step guidance;
- Randomized schedule conditioning during training supports varying inference delays and yields controllable smoothness.
- Smoother trajectories, less spurious multimodal switching, less hesitation.
- Across five real-world manipulation tasks: β10% improvements over RTC in both trajectory smoothness and task completion time.
The clearest instance of RSS 2026's "internalize the inference patch" pattern: RTC (NeurIPS 2025, test-time) β training-time RTC (Ξ¨β adopts it after finding test-time guidance unstable) β Legato (fully native continuation with delay randomization). Together with AR-VLA's persistent-context expert (#85) and Action-to-Action flow (#209), chunk-boundary handling has become its own sub-field β directly relevant to every flow-matching VLA in Review-VLA-Attention's streaming discussion.
β RSS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)