Skip to content

RSS 2026 Legato

hwoo.han edited this page Jul 25, 2026 · 2 revisions

Legato β€” Learning Native Continuation for Action Chunking Flow Policies

Venue: RSS 2026 (Manipulation session) Β· Authors: Yufeng Liu, Hang Yu, Juntu Zhao, Bocheng Li, … Dequan Wang, Yang Gao β€” SJTU Γ— Spirit AI Γ— Tsinghua Γ— Tongji Γ— USTC (work done during an internship at Spirit AI) Β· arXiv: 2602.12978 Β· project Category: Real-time execution for chunked flow policies Trend tag: RSS 2026 thread 7 β€” action representation & inference mechanics

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

Legato vs RTC (Figure 1 of arXiv 2602.12978, Β© the authors)

Figure 1 of the paper. Top: smoothness (NSPARC) vs completion time across five real tasks (bowl, drawer, pickplace, towel, pour) β€” the Legato points (blue) sit left of and below their RTC counterparts (grey) on both axes: shorter execution and smoother trajectories. Bottom: an execution trace on the pour task β€” RTC's y-position trace shows multimodal switching with poor overlap alignment between the executed and discarded chunk continuations (photo inset: visible "hesitation"), while Legato's trace is smooth with well-aligned chunk overlaps ("moving"). Hesitation-induced slowdowns are the concrete cost that native continuation removes.

Problem

Action chunking gives VLAs real-time throughput but discontinuities at chunk boundaries. Real-Time Chunking patches this externally (inference-time guidance) β€” which causes spurious multimodal switching and trajectories that are never intrinsically smooth.

Method

Legato makes continuation native to training:

  • Denoising initializes from a schedule-shaped mixture of known (already-committed) actions and noise, exposing the model to partial action information during training;
  • The learned flow dynamics are reshaped so denoising stays consistent between training and inference under per-step guidance;
  • Randomized schedule conditioning during training supports varying inference delays and yields controllable smoothness.

Results (as reported)

  • Smoother trajectories, less spurious multimodal switching, less hesitation.
  • Across five real-world manipulation tasks: β‰ˆ10% improvements over RTC in both trajectory smoothness and task completion time.

Significance

The clearest instance of RSS 2026's "internalize the inference patch" pattern: RTC (NeurIPS 2025, test-time) β†’ training-time RTC (Ξ¨β‚€ adopts it after finding test-time guidance unstable) β†’ Legato (fully native continuation with delay randomization). Together with AR-VLA's persistent-context expert (#85) and Action-to-Action flow (#209), chunk-boundary handling has become its own sub-field β€” directly relevant to every flow-matching VLA in Review-VLA-Attention's streaming discussion.

← RSS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally