-
Notifications
You must be signed in to change notification settings - Fork 0
Review Realtime Execution
Topic survey (updated Aug 2026) Β· how big policies act smoothly at control rate: chunking, continuation, anytime decoding, async systems. Structure: π trend Β· βοΈ approaches Β·
β οΈ limitations. Companion pages: RTC Β· Legato Β· OAT Β· Fast-in-Slow Β· AsyncVLA Β· VLA Attention Β· Ξ¨β.
A 2β7B VLA takes 100β300 ms per forward pass; robots need 10β50 Hz control. Action chunking (predict H steps, execute while computing the next chunk) is the universal answer β and its failure mode is the chunk boundary: inference delay plus flow-policy multimodality make consecutive chunks disagree, producing visible hesitation, jitter, and collisions.
| Stage | Method | Mechanism | Status |
|---|---|---|---|
| 2025 | RTC (test-time) | Inference-time inpainting/guidance constrains the new chunk to overlap the executing one | Widely adopted; external to the policy β spurious multimodal switching, trajectories never intrinsically smooth |
| Late 2025 | Training-time RTC (PI) | Simulate inference delay during training: expose first d ~ U(0, d_max) actions un-noised, mask from loss | Adopted by Ξ¨β after test-time guidance proved unstable on their model β a documented negative result for test-time steering |
| RSS 2026 | Legato (native continuation) | Denoising initialized from a schedule-shaped mixture of committed actions + noise; flow dynamics reshaped for train/inference consistency; randomized schedules β controllable smoothness, variable delays | Beats RTC by β10% on both smoothness (NSPARC) and completion time across five real tasks |
| RSS 2026 | OAT (anytime decoding) | Ordered token space β every prefix decodes to a valid coarse action; more tokens refine it | A compute-quality dial rather than a boundary fix; AR-policies only |
ICML 2026 β the efficiency cluster (9+ papers), strongest of any venue. Streaming: Reflex achieves 50 Hz stable streaming (2.58Γ speedup, β54% reaction latency) via timestep-invariant attention partitioning with O(1) cache updates. Fewer/faster denoise steps: OMP one-step MeanFlow, STEP warm-started actions (+21.6% over baselines at 2 steps), Sparse ActionGen (4Γ). Token/channel pruning: GridS (β76% FLOPs, no drop), SpecPrune-VLA (1.57β1.70Γ), EcoVLA. Reasoning latency: Latent Reasoning VLA internalizes CoT (β90% inference latency), AVA-VLA adds confidence-gated early exit. Policy-agnostic: Speedup Patch (1.8Γ via safe chunk downsampling). Chunk coherence gets its own treatments (FocalPolicy frequency-optimized chunking).
Parallel line β architectural asynchrony: dual-rate systems (Fast-in-Slow: slow VLM cadence + fast action decoding), AsyncVLA, persistent-context AR experts (AR-VLA, RSS #85: the expert keeps its own history, VL prefixes refresh out-of-band), and action-to-action flow (RSS #209: warm-start denoising from the previous action instead of Gaussian noise).
| Approach | Pros | Cons | Choose when |
|---|---|---|---|
| Test-time RTC | Drop-in, no retraining | Multimodal switching; can be unstable (Ξ¨β's experience) | Frozen checkpoint you can't retrain |
| Training-time continuation (Legato-class) | Intrinsically smooth; delay-robust by randomization; faster completion | Requires retraining; delay range fixed at training | You own the training loop (the 2026 default) |
| Anytime tokens (OAT) | Graceful degradation under compute pressure | AR-family only | Variable compute budget / early-commit control |
| Async dual-rate | Decouples semantic and control rates | System complexity; two-model coherence | Reasoning-heavy tasks with fast reflex needs |
| Single-step action head (IROS 2026 π) | IMLE-VLA replaces the iterative diffusion/flow head with a 1-step cIMLE generator β 55 Hz (3.67Γ Ο0.5), LIBERO 98.0%, keeps multimodality + robustness | Head-swap needs retraining; single-step on high-precision contact untested | You want to kill the multi-step sampling latency at the source |
| Remote/cloud + RTC | Big models off-robot | Network variance (Qwen-RobotNav: server 196 ms avg but spiky vs Jetson-Thor 204 ms stable) | Fleet ops with reliable links; edge for latency-critical tasks |
- History conditioning raises the denoising bill: RobotManip's in-context variant needs 10 steps where the base needs 4 β robustness features and latency budgets trade off directly.
- Chunk length H couples everything: smoothness machinery, RL credit assignment (RECAP's N-step advantages), and reactivity limits (RobotManip lists fixed chunk length as a stated limitation for sub-second reactive control).
- Wiring choice sets the floor: prefix-KV reuse (Ο) vs full re-attention per denoise step (concatenation) β argued but never measured head-to-head (Review-VLM-Action-Connection).
- Latency is the least-reported number in the field. Among 2026 flagships, only scattered figures exist (Ο0.6 63 ms/chunk-class reports; Ξ¨β ~160 ms/pass; Qwen-RobotNav's deployment table); the Qwen manipulation suite reports none. No standard metric (ms/chunk? effective Hz? jerk?) exists.
- Edge deployment measurement is finally starting: the XPU characterization study profiles VLAs across GPUs/NPUs (compute-bound VLM phase, memory-bound expert phase; 2.9β3.3Γ via phase-aware scheduling) β but manipulation-side FP8/TensorRT deployment reports remain rare.
- Smoothness metrics (NSPARC etc.) are young; no benchmark scores hesitation/recovery under forced delay spikes.
- All continuation methods assume the policy is the bottleneck; perception-latency (multi-view encoding) is unaddressed.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)