Skip to content

RSS 2026 AR VLA

hwoo.han edited this page Aug 9, 2026 · 2 revisions

AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: VLA Models Β· paper #85 Authors: Yutong Hu, Jan-Nico Zaech, Nikolay Nikolov, Yuanqi Yao, Sombit Dey, Giuliano Albanese, Renaud Detry, Luc Van Gool, Danda Pani Paudel arXiv: 2603.10126 Β· program page

Summary compiled from the arXiv paper (v2, INSAIT + KU Leuven); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Reactive chunking VLAs vs AR-VLA's autoregressive action stream (Figure 1 of arXiv 2603.10126, Β© the authors)

Figure 1 contrasts (a) the prevalent chunking paradigm β€” each new observation re-encodes the scene, emits a static action chunk, and discards all temporal context after execution β€” with (b) AR-VLA, where an autoregressive action expert emits a continuous causal action stream with a long-lived memory while vision-language conditions (green backbone) are refreshed asynchronously without interrupting the actions.

Problem

VLAs labeled "autoregressive" only autoregress within a single inference step: diffusion/flow policies and chunking VLAs act as reactive, memoryless mappings that "wake up" at every perception step, discarding perception-action history ("Markovian amnesia"). This also couples fast control to slow VLM reasoning, causing inter-chunk gaps, jerky trajectories, and no mechanism for history-dependent tasks.

Method

AR-VLA is a standalone autoregressive Action Expert β€” a Transformer decoder over continuous action tokens (linear projection per timestep, regression head, no discretization) β€” conditioned on refreshable vision-language prefixes from a VLM (Paligemma-3B in experiments; ~3B + 300M scale). A Hybrid KV cache holds two streams under different update rules: a long rolling FIFO of proprioceptive/action history and a single-slot refreshable VL prefix. Dynamic Temporal Re-anchoring (DTR) uses RoPE shift-invariance to index atemporal VL tokens at their capture timestep, so attention depends only on staleness Ξ”t, bridging the training/inference index gap. Training is two-phase: action-only autoregressive pretraining of "kinematic syntax," then VL-action alignment with stochastic historical dropout to prevent over-reliance on history. A dual-thread deployment runs the action thread at high frequency while the perception thread updates the VL prefix asynchronously.

Results

On SimplerEnv (BridgeV2 pretraining, identical VLM backbone), AR-VLA averages 61.5% vs reproduced Pi-0-Fast* 49.0%, Pi-0.5* 51.0%, and CogACT 52.1%. Zero-shot real-world WidowX: 89% average across five challenging tasks (100% on cup-on-plate and lobster), far above the Pi reimplementations. As a specialist action head it hits 97.33%/67.33% on ALOHA cube transfer (scripted/human) vs ACT's 86%/50%. Latency: 46.25 ms effective per action (28.86 ms expert) with the lowest average and max jerk among OpenVLA/Fast/flow-matching baselines, holding ~29 ms per-action control even with a 70 ms perception backbone. On history-dependent tasks, AR reaches 66.7% on PushT2 (vs DP 44.0, ACT 34.0) and 81.2% on Stack3 with a 40-step context (vs 56.3% flow matching).

Significance

Argues that the action head β€” not just the backbone β€” should be truly autoregressive across time, offering context-awareness and smoothness as structural properties rather than add-on memory modules. Connects to the architecture debate in Review-VLA-Architecture and the asynchronous fast/slow execution themes in Review-Realtime-Execution.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally