-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 AsyncVLA
Authors: Noriaki Hirose Β· Catherine Glossop Β· Dhruv Shah Β· Sergey Levine Affiliations: UC Berkeley Β· Toyota Motor North America (Hirose) Β· Princeton (Shah) Venue: arXiv preprint Β· Feb 13, 2026 Β· arXiv 2602.13476 Category: VLA Architecture (dual-system / hierarchical) Β· Navigation Trend tag: Trend 1 (system architecture) Β· Trend 7 (real-time deployment)
flowchart LR
subgraph Remote[Remote workstation Β· RTX 4090]
OBS1["Past observation I_{t-k}"] --> BVLA
INST[Language / 2D-pose goal] --> BVLA[OmniVLA<br/>8.26B params Β· 5 Hz]
BVLA -- delayed action token<br/>embeddings via WiFi --> NET((WiFi Β· 0.28β6.0 s))
end
subgraph Edge[Robot Β· Jetson Orin 30W]
NET --> EA[Edge Adapter<br/>76M params Β· 8 Hz]
OBS2[Recent observation I_t] --> EA
EA --> ACT[Refined action chunk]
ACT --> PD[PD controller Β· 10 Hz]
PD --> ROB[Vizbot ground robot]
end
classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef net fill:#fff9c4,stroke:#f57f17,color:#000
class BVLA,OBS1,INST remote
class EA,OBS2,ACT,PD,ROB edge
class NET net
Robotic foundation models (OmniVLA, NaVILA-class) achieve strong open-world generalization by leveraging internet-scale vision-language pretraining, but their 8B-parameter inference cost creates a fundamental bottleneck for real-time deployment: a single forward pass can take hundreds of milliseconds on a workstation GPU, and multi-second WiFi round-trips push end-to-end latency to 0.28β6 seconds in realistic mobile-robot settings. With a moving robot in a dynamic environment (pedestrians, doors, other vehicles), the foundation model's "current" action chunk is already stale by the time it lands on the edge β collisions follow.
Two prior responses both fall short:
- Run the small model only (OmniVLA-edge, 108 M params, 6 Hz onboard) β fast and reactive, but limited semantic capability; struggles with language-conditioned navigation and out-of-distribution prompts.
- Run the big model only β full semantic capability, but at 5 Hz on workstation + WiFi delay it cannot react to dynamic obstacles.
Prior dual-system work (Fast-in-Slow, Ο0.5, Hi Robot) uses synchronous coupling with sub-300 ms hand-offs β none has been demonstrated under multi-second edge-to-cloud latency.
Two policies, asynchronously coupled.
-
Base VLA (OmniVLA, 8.26 B): SigLIP + DINOv2 vision encoders + LLaMA-2-7B language model. Runs on a remote RTX 4090 workstation at 5 Hz. Emits action-token embeddings conditioned on a (delayed) past observation
$I_{t-k}$ and the goal (language or 2D pose). -
Edge Adapter (76 M params): Lightweight onboard policy on Jetson Orin (30 W). Runs at 8 Hz, consumes the most recent observation
$I_t$ , the delayed past observation$I_{t-k}$ , and the delayed action-token embeddings from the base VLA, and emits a refined action chunk that the PD controller (10 Hz) executes.
Why both observations. The Edge Adapter receives the same
Trajectory re-weighting during training. Sequences with large intra-chunk pose changes (final-pose distance >
Two-stage end-to-end finetuning.
- Stage 1: Train Edge Adapter (ΞΈ, Ο) with the base VLA frozen.
- Stage 2: Fine-tune the entire system (including base VLA Ο) jointly so the two policies align their representations across the WiFi delay.
Both stages use an imitation loss + an action-smoothing loss.
Vizbot ground robot. Two task families. Strong dynamic-obstacle setting (pedestrians, cluttered indoor + outdoor).
| AsyncVLA | OmniVLA (8.26 B remote, 5 Hz) | OmniVLA-edge (108 M onboard, 6 Hz) | |
|---|---|---|---|
| Success rate | 85% | 45% | 25% |
| Static collisions | 0.10 | 0.30 | 0.60 |
| Dynamic collisions | 0.10 | 1.05 | 1.00 |
| Time to goal (s) | 59.18 | 70.73 | 80.07 |
β +40 pp success vs. SOTA baselines (85% vs. 45% for remote OmniVLA, 25% for OmniVLA-edge), >10Γ fewer dynamic collisions (0.10 vs. ~1.0β1.05), faster goal-reaching (59.18 s vs. 70.73β80.07 s). Holds with WiFi latency injected up to 6 seconds. Note: the strong "Ours (workstation)" ablation (no edge, 89.79 s, 0.50 static / 0.67 dynamic collisions) underperforms the full async system β the edge adapter is what delivers the reactive gains.
Robust to:
- 2D-pose-conditioned navigation (12β30 m, 10 environments).
- Language-conditioned navigation (5β20 m, 12 environments β offices, kitchens, halls), including out-of-distribution prompts.
The first hierarchical VLA designed for seconds-scale latency, not milliseconds. Prior dual-system VLAs (Fast-in-Slow, Ο0.5/Hi Robot, GR00T, ChatVLA-2) all assume the System-2 β System-1 hand-off is < 300 ms β i.e., both systems are co-located on the same machine or rack. AsyncVLA breaks that assumption: System 2 lives in the cloud, System 1 lives on the robot, and they communicate over WiFi.
Two structural ideas are new:
-
The edge model conditions on both observations (
$I_t$ +$I_{t-k}$ ), not just the recent one β that's what lets it correctly interpret stale guidance. - End-to-end fine-tuning across the WiFi gap β both policies are jointly optimized despite being non-co-located at deployment, so the foundation model learns to emit guidance that is robust to being interpreted by a 76 M edge model under delay.
Compared to neighboring 2025β2026 latency work:
- Real-Time Chunking β handles chunk-boundary latency within one model via async inpainting. AsyncVLA handles inter-system latency between two models. Complementary, not overlapping.
- Ο0.7 β uses async subgoal-image refresh (every ~4 s) but the action expert and VLM are co-located. AsyncVLA is the WiFi-separated cousin.
- Fast-in-Slow β ms-scale embedded dual-system. AsyncVLA is the s-scale physically separated dual-system.
- Steerable Policies β adds a richer S2βS1 vocabulary; AsyncVLA decouples where S2 and S1 run.
Cross-domain note: this is a navigation paper, not manipulation. But the architecture template β large remote VLA + small edge adapter, end-to-end fine-tuned, observation-aware re-conditioning of stale embeddings β is the most plausible deployment story for any robot fleet that wants the semantic strength of an 8 B-class VLA without a workstation onboard. Direct read-across to mobile-manipulation Ο0.5-class systems and to humanoid stacks where the S0/S1 layer must remain on-board.
- arXiv: 2602.13476
- HTML: https://arxiv.org/html/2602.13476
- OmniVLA (base model, 2024): arXiv 2406.04823
- Lineage: ViNT, NoMaD, LeLaN, GNM, SACSoN (Levine lab navigation series)
- In-depth review: AsyncVLA (long-form)
- Real-Time Chunking β intra-model async inpainting
- Fast-in-Slow β embedded ms-scale dual-system
- Ο0.7 β async subgoal-image refresh
- Steerable Policies β richer S2βS1 vocabulary
- VLA Architectures Review β dual-system category
- System 0/1/2 Review β hierarchical-control taxonomy
- Survey: VLA & Manipulation
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)