Skip to content

ICLR 2026 AsyncVLA

hwoo.han edited this page Jun 11, 2026 · 3 revisions

AsyncVLA β€” Asynchronous Foundation-Model + Edge Adapter for Navigation Under 6-Second Latency

Authors: Noriaki Hirose Β· Catherine Glossop Β· Dhruv Shah Β· Sergey Levine Affiliations: UC Berkeley Β· Toyota Motor North America (Hirose) Β· Princeton (Shah) Venue: arXiv preprint Β· Feb 13, 2026 Β· arXiv 2602.13476 Category: VLA Architecture (dual-system / hierarchical) Β· Navigation Trend tag: Trend 1 (system architecture) Β· Trend 7 (real-time deployment)

Approach diagram

flowchart LR
  subgraph Remote[Remote workstation Β· RTX 4090]
    OBS1["Past observation I_{t-k}"] --> BVLA
    INST[Language / 2D-pose goal] --> BVLA[OmniVLA<br/>8.26B params Β· 5 Hz]
    BVLA -- delayed action token<br/>embeddings via WiFi --> NET((WiFi Β· 0.28–6.0 s))
  end

  subgraph Edge[Robot Β· Jetson Orin 30W]
    NET --> EA[Edge Adapter<br/>76M params Β· 8 Hz]
    OBS2[Recent observation I_t] --> EA
    EA --> ACT[Refined action chunk]
    ACT --> PD[PD controller Β· 10 Hz]
    PD --> ROB[Vizbot ground robot]
  end

  classDef remote fill:#bbdefb,stroke:#1565c0,color:#000
  classDef edge fill:#c8e6c9,stroke:#2e7d32,color:#000
  classDef net fill:#fff9c4,stroke:#f57f17,color:#000
  class BVLA,OBS1,INST remote
  class EA,OBS2,ACT,PD,ROB edge
  class NET net
Loading

Problem

Robotic foundation models (OmniVLA, NaVILA-class) achieve strong open-world generalization by leveraging internet-scale vision-language pretraining, but their 8B-parameter inference cost creates a fundamental bottleneck for real-time deployment: a single forward pass can take hundreds of milliseconds on a workstation GPU, and multi-second WiFi round-trips push end-to-end latency to 0.28–6 seconds in realistic mobile-robot settings. With a moving robot in a dynamic environment (pedestrians, doors, other vehicles), the foundation model's "current" action chunk is already stale by the time it lands on the edge β€” collisions follow.

Two prior responses both fall short:

  • Run the small model only (OmniVLA-edge, 108 M params, 6 Hz onboard) β€” fast and reactive, but limited semantic capability; struggles with language-conditioned navigation and out-of-distribution prompts.
  • Run the big model only β€” full semantic capability, but at 5 Hz on workstation + WiFi delay it cannot react to dynamic obstacles.

Prior dual-system work (Fast-in-Slow, Ο€0.5, Hi Robot) uses synchronous coupling with sub-300 ms hand-offs β€” none has been demonstrated under multi-second edge-to-cloud latency.

Method

Two policies, asynchronously coupled.

  1. Base VLA (OmniVLA, 8.26 B): SigLIP + DINOv2 vision encoders + LLaMA-2-7B language model. Runs on a remote RTX 4090 workstation at 5 Hz. Emits action-token embeddings conditioned on a (delayed) past observation $I_{t-k}$ and the goal (language or 2D pose).
  2. Edge Adapter (76 M params): Lightweight onboard policy on Jetson Orin (30 W). Runs at 8 Hz, consumes the most recent observation $I_t$, the delayed past observation $I_{t-k}$, and the delayed action-token embeddings from the base VLA, and emits a refined action chunk that the PD controller (10 Hz) executes.

Why both observations. The Edge Adapter receives the same $I_{t-k}$ that the base VLA used so it can interpret the (delayed) embedding correctly, then conditions its refinement on the current $I_t$ to react to anything that has changed since.

Trajectory re-weighting during training. Sequences with large intra-chunk pose changes (final-pose distance > $d_\text{th} = 1.0$ m) are up-weighted, prioritizing samples that exhibit reactive behavior (collision avoidance, yielding to pedestrians) β€” i.e., training emphasizes exactly the cases where async correction matters.

Two-stage end-to-end finetuning.

  • Stage 1: Train Edge Adapter (ΞΈ, Ο†) with the base VLA frozen.
  • Stage 2: Fine-tune the entire system (including base VLA ψ) jointly so the two policies align their representations across the WiFi delay.

Both stages use an imitation loss + an action-smoothing loss.

Results

Vizbot ground robot. Two task families. Strong dynamic-obstacle setting (pedestrians, cluttered indoor + outdoor).

AsyncVLA OmniVLA (8.26 B remote, 5 Hz) OmniVLA-edge (108 M onboard, 6 Hz)
Success rate 85% 45% 25%
Static collisions 0.10 0.30 0.60
Dynamic collisions 0.10 1.05 1.00
Time to goal (s) 59.18 70.73 80.07

β†’ +40 pp success vs. SOTA baselines (85% vs. 45% for remote OmniVLA, 25% for OmniVLA-edge), >10Γ— fewer dynamic collisions (0.10 vs. ~1.0–1.05), faster goal-reaching (59.18 s vs. 70.73–80.07 s). Holds with WiFi latency injected up to 6 seconds. Note: the strong "Ours (workstation)" ablation (no edge, 89.79 s, 0.50 static / 0.67 dynamic collisions) underperforms the full async system β€” the edge adapter is what delivers the reactive gains.

Robust to:

  • 2D-pose-conditioned navigation (12–30 m, 10 environments).
  • Language-conditioned navigation (5–20 m, 12 environments β€” offices, kitchens, halls), including out-of-distribution prompts.

Significance

The first hierarchical VLA designed for seconds-scale latency, not milliseconds. Prior dual-system VLAs (Fast-in-Slow, Ο€0.5/Hi Robot, GR00T, ChatVLA-2) all assume the System-2 ↔ System-1 hand-off is < 300 ms β€” i.e., both systems are co-located on the same machine or rack. AsyncVLA breaks that assumption: System 2 lives in the cloud, System 1 lives on the robot, and they communicate over WiFi.

Two structural ideas are new:

  1. The edge model conditions on both observations ($I_t$ + $I_{t-k}$), not just the recent one β€” that's what lets it correctly interpret stale guidance.
  2. End-to-end fine-tuning across the WiFi gap β€” both policies are jointly optimized despite being non-co-located at deployment, so the foundation model learns to emit guidance that is robust to being interpreted by a 76 M edge model under delay.

Compared to neighboring 2025–2026 latency work:

  • Real-Time Chunking β€” handles chunk-boundary latency within one model via async inpainting. AsyncVLA handles inter-system latency between two models. Complementary, not overlapping.
  • Ο€0.7 β€” uses async subgoal-image refresh (every ~4 s) but the action expert and VLM are co-located. AsyncVLA is the WiFi-separated cousin.
  • Fast-in-Slow β€” ms-scale embedded dual-system. AsyncVLA is the s-scale physically separated dual-system.
  • Steerable Policies β€” adds a richer S2β†’S1 vocabulary; AsyncVLA decouples where S2 and S1 run.

Cross-domain note: this is a navigation paper, not manipulation. But the architecture template β€” large remote VLA + small edge adapter, end-to-end fine-tuned, observation-aware re-conditioning of stale embeddings β€” is the most plausible deployment story for any robot fleet that wants the semantic strength of an 8 B-class VLA without a workstation onboard. Direct read-across to mobile-manipulation Ο€0.5-class systems and to humanoid stacks where the S0/S1 layer must remain on-board.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally