-
Notifications
You must be signed in to change notification settings - Fork 0
Review Qwen VLA
In-Depth Review β Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Paper: Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Authors: Qiuyue Wang*, Mingsheng Li*, Jian Guan*, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Baiβ (corresponding), Jingren Zhou, + 24 contributors Affiliation: Qwen Team, Alibaba Group (Tongyi Lab) arXiv: 2605.30280 Β· v1 May 28, 2026 Β· v2 Jun 1, 2026 (34 pages, cs.RO) β this review is verified against v2; the substantive v1βv2 additions are the no-T2A baseline (60.9%, making T2A worth +10.2 pp) and the T2A action-representation clarification in Β§4.2 Code: github.com/QwenLM/Qwen-VLA Β· Blog: qwen.ai/blog?id=qwenvla Weights / License: model weights and license terms not stated in the paper as of this review
This is the long-form companion to the per-paper summary. Companion reviews to read alongside: Ο series evolution Β· GR00T series Β· LBM Co-training Β· Knowledge Insulation Β· VLA Architectures Β· VLMβAction Connection.
- The Qwen team's first dedicated VLA. A Qwen3.5-4B vision-language backbone with early multimodal fusion is paired with a separate 1.15B DiT flow-matching action expert. Architecturally this is the same family as Ο0.6 / Ο0.7 (Category B in Review-VLA-Architecture) β not a same-stack MoE or a cross-attended dual-system head like GR00T.
- One generalist, many embodiments. A single set of weights handles WidowX, Google Robot, Franka Panda, ARX5, Fourier GR-1, Mobile ALOHA, AgiBot A2-D, Galaxea R1, AIRBOT MMK2, TienKung, and human MANO hands. The platform is selected by embodiment-aware prompt conditioning (plain text describing arms / waist / mobile base / control frequency / chunk size) β no per-embodiment heads, no architectural switches.
- A four-stage training recipe built around a "compression" thesis: (I) T2A β text-only DiT pretraining with images suppressed, (II) CPT β joint multimodal continued pretraining, (III) SFT in two parallel branches (multi-task / real-robot), (IV) PPO RL in SimplerEnv with sparse binary rewards and an analytic flow-matching log-probability via an ODEβSDE conversion.
- Headline numbers as a single generalist: 97.9% LIBERO, 73.7% Simpler-WidowX, 86.1 / 87.2% RoboTwin-Easy / Hard, 56.7% RoboCasa-GR1, 69.0% OSR / 57.5% SR on R2R, 59.6% SR on RxR, 76.9% average OOD success on ALOHA real-robot, 26.6% zero-shot SR on DOMINO dynamic manipulation. The DOMINO number beats every published fine-tuned baseline including PUMA (17.2%) without any DOMINO data.
- Recipe verdicts that align with TRI's LBM study. VL co-training helps fine-grained-recognition benchmarks (+4.9 pp RoboCasa-GR1, +4.6 pp RoboTwin) and never interferes; an ablation shows a pretrained DiT outperforms a from-scratch DiT throughout SFT. Explicit proprioceptive state in either the VLM prompt or the DiT yields at most +1.3 pp β the embodiment text prompt is sufficient.
By Q2 2026 the VLA field has three publicly comparable "VLM-developer-as-VLA-author" lineages:
| Lineage | VLM developer | Their VLA | Architecture family |
|---|---|---|---|
| Google DeepMind | Gemini | Gemini Robotics 1.5 (Sep 2025) | Closed dual-system |
| AllenAI | Molmo | MolmoAct (2508.07917) | Reasoning-augmented (Cat G) |
| Alibaba Qwen | Qwen3-VL / Qwen3.5 | Qwen-VLA (May 2026, this paper) | Flow-matching expert (Cat B) |
Until May 2026 the Qwen3-VL backbone had been a building block for other people's VLAs β GR00T N1.7 uses Qwen3-VL-2B as Cosmos-Reason2-2B; NORA uses Qwen2.5-VL-3B; the StarVLA family uses Qwen2.5/3-VL across multiple action-head variants. Qwen-VLA is the first time the Qwen team itself ships a VLA, on a 4B Qwen3.5 (not Qwen3-VL) backbone with native multimodal early-fusion. This matters for two practical reasons:
- Backbone identity is no longer ambiguous. Reading the paper, the Qwen team picks the Qwen3.5 family β a hybrid of gated linear attention and grouped-query softmax attention (Bai et al., Qwen3-VL technical report, arXiv 2511.21631) β rather than Qwen3-VL. The 4B size matches what PI and TRI have settled on for their LBM-class models (Gemma3-4B for Ο0.6/Ο0.7 and PaliGemma2-3B for LBM). The convergence is striking: three independent teams, three different VLM lineages, similar ~4B size, all paired with similar-sized flow-matching action experts.
-
Open-source posture is at least partially open. Code is published at
github.com/QwenLM/Qwen-VLA. The paper itself does not state the weight license, and as of this review the repository content is documentation rather than full weights, so the practical openness story is still developing. This puts Qwen-VLA somewhere between fully-closed (Ο series, Gemini Robotics) and fully-open (GR00T N1.7 under Apache 2.0).
The other reason it matters is the DOMINO 26.6% zero-shot SR. DOMINO is a 2026 dynamic-manipulation benchmark (Fang et al., 2026) that other models fine-tune on β and Qwen-VLA-Instruct, evaluated zero-shot with current-frame observations only, surpasses the best fine-tuned baseline PUMA (17.2%) by 9.4 percentage points. That single number is the most defensible "the joint-pretraining bet paid off" claim in the paper.
flowchart TB
subgraph PROMPT["Embodiment-aware prompt + task instruction"]
P1["The robot is {robot_tag} with {single/dual arms}[, waist][, and mobile base]."]
P2["The control frequency is {FPS} Hz."]
P3["Please predict the next {chunk_size} control actions to execute the following task: {ori_instruction}."]
end
subgraph IMG["Multi-view observations with view-tag tokens"]
V1["ego camera"]
V2["cam_left_wrist"]
V3["cam_right_wrist"]
VTAGS["wrapped as <|tag_start|> image <|tag_end|>"]
end
PROMPT --> VLM["Qwen3.5 (4B) β natively multimodal<br/>ViT with spatial merging Β· interleaved visual + text tokens<br/>Gated linear attention (majority of layers)<br/>+ GQA softmax attention at intervals<br/>+ M-RoPE"]
IMG --> VLM
VLM -- hidden states --> CAT["Concatenate VLM hidden states with noisy action chunk"]
NOISE["Noisy action Y_tau in R^HxK"] --> CAT
CAT --> DIT["DiT-style flow-matching action expert Β· 1.15B<br/>16 blocks Β· joint self-attention<br/>AdaLN timestep Β· multi-section RoPE aligned with backbone"]
TIME["timestep tau"] --> ADALN["AdaLN modulation"]
ADALN --> DIT
DIT --> VEL["velocity field v_theta"]
VEL --> EULER["Few Euler integration steps<br/>tau = 1 to 0"]
EULER --> ACT["Action chunk Y_0 in R^HxK<br/>(zero-padded, mask-aware)"]
VLM -. next-token CE on auxiliary text .-> LMHEAD["LM head"]
LMHEAD --> VLLOSS["L_vl"]
VEL --> FMLOSS["L_act"]
classDef vlm fill:#bbdefb,stroke:#1565c0,color:#000
classDef act fill:#c8e6c9,stroke:#2e7d32,color:#000
classDef inp fill:#fff9c4,stroke:#f57f17,color:#000
class VLM vlm
class DIT,ADALN,VEL,EULER,ACT act
class PROMPT,IMG inp
- Qwen3.5-4B (the paper cites Team, 2026 β pointing at the still-in-progress Qwen3.5 line, not the Nov 2025 Qwen3-VL technical report).
- Natively multimodal with early fusion. Visual tokens are produced by a ViT with spatial merging and interleaved directly into the text token stream. There is no separate vision encoder bolted on after the fact.
- Hybrid attention. Majority of layers use gated linear attention (cf. Qwen3.5's Gated DeltaNet choice); at regular intervals layers use grouped-query (GQA) softmax attention for full global reasoning. This is the Qwen3.5 standard recipe, inherited unchanged.
- Multi-section RoPE (M-RoPE) for the image, text, and action token sub-sequences.
- No separate vision encoder named (no SigLIP / DINOv2 / Eagle citation) β the ViT is part of the Qwen3.5 stack.
- Single-stream DiT (Esser et al., 2024) with flow-matching objective (Lipman et al., 2023).
- Parameter budget β exact breakdown from Β§2.2 of the paper:
| Component | Parameters |
|---|---|
| 16 DiT blocks (70.8M each) | 1.13B |
| Action projection MLPs (raw action dim β DiT latent) | 4.9M |
| VLM hidden states β DiT channel linear | 3.9M |
| Timestep embedding | 2.8M |
| Output AdaLN modulation | 4.7M |
| Total action expert | ~1.15B |
- Connection mechanism: the action expert concatenates VLM hidden states with the noisy action chunk into one sequence and processes them via joint self-attention with AdaLN-injected timestep conditioning. There is no cross-attention into the VLM; there is no shared transformer either. This sits between Cat B (Ο-style separate-expert-with-attention-routing) and Cat C (GR00T-style cross-attention-into-VLM-features), and the paper itself frames it as a decoupled design that "lets the action expert specialize in fine-grained action generation".
- Inference: Few Euler integration steps from Ο = 1 to Ο = 0. The paper does not give a single canonical step count, but its DOMINO discussion describes "coherent action chunks" so the recipe matches the small-K (β5) regime of Ο0.6/Ο0.7 rather than the K=50 of original Ο0.
The single design choice that lets one DiT handle 11+ embodiments is:
- Fixed tensor interface Y β R^(HΓK). H is a fixed prediction horizon; K is a fixed channel dimension shared across all control modes.
- Active channels and zero padding. A control mode uses c β€ K channels. The c task-relevant values live in the leading c dims; the remaining K β c are zero-padded.
- Per-channel binary mask M β {0,1}^(HΓK): M_{h,k} = 1 iff k < c and h < H_task. The mask is used to exclude padded entries from the gradient.
- No embodiment-specific output head. One DiT, one set of weights, switched only by the embodiment prompt + dataset-specific quantile normalization.
This is the same family of "shared latent, per-embodiment projection" idea that GR00T implements with MultiEmbodimentActionEncoder + CategorySpecificMLP, but Qwen-VLA simplifies further: the Β§5.2.2 ablation shows that zero-padding with a single shared MLP performs within 1.2 pp of per-embodiment Multi-MLP or Concatenation projections on Bridge + Robocasa, so the team picks zero-padding as the default.
In the Review-VLA-Architecture taxonomy this is Category B (Flow-matching action expert), but with a concatenation attention routing rather than the prefix-KV / same-stack-MoE routing of the Ο series. Specifically:
- Ο0.6 / Ο0.7: same-stack MoE β action expert is a parallel branch within the same transformer attending to a prefix-KV cache from the VLM.
- GR00T N1 β N1.7: cross-attention from a separate DiT into truncated VLM hidden states (Cat C).
- TRI LBM: adaLN-conditioned 8-layer flow expert reading a single observation token aggregated from the last 4 VLM layers.
- Qwen-VLA: concatenate VLM hidden states + noisy action tokens β joint self-attention in a 16-block DiT. No cross-attention into VLM; no shared backbone transformer; one-direction information flow (VLM features feed in, no action-gradient back-prop into VLM during T2A; backbone unfrozen during CPT).
This makes Qwen-VLA's wiring a third-way that is closest in spirit to the simpler "DiT receives VLM features as context tokens" pattern that DiT4DiT and several 2026 VAM-class systems use, applied to a fully VLA setting.
The paper does not state a single canonical inference setting. From the experimental sections:
- Action chunk H = 16 for all sim manipulation benchmarks (LIBERO, RoboCasa-GR1, Simpler-WidowX, RoboTwin 2.0).
- Action chunk H = 8 waypoints for navigation (R2R, RxR).
- RL stage uses H = 16 and 128 parallel environments.
- Temperature: Ο = 1.0 during PPO rollouts, Ο = 0.6 at evaluation time, to sharpen the action distribution.
- No published latency / ms-per-chunk number β a gap relative to Ο0.6 (63 ms / chunk on H100) and GR00T N1 (63.9 ms / 16-action chunk on L40).
The paper's training recipe is built on an explicit thesis:
"A language instruction such as 'pick up the red cup' together with an embodiment prompt compactly encodes the task intent in a handful of tokens, yet the corresponding action trajectory may span hundreds of high-dimensional joint-position values. Bridging this dimensionality gap is a structured decompression problem."
The four-stage recipe is built so each stage closes one specific gap from the one before it:
flowchart LR
S1["Stage I β T2A<br/>VLM FROZEN<br/>Train DiT only<br/>Text + embodiment prompt β action<br/>No images (deliberately)"]
S2["Stage II β CPT<br/>VLM UNFROZEN<br/>Joint multimodal training<br/>Heterogeneous mixture (Table 1)<br/>Both sim + real-robot data"]
S3a["Stage III-a β Multi-task SFT<br/>VLM + DiT unfrozen<br/>VQA + spatial grounding<br/>+ manipulation + navigation<br/>Embodiment- and task-balanced"]
S3b["Stage III-b β Real-robot SFT<br/>In-house ALOHA teleop<br/>Tests CPT-to-hardware transfer"]
S4["Stage IV β RL<br/>PPO + GAE<br/>Sparse binary rewards<br/>SimplerEnv rollouts only<br/>β Qwen-VLA-Instruct"]
S1 --> S2
S2 --> S3a
S2 --> S3b
S3a --> S4
classDef ph fill:#fff9c4,stroke:#f57f17,color:#000
class S1,S2,S3a,S3b,S4 ph
The most distinctive design choice. The VLM is frozen; only the DiT trains; images are deliberately suppressed. The decoder must reconstruct action distributions from language + embodiment text alone. This is the "decompression prior" β the DiT learns how language indexes regions of action space before any visual grounding is available.
The Β§5.2.1 ablations are unusually thorough and produce four specific recommendations:
- T2A itself is worth +10.2 pp (added in v2): the no-T2A baseline scores 60.9%, vs 71.1% for the best T2A configuration β the most direct quantification of the warm-start's value.
- Data composition. Pure synthetic gets 64.1% downstream SFT, pure real gets 51.0%; the optimum is ~20% synthetic + 80% real, reaching 71.1%. Real anchors the prior in plausible dynamics; synthetic broadens languageβaction coverage.
- Full-sequence prediction beats chunk prediction. +4.9 pp at 10% synthetic data, +2.9 pp at 0% synthetic. The decoder needs to see trajectory-level coherence; chunks fragment it.
- Sigmoid-Normal timestep distribution at T2A is +5.7 pp over Beta. Without visual conditioning, intermediate noise levels carry the most learning signal for the language-action prior. Beta is then used at CPT/SFT where rich VLM conditioning is available.
- T2A converges fast: 2,000 steps is optimal. Performance plateaus through 10k steps (67.5% / 67.2%); at 40k steps overfitting kicks in (60.4%, vs. 71.1% at 2k).
- Action representation is identical between T2A and downstream stages (clarified in v2): actions are delta end-effector displacements relative to the first frame of each chunk, with the embodiment prompt carrying the platform, normalisation convention, and prediction horizon β so the language-only prior transfers directly into the visually-conditioned stages.
The "compression" framing is original β no other major VLA paper organizes the warm-start this way. It is functionally related to the language-as-warm-start arguments in OpenVLA-OFT and the action-prior pretraining of GR00T-N1's pseudo-action latent stage, but it is more aggressive: Qwen-VLA forces the prior to be purely language-conditioned.
- Both VLM and DiT unfrozen.
- Heterogeneous data mixture (Table 1 of the paper):
| Source | Proportion |
|---|---|
| Robot Manipulation Trajectories | 74.2% |
| Navigation Trajectories | 7.5% |
| Human Egocentric Trajectories | 6.0% |
| Synthetic Simulation Trajectories | 3.7% |
| General Vision-Language Data | 3.4% |
| Spatial Grounding (2D) | 2.5% |
| Autonomous Driving VQA | 2.4% |
| Fine-Grained Embodied Action Caption | 0.2% |
This is the data-axis identity of Qwen-VLA: heavily robot-trajectory-weighted (74.2%), with VL co-training as a minority but explicitly-purposeful contributor.
- Multi-task SFT. Joint fine-tune on VQA + spatial grounding + manipulation + navigation under embodiment-balanced + task-balanced sampling. SFT loss weights: 0.1 for vision-language next-token prediction, 1.0 for both manipulation and navigation action prediction. The 10Γ ratio gives gradient capacity to action while preserving language grounding.
- Real-robot SFT. In-house ALOHA teleop. Tests whether CPT's cross-domain priors transfer to physical hardware.
This is the second distinctive contribution after T2A. Key details:
- Algorithm: PPO with GAE (Ξ³ = 0.99, Ξ» = 0.95, Ο΅ = 0.2). Four optimization epochs per rollout batch.
- Value head: A lightweight value head is attached directly to the VLM backbone. The head mean-pools all VLM hidden states and maps them to a scalar via a linear projection. Stop-gradient on the VLM hidden states before they enter the value head β value-function gradients do not back-prop through the pretrained backbone. The value head uses LR = 10β»β΄, β20Γ the actor LR of 5Γ10β»βΆ, so it converges fast while the policy updates remain conservative.
- Log-probability under flow matching. The headline RL contribution: the deterministic probability-flow ODE is converted into a corresponding SDE by injecting controlled noise at each Euler denoising step (Song et al., 2021). Each transition then becomes an explicit Gaussian whose log-probability can be computed analytically without numerical ODE integration. During rollout the intermediate denoising states are stored; at the PPO update the velocity field is re-evaluated under the current parameters and Gaussian log-probability is recomputed. By default a single denoising step is randomly sampled per rollout for the log-prob estimate, requiring only one extra DiT forward pass during recomputation.
- Reward: sparse binary β R = 1 if the task goal is achieved at episode end, 0 otherwise. No learned reward model.
- Rollouts: 128 parallel environments, 8 rollout epochs Γ 128 environment steps per iteration = 8,192 transition chunks per iteration (with chunk length H = 16).
- Training distribution: RL rollouts collected exclusively in SimplerEnv.
This is a meaningful contribution. Flow-matching policies have notoriously awkward log-probabilities β the closest 2025β2026 publications are ReinFlow (learnable noise injection for exact likelihoods) and PI's RECAP (advantage-conditioned RL that avoids needing a log-prob). Qwen-VLA picks a third path: a per-step Gaussian-stochasticization that lets standard PPO clip the importance ratio without modifying the policy parameterization.
The flow-matching action loss with the per-channel mask M (Eq. 1β2 of the paper):
The two-level averaging (first across timesteps for each active channel, then across active channels) ensures each control dimension contributes equally regardless of how many channels a given embodiment uses, and padded positions are fully excluded.
The next-token vision-language loss (Eq. 3):
Joint loss (Eq. 4):
The exact pretraining Ξ» values are not stated. The SFT-stage weights are Ξ»_vl = 0.1, Ξ»_act = 1.0 (action β« VL by 10Γ).
| Detail | What the paper says |
|---|---|
| Optimizer | Not explicitly stated (AdamW assumed from context) |
| Peak LR | RL stage: 5Γ10β»βΆ (actor), 1Γ10β»β΄ (value head). Pretraining/SFT LRs not stated as a single number; the paper says "cosine-decayed learning rates with separate group-wise schedules for VLM backbone and action decoder" |
| Total steps | Not stated. T2A converges in 2k steps; SFT ablation curves go up to 80k steps |
| Batch size | RL: 8,192 transition chunks per iteration. Pretraining/SFT BS not stated |
| Hardware | Not stated |
| GPU-hours | Not stated |
| Inference latency (ms/chunk) | Not stated |
This is the biggest reproducibility gap. The paper is unusually detailed on the RL stage and the T2A ablations but does not give a single canonical pretrain compute budget.
Section 3.2 of the paper lists every source explicitly. Distilled:
- Real-robot public datasets (74.2% mixture, including all robot-trajectory tiers): RobotSet, Galaxea, AgiBot World, RoboCOIN, RoboMIND V1/V2, RDT-1B, DROID, BridgeData V2, RH20T, RT-1, BC-Z β over 10,000 hours total.
- In-house real-robot: over 1,000 hours, ~20% of the total mixture.
- Simulation trajectories (3.7% of mixture, "over 8M synthetic" in total): InternData-A1 + GR00T-X-Embodiment-Sim (public sim) + the team's own RoboInf vision-conditioned synthesis pipeline (random-placement scene generation: 20 tabletop scenes Γ 10 init configurations = 200 base scene configs, 450 tasks, 300 trajectories each, plus randomization over 3K backgrounds and 1K table textures), which yields 359,848 full successful trajectories including subtask segments (Β§3.2.3). The "over 8M synthetic" headline figure is the total synthetic count and is dominated by the 7.2M language-only trajectories below, not by RoboInf vision-conditioned data.
- Stage-I-specific language-only action data: six task templates (pick-and-place, push, pull, rotation with reposition, rotation toward viewpoint, swap) Γ six robots (Franka Panda, UR10e, UR5e, Kinova Gen3, TM12, xArm7) Γ ~200k trajectories each = ~7.2M trajectories, >14,000 hours of motion-planned simulated robot trajectories at 50 Hz, without any rendering or physics simulation.
- Egocentric human (6.0%): Ego4D + EPIC-KITCHENS subsets processed by VITRA (atomic manipulation segments + fine-grained captions + 3D hand & camera trajectories); EgoDex (829h dexterous, Apple Vision Pro, 194 tasks); EgoVerse (1,300+ hours, 1,965 tasks, 240 scenes); Xperience (depth + hand + body mocap + hierarchical annotations). Hand articulation is compressed to 10 PCA eigengrasp coefficients per hand from the 45-DoF axis-angle MANO pose; per-hand wrist motion is 6 SE(3) parameters β 32 dims per time step for ego-human data.
- Navigation (7.5%): instruction-following (4.3%) + object searching (2.3%) + target tracking (1.0%); ~2 FPS video sampling; 3-DoF mobile robot (translation + heading).
- Auxiliary VL (8.5% combined): general VL (3.4%) + spatial grounding (2.5%) + autonomous driving VQA (2.4%) + fine-grained embodied action caption (0.2%; ~48k video-caption pairs annotated along 13 dimensions by a two-stage Qwen3.6-plus pipeline).
The autonomous-driving VQA subset is particularly notable β it includes LingoQA, DriveAction, MMAU, Impromptu-VLA, nuScenes-QA, nuScenes-MQA, MapLM, WaymoQA, CODA-LM, Talk2Car, DrivingVQA, DriveLM, W3DA, GRAID, Bench2Drive-VL, DriveGPT4, OmniDrive, Senna, NAVSIM-RecogDrive, NAVSIM-Traj. The motivation is "viewpoint-robust localization", "temporal scene understanding", "language-grounded localization", and "planning-aware reasoning" β i.e. AD data is being recycled as embodied-perception co-training, not because Qwen-VLA targets driving.
Every training sample is prefixed with:
The robot is {robot_tag} with {single arm / dual arms}[, waist][, and mobile base]. The control frequency is {FPS} Hz. Please predict the next {chunk_size} control actions to execute the following task: {ori_instruction}.
Eleven representative embodiments are listed in Table 2 of the paper:
| Robot | Arms | Action type |
|---|---|---|
| WidowX | Single | ΞEEF + G |
| Google Robot | Single | ΞEEF + G |
| Franka Panda | Single / Dual | ΞEEF + G; Abs Joint + G |
| ARX5 | Dual | ΞEEF + G |
| Fourier GR-1 | Dual | ΞEEF + G |
| Mobile ALOHA | Dual | ΞEEF + G; Abs Joint + G |
| AgiBot A2-D | Dual | Abs Joint + G; Abs Joint + DH |
| Galaxea R1 | Dual | Abs Joint + G |
| AIRBOT MMK2 | Dual | Abs Joint + DH |
| TienKung | Dual | Abs Joint + G; Abs Joint + DH |
| Real Human | Dual | ΞEEF (from MANO) |
Action values are normalized per-dataset using 1st / 99th percentile quantile mapping to [β1, 1] (Eq. 5 of the paper). Each dataset preserves its native control convention; the embodiment prompt + quantile normalization carries the platform-specific semantics.
This section answers the central question β where does Qwen-VLA sit relative to the leading 2026 VLA recipes?
| Axis | Qwen-VLA | Ο0.6 / Ο0.7 | TRI LBM | GR00T N1.7 |
|---|---|---|---|---|
| Backbone | Qwen3.5-4B (natively multimodal, gated-linear + GQA hybrid, M-RoPE) | Gemma3-4B (SigLIP 400M + Gemma3-4B LLM) | PaliGemma2-3B (SigLIP + Gemma2 LLM) | Qwen3-VL-2B / Cosmos-Reason2-2B |
| Backbone params | ~4B | ~4.4B (4B LLM + 400M SigLIP) | ~3B | 2.44B |
| Vision encoder | ViT with spatial merging, integrated into backbone | SigLIP 400M, frozen-then-unfrozen | SigLIP, frozen | Qwen3 native vision encoder (truncated at last layer) |
| Backbone status during action training | Frozen at T2A, unfrozen at CPT (joint training) | Unfrozen but gradient-insulated via Knowledge Insulation | Frozen throughout action training |
Top-N LLM layers unfrozen (tune_top_llm_layers=4 for N1.6, inherited at N1.7) |
| Attention recipe | Gated linear attention majority + GQA at intervals | Full-attention transformer | Full-attention transformer | Full-attention transformer + vlln + vl_self_attention (N1.7) |
| Native multimodal early fusion | Yes (image tokens interleaved in text stream) | No (SigLIP β projection β LLM) | No (PaliGemma2 standard fusion) | No (Qwen3-VL native vision; not fully early-fused) |
| Cross-embodiment design ethos | Single set of weights + embodiment text prompt | Cross-embodiment training, no per-embodiment head | Per-target-platform specialization in phase 3 |
MultiEmbodimentActionEncoder + CategorySpecificMLP
|
Two non-trivial observations:
- The Qwen team chose Qwen3.5 over Qwen3-VL. This is consequential: Qwen3.5 uses a different attention recipe (gated linear attention + GQA intervals) than Qwen3-VL (pure GQA + QK-Norm β see ML-Attention table). Whether this matters for embodied tasks is not ablated in the paper.
- Qwen-VLA is the only one of the four that explicitly froze the backbone at the warm-start stage (T2A) β Ο series freezes nothing (KI gradient-routes instead); LBM keeps PaliGemma2 frozen throughout; GR00T does partial top-N unfreeze. The Qwen-VLA recipe sits between "all gradients allowed everywhere" and "VLM is forever frozen".
Three competing patterns are now well-documented in the field; Qwen-VLA picks a fourth:
| Pattern | Exemplar | Mechanism |
|---|---|---|
| Same-stack MoE + prefix-KV | Ο0.6 / Ο0.7 | Action expert is a parallel branch within the same transformer; attends to prefix-KV from VLM |
| Cross-attention into truncated VLM | GR00T N1 β N1.7 | Separate DiT cross-attends to VLM hidden states; vlln + vl_self_attention preprocessing in N1.7 |
| AdaLN-conditioned on aggregated VLM features | TRI LBM | 8-layer ActionFT consumes a single observation token built from last 4 VLM layers, gated by adaLN |
| Concatenation + joint self-attention | Qwen-VLA | Noisy action chunk concatenated with VLM hidden states into one sequence; DiT processes them through joint self-attention with AdaLN timestep injection |
The Qwen-VLA choice is closer to the DiT-style image diffusion lineage (Esser et al., 2024 β Stable Diffusion 3 / SD3 family) than to any 2025 VLA. The action expert is not a separate cross-attender; it is a transformer reading a concatenated [VLM_features | noisy_action] sequence. This is the simplest possible "feed in the VLM features and let attention sort it out" approach, and the Β§5.2.2 ablation (pretrained DiT outperforms from-scratch DiT throughout SFT) is the empirical justification.
The cost of this choice: every action-expert forward pass re-attends over the full VLM hidden state. Compared to Ο0.6/Ο0.7's prefix-KV cache, this is in principle a higher inference cost per chunk. The paper does not publish latency, so this is unverified.
| System | Action head class | Objective | Tokenizer / discretization |
|---|---|---|---|
| OpenVLA / Ο0-FAST | Cat A (AR tokens) | Cross-entropy | FAST DCT tokens (vocab 2048) |
| DDVLA / dVLA | Cat D (discrete diffusion in VLM) | Masked cross-entropy | Discrete action vocabulary |
| Ο0.6 / Ο0.7 / LBM / GR00T | Cat B (separate flow-matching expert) | Flow matching MSE on continuous targets | Continuous (no action tokenizer) β Ο series adds FAST as VLM CE target via KI |
| Qwen-VLA | Cat B (separate flow-matching expert) | Flow matching MSE + standard NTP on language | Continuous (no action tokenizer); auxiliary CE only on language |
Qwen-VLA is unambiguously Category B in the Review-VLA-Architecture taxonomy. There is no FAST head, no VQ-VAE head, no latent action token head, no auxiliary CE on discretized actions β only the flow-matching loss on continuous targets and the standard next-token CE on auxiliary language data. This aligns with TRI LBM's finding that discrete action tokens (FAST, VQ-VAE) provide no significant benefit at LBM scale and sometimes hurt β Qwen-VLA never tried them. Whether this is because the team independently arrived at the same conclusion or because they read TRI's Feb 2026 study is not addressed.
| System | Robot teleop | Cross-embodiment OXE | Human ego-video | Web VL | Auxiliary | Total scale |
|---|---|---|---|---|---|---|
| Qwen-VLA | >1,000 h in-house + >10,000 h public (74.2% mix) | Same OXE bucket β not separated | ~3,000+ h (Ego4D, EPIC-KITCHENS via VITRA + 829 h EgoDex + 1,300 h EgoVerse + Xperience) at 6.0% | 3.4% + spatial grounding 2.5% + AD VQA 2.4% + caption 0.2% = 8.5% | "over 8M" synthetic total = 359,848 RoboInf vision-conditioned + 7.2M language-only (3.7% mix) | >10,000 h robot + ~3,000+ h human + 8.5% VL |
| Ο0.6 / Ο0.7 | PI internal (very large, not disclosed) | OXE + DROID + community | Egocentric human (introduced in Ο0.7) | Web multimodal + detection + captioning | RL rollouts + failures (Ο0.7) | Not separately disclosed |
| TRI LBM | 523 h TRI-Ramen (target) | 1,150 h OXE-Ramen (12 robots, 924 tasks, 466k demos) | 2,271 h (Ego4D + EgoDex + Sth-Sth V2 + Epic Kitchen + HoloAssist) | 50M VL samples (RoboPoint + RefSpatial) | GPT-5 robot captions + GPT-5 human captions | 523 + 1,150 + 2,271 h + 50M VL |
| GR00T N1.7 | Several thousand hours teleop (YAM + AGIBot Genie1 + Galaxea R1 Pro + Unitree G1 + BEHAVIOR sim + DROID) | Cross-embodiment throughout, all stages | 20,854 h EgoScale (Apple Vision Pro EgoDex 829 h + in-the-wild ~20,000 h) | (less prominent) | DexMimicGen 780k sim + DreamGen neural trajectories | ~3,000+ h robot + ~20,854 h human |
Where Qwen-VLA sits on the scale axis: comparable to Ο / GR00T on robot trajectories (>10,000 h), comparable to LBM on human-video (~3,000 h is between LBM's 2,271 h and N1.7's 20,854 h), and lower than LBM on VL data (LBM's 50M RoboPoint + RefSpatial is the entire mixture; Qwen-VLA's 8.5% combined VL share is closer to Ο series' "co-training as side channel" framing). The data thesis is robot-trajectory-heavy with VL as regularizer, which is closest to Ο series rather than to LBM.
| Axis | Qwen-VLA | Ο0.6 / Ο0.7 | TRI LBM | GR00T N1.6 / 1.7 |
|---|---|---|---|---|
| Number of phases | 4 (T2A β CPT β SFT-multi-task β₯ SFT-real-robot β RL) | 1 (joint with KI) | 3 (pure co-train β joint specialization β target specialization) | 1 (joint cross-embodiment) |
| Backbone freezing | Frozen at T2A, unfrozen at CPT | Unfrozen, gradient-insulated (KI) | Frozen throughout | Top-N LLM layers unfrozen (N=4 reported in N1.6) |
| Action-head warm-start | T2A: text-only DiT pretraining | Joint from scratch | Joint from scratch (PaliGemma frozen, ActionFT random) | Joint from scratch |
| Auxiliary CE objective | NTP on language only (no FAST / discrete action target) | FAST tokens CE as VLM training signal (via KI) | NTP on language (RoboPoint/RefSpatial QA, GPT-5 captions, human captions); FAST/VQ-VAE/latent-video heads ablated and not used in the final recipe | Co-train with various VLM losses + FLARE (N1.5 only) world-model latent alignment auxiliary |
| VL co-training role | "Stabilizes language grounding and prevents catastrophic forgetting" | Co-training for VLM-side supervision; KI keeps stable | "Largest single-modality contribution; restores MMBench/GQA scores robot-only training degrades" | Web data via SigLIP pretrain; less explicit VL co-train at action time |
| Loss weights | SFT: Ξ»_vl = 0.1, Ξ»_act = 1.0 | Not disclosed; KI structurally separates the two | w_CE = 0.02 for CE on language; rest is flow matching | Not disclosed |
| RL stage | Yes β PPO + GAE on flow-matching log-prob via ODEβSDE, SimplerEnv-only | RECAP (advantage-conditioned, separate Ο*0.6 release) | No (pure supervised) | DreamGen-style synthetic data, no canonical RL stage in the public N1.7 recipe |
The biggest methodological points of comparison:
- Phase count. Qwen-VLA at four phases is the longest pipeline of the four (Ο is single-phase, LBM is three, GR00T is single-phase). The complexity buys two things: a language-conditioned action prior (T2A) and a task-success refinement (RL), neither of which the other three pursue together.
- Backbone freezing strategy. Qwen-VLA is the only one of the four that uses a time-varying freezing policy (frozen β unfrozen), and the only one that does it both for the warm-start and for the RL value head (the value head uses stop-gradient on VLM hidden states throughout RL). It is closer in spirit to LBM than to Ο / GR00T on this axis.
- No discrete action token target. Both Qwen-VLA and the LBM study's final recipe drop discrete action tokens; Ο series (KI) and Ο0-FAST use them. This is the cleanest 2026 alignment with LBM's empirical finding.
| Lab | Philosophy |
|---|---|
| PI | Internal-data-first; OXE and community data are background; cross-embodiment via shared continuous-action space + heterogeneous training |
| TRI LBM | OXE used in phase 1 and phase 2; drop OXE in phase 3 to specialize the target dual-Franka; cross-embodiment as warm-start, not as final spec |
| GR00T | Cross-embodiment throughout all stages; MultiEmbodimentActionEncoder + CategorySpecificMLP per-embodiment heads; data pyramid spans 7 ego-video datasets, sim, real teleop, neural trajectories |
| Qwen-VLA | Cross-embodiment throughout (11 embodiments + human MANO in CPT mixture, all kept through SFT, no platform-specialization phase). Embodiment-aware text prompt as the sole interface; zero-padding projection ablated to be on par with per-embodiment Multi-MLP |
Qwen-VLA is the most extreme "single weights, all embodiments, embodiment by prompt" philosophy of the four. Where LBM treats cross-embodiment as a phase-1 supplement and drops it for the target, Qwen-VLA explicitly aims for a generalist policy that serves all 11 platforms simultaneously β and its Table 4 demonstrates this by evaluating a single Qwen-VLA-Instruct on LIBERO + RoboCasa-GR1 + Simpler-WidowX + RoboTwin-Easy + RoboTwin-Hard without per-benchmark adaptation and comparing it head-to-head with specialist baselines that were fine-tuned per-benchmark. The single generalist beats the specialists on RoboCasa-GR1 (56.7 vs Ο0.5's 37.0), Simpler-WidowX (73.7 vs StarVLA-OFT's 64.6), and RoboTwin-Easy / Hard (86.1 / 87.2 vs ABot-M0's 86.0 / 85.0).
The most direct head-to-head numbers from Table 4:
| Benchmark | Ο0 | Ο0.5 | StarVLA-OFT | GR00T N1.6 | ABot-M0 | Being-H0.5 | Qwen-VLA-Base | Qwen-VLA-Instruct |
|---|---|---|---|---|---|---|---|---|
| LIBERO | 94.4 | 97.6 | 96.6 | 97.2 | 98.6 | 97.6 | 90.8 | 97.9 |
| RoboCasa-GR1 | β | 37.0 | 48.8 | 49.9 | 58.3 | 53.3 | 40.4 | 56.7 |
| Simpler-WidowX | β | 46.9 | 64.6 | 63.2 | β | β | 64.3 | 73.7 |
| RoboTwin-Easy | 65.9 | 82.7 | 50.4 | 47.6 | 86.0 | β | 64.3 | 86.1 |
| RoboTwin-Hard | 58.4 | 76.8 | β | β | 85.0 | β | 66.4 | 87.2 |
Qwen-VLA-Instruct does not beat ABot-M0 on LIBERO (98.6 vs 97.9) but matches it on RoboTwin and surpasses every model published as of May 2026 on RoboCasa-GR1 and Simpler-WidowX as a single generalist. The Base β Instruct delta (+7.1 LIBERO, +16.3 RoboCasa, +9.4 Simpler, +21.8 RoboTwin-Easy, +20.8 RoboTwin-Hard) is the empirical evidence that "instruction tuning yields substantial gains" on top of large-scale pretraining.
For real-world ALOHA (Tables 5β6 of the paper), the Qwen-VLA-aloha-with-pretrain variant outperforms Ο0.5 (Black et al. 2025) on the in-domain six-task average (83.6% vs 71.6%) and on every one of the five OOD axes, with the most striking gap being instruction generalization (84.6% vs Ο0.5's 42.3%) and background generalization (80.8% vs 26.9%). GR00T N1.6 trails far behind (28.6% in-domain, 25.4% OOD average), though it should be noted GR00T was not specifically fine-tuned for the ALOHA tasks β the comparison is "single model evaluation across the same suite", not a per-platform fairness adjustment.
Two additional empirical points worth flagging:
- SimplerEnv-OOD (Table 8). Qwen-VLA-Instruct 32.0% > Ο0.5 12.6% on average across 6 OOD task types where fine-tuning was restricted to simple Bridge pick-and-place. Ο0.5 fails completely on MoveRight (0.0%) and PlaceNear (0.0%) β i.e. its training distribution excludes the spatial-instruction parsing these tasks require β while Qwen-VLA-Instruct scores 33.3% and 39.6% respectively. The structural reading: Qwen-VLA's heavier VL co-training (8.5%) and spatial-grounding share (2.5%) carry spatial-language priors that translate to spatial-instruction generalization.
- DOMINO zero-shot (Table 9). Qwen-VLA-Instruct's 26.6% SR / 39.5 MS beats every fine-tuned baseline on a benchmark where moving-object dynamics are not part of the Qwen-VLA training data. The strongest fine-tuned baseline (PUMA, 17.2% SR / 35.0 MS) uses DOMINO-specific fine-tuning and temporal motion inputs; Qwen-VLA-Instruct uses neither and surpasses it by 9.4 pp SR. This is the most defensible claim that the joint pretraining produces generalizable spatial-to-kinematic priors, not just multi-task amortization.
The user-specific ask. In plain language:
Genuinely novel:
- The text-to-action (T2A) decompression warm-start. Freezing the VLM and pretraining the DiT on language + embodiment text alone β with images deliberately suppressed β is a new training-stage idea. The closest prior art is OpenVLA's initial fine-tune on FAST tokens or GR00T-N1's LAPA pseudo-action stage, but neither prohibits visual input. T2A's explicit thesis (the decoder should learn a structured language-indexed action prior before visual grounding) is original and the Β§5.2.1 ablations on Sigmoid-Normal vs. Beta timestep, full-sequence vs. chunk, and the rapid 2k-step convergence are unique contributions.
- Analytic flow-matching log-probability via ODEβSDE conversion for PPO. The standard trick of injecting per-step Gaussian noise during Euler integration exists in the literature (Song et al., 2021) but applying it to flow-matching action policies for PPO clipping β with a per-step random sub-sampling so only one extra DiT forward pass is needed during importance-ratio recomputation β is a clean RL contribution. This is a genuinely different solution to the "flow-matching policies don't have tractable log-probs" problem than ReinFlow's learnable-noise-injection or PI's RECAP (advantage-conditioned, avoids log-probs entirely).
- The RoboInf vision-synthesis pipeline (359,848 trajectories) + 7.2M-trajectory rendering-free language-only synthetic stage (the "over 8M synthetic" total is dominated by the language-only set, not RoboInf). Both pipelines are described in detail. The rendering-free language-only pipeline (six task templates Γ six robots Γ ~200k trajectories each at 50 Hz, via cuRobo motion planning, without physics simulation or scene rendering) is a deliberately cheap T2A-targeted data factory and a novel scale of language-action-only supervision.
- The 11-embodiment + human-MANO single-policy demonstration. No prior paper has shipped a single VLA evaluated head-to-head against 7+ specialist baselines on 4 distinct manipulation benchmarks plus navigation on R2R/RxR. GR00T N1.7 covers more humanoid embodiments but does not publish the corresponding sim-benchmark head-to-head; Ο0.6/Ο0.7 cover fewer platforms.
Recombination of known recipes:
- VL co-training to prevent VLM forgetting. This is the TRI LBM finding (Review-LBM-Cotraining) β robot-only training erodes MMBench/GQA; balanced VL co-training restores them. Qwen-VLA's Β§5.2.2 confirms VL+VLA matches VLA-only on LIBERO/Simpler-WidowX and beats it on RoboCasa-GR1 (+4.9 pp) and RoboTwin-2.0 (+4.6 pp). This is a known recipe applied; the contribution is the validation at Qwen-3.5 scale.
- Continuous flow-matching action expert. The architectural family is Ο0 / Ο0.6 / GR00T / LBM. The specific size (1.15B), block count (16), and AdaLN timestep injection are conventional choices.
- Embodiment-aware prompt conditioning. Specifying the platform via plain text is the GR00T series / Ο0.7 approach. The Qwen-VLA prompt template (robot tag + arm count + waist + mobile base + FPS + chunk size) is a particular instance, not a new idea.
- PPO + GAE + sparse binary reward for VLA RL. Standard recipe; the only novel piece is the flow-matching log-prob (see above).
Where Qwen-VLA aligns with PI's bets:
- Flow matching as the action objective (not discrete diffusion).
- Single-model cross-embodiment with text-based platform selection.
- VL co-training as a side channel (low mixture weight, throughout training) rather than as a primary signal.
- Distinct VL and action loss weights (Qwen-VLA's 0.1 / 1.0 echoes LBM's 0.02 / 1.0 in spirit β much smaller language weight to avoid VLM gradient domination).
Where Qwen-VLA aligns with TRI LBM's empirical findings:
- No discrete action tokens. LBM found FAST / VQ-VAE / latent video tokens do not help and sometimes hurt at LBM scale; Qwen-VLA simply does not include them.
- VL co-training helps and never interferes. LBM's largest single-modality gain finding is re-confirmed at smaller scale (8.5% combined VL share vs. LBM's much larger).
- The general framing of "co-training as anti-forgetting" rather than purely as generalization gain. Qwen-VLA's Β§2.5 explicitly says "this objective stabilizes language grounding under heavy embodied co-training and prevents catastrophic forgetting".
Where Qwen-VLA disagrees with either lab:
- vs. LBM on cross-embodiment phasing. LBM finds cross-embodiment data must be dropped in phase 3 to let the target platform specialize. Qwen-VLA keeps all 11 embodiments + human MANO through every stage including the SFT branches. The paper does not run the LBM-style "drop OXE in phase 3" ablation, so we cannot tell whether the LBM finding generalizes. But Qwen-VLA's Table 11 shows the final RL-refined model preserves performance on all benchmarks β so at minimum, keeping cross-embodiment in late stages does not collapse the target-task performance the way LBM warned it would on the dual-Franka platform.
- vs. PI on backbone freezing. PI / KI never freezes the backbone β KI's gradient-routing trick lets the backbone train continuously with stable CE-only gradients. Qwen-VLA's T2A does freeze the backbone, then unfreezes at CPT. This is a third-way that the LBM/KI camps had not explicitly considered.
-
vs. GR00T on cross-attention vs. concatenation. GR00T's whole architectural identity is cross-attention from DiT into VLM hidden states with an evolving preprocessing pipeline (raw β
vllnβvlln+vl_self_attention). Qwen-VLA picks concatenation + joint self-attention β no cross-attention, no preprocessing β and the Β§5.2.2 (b) ablation showing that a pretrained DiT outperforms a from-scratch one throughout SFT is the indirect evidence that the simpler routing suffices.
The Qwen-team-specific contribution:
The paper does not explicitly capitalize on Qwen3.5-specific features (gated linear attention, M-RoPE, QK-Norm) beyond inheriting them via the backbone. There is no ablation isolating "would this work as well with Qwen3-VL instead of Qwen3.5?" or "does the gated-linear-attention layer help vs. full GQA?". So the Qwen-developer-side advantage is latent in the backbone choice but is not the paper's headline contribution. The genuine team-specific contributions are (a) the data pipeline (RoboInf synthesis + language-only synthetic), (b) the 4-stage recipe with T2A warm-start, and (c) the analytic flow-matching log-prob for PPO.
Single generalist vs. specialists fine-tuned per-benchmark:
| Method | Type | LIBERO | RoboCasa-GR1 | Simpler-WidowX | RoboTwin-Easy | RoboTwin-Hard |
|---|---|---|---|---|---|---|
| Ο0 | Specialist | 94.4 | β | β | 65.9 | 58.4 |
| StarVLA-OFT | Specialist | 96.6 | 48.8 | 64.6 | 50.4 | β |
| GR00T N1.6 | Specialist | 97.2 | 49.9 | 63.2 | 47.6 | β |
| Ο0.5 | Specialist | 97.6 | 37.0 | 46.9 | 82.7 | 76.8 |
| ABot-M0 | Specialist | 98.6 | 58.3 | β | 86.0 | 85.0 |
| Being-H0.5 | Specialist | 97.6 | 53.3 | β | β | β |
| Qwen-VLA-Base | Generalist | 90.8 | 40.4 | 64.3 | 64.3 | 66.4 |
| Qwen-VLA-Instruct | Generalist | 97.9 | 56.7 | 73.7 | 86.1 | 87.2 |
| Model | Pick&Place | Table Cleaning | Bowl Stacking | Bowl Pick&Place | Towel Folding | Fine-grained | Avg. |
|---|---|---|---|---|---|---|---|
| GR00T N1.6 | 30.8 | 38.5 | 53.8 | 19.2 | 19.2 | 10.3 | 28.6 |
| Ο0.5 | 73.1 | 84.6 | 88.5 | 69.2 | 80.8 | 33.3 | 71.6 |
| Qwen-VLA-aloha (w/o pretrain) | 30.8 | 53.8 | 61.5 | 64.1 | 50.0 | 30.8 | 48.5 |
| Qwen-VLA-aloha (w/ pretrain) | 96.2 | 92.3 | 98.7 | 87.2 | 65.4 | 61.5 | 83.6 |
| Model | Color | Instance | Position | Background | Instruction | Avg. |
|---|---|---|---|---|---|---|
| GR00T N1.6 | 46.2 | 38.5 | 3.8 | 19.2 | 19.2 | 25.4 |
| Ο0.5 | 57.7 | 61.5 | 19.2 | 26.9 | 42.3 | 41.5 |
| Qwen-VLA-aloha (w/o pretrain) | 42.3 | 30.8 | 34.6 | 30.8 | 42.3 | 36.2 |
| Qwen-VLA-aloha (w/ pretrain) | 88.5 | 76.9 | 53.8 | 80.8 | 84.6 | 76.9 |
The +35.4 pp delta vs. Ο0.5 and +40.7 pp delta vs. the from-scratch ablation are the cleanest evidence in the paper that the Qwen-VLA-Base pretraining contributes β not just the architecture.
R2R Val-Unseen and RxR Val-Unseen:
| Method | R2R NEβ | R2R OSβ | R2R SRβ | R2R SPLβ | RxR NEβ | RxR SRβ | RxR SPLβ | RxR nDTWβ |
|---|---|---|---|---|---|---|---|---|
| NaVid | 5.7 | 49.2 | 41.9 | 36.5 | 5.7 | 45.7 | 38.2 | β |
| Uni-NaVid | 5.6 | 53.3 | 47.0 | 42.7 | 6.2 | 48.7 | 40.9 | β |
| NaVILA | 5.2 | 62.5 | 54.0 | 49.0 | 6.8 | 49.3 | 44.0 | 58.8 |
| StreamVLN | 5.0 | 64.2 | 56.9 | 51.9 | 6.2 | 52.9 | 46.0 | 61.9 |
| Qwen-VLA-Base | 5.2 | 61.7 | 53.8 | 49.4 | 6.4 | 55.1 | 45.8 | 56.2 |
| Qwen-VLA-Instruct | 5.1 | 69.0 | 57.5 | 51.2 | 5.8 | 59.6 | 47.8 | 57.1 |
Qwen-VLA-Instruct leads on most metrics; SPL and nDTW on the two splits are within 1 pp of the StreamVLN best.
All models fine-tuned only on the Bridge training split (simple pick-and-place with fixed object pairings); evaluation on unseen spatial-instruction tasks.
| Method | MoveAway | MoveRight | PlaceNear | PlaceRight | PutFront | StackYellow | Avg. |
|---|---|---|---|---|---|---|---|
| Ο0.5 | 26.1 | 0.0 | 0.0 | 32.1 | 13.0 | 4.2 | 12.6 |
| Qwen-VLA-Base | 31.3 | 31.6 | 16.7 | 47.1 | 6.3 | 18.8 | 25.3 |
| Qwen-VLA-Instruct | 43.8 | 33.3 | 39.6 | 47.9 | 4.2 | 22.9 | 32.0 |
| Method | SRβ | MSβ | Notes |
|---|---|---|---|
| OpenVLA (fine-tuned) | 1.5 | 6.1 | DOMINO fine-tune |
| RDT-1B (fine-tuned) | 5.3 | 17.7 | DOMINO fine-tune |
| Ο0 (fine-tuned) | 8.2 | 24.0 | DOMINO fine-tune |
| Ο0.5 (fine-tuned) | 9.6 | 26.2 | DOMINO fine-tune |
| InternVLA-M1 (fine-tuned) | 5.4 | 27.6 | DOMINO fine-tune |
| VLA-Adapter (fine-tuned) | 4.4 | 24.3 | DOMINO fine-tune |
| Ο0-FAST (fine-tuned) | 3.5 | 20.9 | DOMINO fine-tune |
| OpenVLA-OFT (fine-tuned) | 9.1 | 24.1 | DOMINO fine-tune |
| StarVLA-OFT (fine-tuned) | 10.9 | 30.5 | DOMINO fine-tune |
| PUMA (fine-tuned) | 17.2 | 35.0 | DOMINO fine-tune + temporal motion inputs |
| OpenVLA-OFT (zero-shot) | 6.7 | 20.0 | zero-shot |
| Ο0.5 (zero-shot) | 7.5 | 20.4 | zero-shot |
| LingBot-VLA w/ depth (zero-shot) | 11.8 | 26.7 | zero-shot |
| LingBot-VA (zero-shot) | 24.1 | 36.1 | zero-shot, WAM-style |
| Qwen-VLA-Base (zero-shot) | 21.1 | 37.4 | zero-shot, current frame only |
| Qwen-VLA-Instruct (zero-shot) | 26.6 | 39.5 | zero-shot, current frame only |
Qwen-VLA-Instruct is the best zero-shot model and the best overall model β including beating every fine-tuned baseline. This is the headline empirical result of the paper.
| Stage | Simpler | RoboCasa | RoboTwin-E | RoboTwin-H | LIBERO | Simpler-OOD | DOMINO SR | DOMINO MS |
|---|---|---|---|---|---|---|---|---|
| CPT (Base) | 64.3 | 40.4 | 64.3 | 66.4 | 90.8 | 25.3 | 21.1 | 37.4 |
| + SFT | 70.8 | 56.0 | 86.3 | 87.1 | 97.8 | 31.6 | 25.7 | 39.1 |
| + RL (Instruct) | 73.7 | 56.7 | 86.1 | 87.2 | 97.9 | 32.0 | 26.6 | 39.5 |
RL gives a +2.9 pp boost on SimplerEnv (where the rollouts come from) and 0.1β0.9 pp positive transfer to every other benchmark including DOMINO β i.e., no catastrophic forgetting from RL, even on held-out task families. The paper attributes this to (a) CPT having already exposed the model to mixed sim+real distributions and (b) RL optimizing task-success which generalizes across visual domains.
The paper publishes a substantial ablation set (Β§5.2.1β5.2.4). Summary:
Five design choices, all measured by SFT success rate on Simpler-WidowX:
| Choice | Best setting | Effect of going off-best |
|---|---|---|
| T2A stage itself (baseline added in v2) | With T2A | Without T2A: 60.9% vs 71.1% (β10.2 pp) |
| Data composition | 20% synthetic + 80% real | Pure real: β20 pp; pure synthetic: β7 pp |
| Sequence prediction mode | Full-sequence | Chunk: β4.9 pp at 10% synthetic |
| Visual input during T2A | None | Adding images: β2.9 pp |
| Flow-matching timestep distribution | Sigmoid-Normal at T2A, Beta at SFT | All other combinations: β5.7 to β11.7 pp |
| T2A training duration | 2,000 steps | 40k steps: β10.7 pp (overfitting) |
| Benchmark | VLA-only | VL + VLA | Delta |
|---|---|---|---|
| LIBERO | β same | β same | ~0 |
| Simpler-WidowX | β same | β same | ~0 |
| RoboCasa-GR1 | 51.1 | 56.0 | +4.9 |
| RoboTwin-2.0 | 81.8 | 86.4 | +4.6 |
VL co-training never hurts; helps on tasks requiring fine-grained object recognition and compositional instruction parsing. The pattern is consistent with TRI LBM's findings on object-recognition-heavy generalization.
The pretrained DiT (after T2A) outperforms a from-scratch DiT throughout SFT training on RoboCasa-GR1 β converges faster in early stages, reaches a higher peak. This is the direct validation of the T2A stage's structural value.
| Design | Bridge | Robocasa |
|---|---|---|
| Single-embodiment baseline | 62.8 / 53.4 | β |
| Multi-MLP | 63.3 | 52.1 |
| Concatenation | 63.0 | 52.8 |
| Zero-Padding | 63.0 | 53.2 |
All three within 1.2 pp; Zero-Padding chosen for parameter efficiency (2h Γ d_max vs. 2h Γ Ξ£d for the others).
| Conditioning | RoboTwin-Easy | RoboTwin-Hard |
|---|---|---|
| No state | 88.7 | 87.4 |
| State in VLM prompt (256-bin discretized) | 89.3 | 88.7 |
| State in DiT (continuous-valued) | 89.4 | 88.3 |
β€ +1.3 pp benefit from either injection; the multi-view visual observations + flow-matching predicting relative displacements together obviate explicit proprioception. Embodiment text prompt as the sole platform-specific interface is the final design choice. This is one of the cleaner "less is more" findings in 2026 VLA literature.
Already covered above. Key observation: RL transfers without catastrophic forgetting across all six in-distribution + OOD + dynamic benchmarks.
- Embodied action data remains far smaller and less diverse than VL pretraining data. Robustness to long-tail objects, environments, embodiments, and contact-rich interactions is limited.
- Joint training introduces optimization trade-offs. Action-oriented training "can modestly regress some pure vision-language and navigation evaluations" β the paper acknowledges this is a balancing problem they did not fully solve.
- Evaluations are largely short-horizon and benchmark-driven. Long-duration, failure-prone real-world deployment is an open challenge.
- No latency / inference-speed publication. Ο0.6 reports 63 ms / chunk on H100; GR00T N1 reports 63.9 ms / 16-action chunk on L40; LBM reports 0.146 s avg. Qwen-VLA gives no comparable number, which is significant given its concatenation-of-VLM-features-into-DiT-input pattern is a priori more expensive than prefix-KV or single-token-summary patterns.
- Hyperparameter reproducibility gap. Total pretraining steps, batch size, optimizer choice, total GPU-hours, peak LR for pretrain/SFT β none are stated as a single canonical number. The RL stage is well-specified; the supervised pretraining is under-documented.
- Weights and license status not addressed in the paper. The GitHub repo exists but the paper does not publicly commit to a license or weight release. Without weights, the public reproducibility of the 97.9% LIBERO, 73.7% Simpler-WidowX, and 26.6% DOMINO numbers depends on the team's future release decision.
- No head-to-head against Ο0.6/Ο0.7. The paper compares against Ο0.5 (April 2025) and GR00T N1.6 (Sep 2025), but not against the November 2025 Ο0.6 or April 2026 Ο0.7. Both are closed-weights so the omission is partly understandable, but the paper's claim that Qwen-VLA-Instruct is a leading 2026 VLA would land harder with at least the published Ο0.6 / Ο0.7 numbers.
- No discrete-diffusion VLA baseline. Discrete Diffusion VLA hits 96.3% on LIBERO; Qwen-VLA-Instruct hits 97.9%. The two architectural families (Cat B vs. Cat D) are not directly compared. A LIBERO head-to-head with a single shared eval protocol would settle a long-running debate.
- The "single generalist beats specialists" framing has caveats. Several "specialist" baselines (e.g., ABot-M0 on LIBERO 98.6 vs Qwen-VLA's 97.9) actually beat Qwen-VLA. The fairer claim is "matches specialists on the easiest benchmarks, beats them on the harder ones".
- VL co-training ablation is limited. Β§5.2.2 shows VL+VLA matches VLA-only on Libero/Simpler and helps on RoboCasa/RoboTwin, but does not run the LBM-style "what fraction of VL data?" sweep that would settle the 8.5% choice. It also does not measure VLM-benchmark preservation (MMBench/GQA) the way LBM does β so the "prevents catastrophic forgetting" claim in Β§2.5 is asserted, not measured.
- T2A ablations are run only on Simpler-WidowX. The 71.1% peak is a single-benchmark result. Whether the 20%-synthetic-80%-real, Sigmoid-Normal-then-Beta, full-sequence, 2k-step optimum is the optimum for all embodiments and benchmarks is not verified β it is presented as a global recipe.
- The 8.5% combined VL share is much lower than LBM's recipe. LBM finds VL co-training is the largest single-modality contribution. Qwen-VLA's 8.5% is consistent with Ο series posture, but the LBM finding suggests there may be unrealized gain available from increasing it. The paper does not run this sweep.
- No published failure-mode analysis. Real-world experiments are short-horizon (β€ Table 5's six categories). Long-horizon deployment with recovery, replanning, and observation-of-failure is acknowledged as future work, not measured.
The May 2026 release positions Qwen-VLA as:
- First Qwen team VLA. Backbone-first VLM developers shipping their own VLA is a 2025β2026 trend (Google β Gemini Robotics; AllenAI β MolmoAct; Alibaba β Qwen-VLA). The technical contributions here are the four-stage recipe and the unified action representation; the strategic contribution is making Qwen a credible end-to-end VLA developer rather than just a backbone vendor.
- The most-embodiment-diverse single-policy demonstration of 2026. 11 robot platforms + human MANO + navigation + autonomous driving VQA + spatial grounding all routed through one model. GR00T N1.7 covers more humanoid embodiments specifically; Ο0.7 covers fewer but with deeper post-training. No prior paper has demonstrated this breadth at this benchmark density.
- The strongest "joint pretraining as transferable prior" claim of 2026. The 26.6% DOMINO zero-shot result β beating every fine-tuned baseline including PUMA β is the most defensible empirical result. It would have been even stronger had the paper published an ablation showing which pretraining sources matter for the DOMINO transfer.
- Aligned with TRI LBM's empirical findings on three axes (discrete action tokens not needed; VL co-training helps and doesn't interfere; the action expert benefits from a structured warm-start).
- A third-way between PI's KI and TRI LBM's frozen backbone for the VLM training policy: freeze for warm-start, unfreeze for joint training, stop-gradient for the RL value head.
10.2 Does it change the "which architecture should I pick?" calculus from Review-VLA-Architecture?
Not radically, but with three refinements:
- Category B (flow-matching expert) gets a new wiring pattern. Concatenation + joint self-attention is a viable third option alongside same-stack MoE (Ο) and adaLN-conditioning (LBM). Whether it has different latency / scaling properties is not yet measured publicly.
- The T2A warm-start is a new training-recipe primitive. It is generalizable in principle β any backbone-frozen, language-only DiT pretrain could be applied to any Cat B VLA. Whether it transfers to Ο0.6 / LBM-class systems remains to be seen.
- The ODEβSDE flow-matching log-prob is a new RL primitive. Applicable to any flow-matching VLA. This may be the most reusable contribution of the paper for the RL community.
Where Qwen-VLA does not change the calculus:
- It does not displace Ο series on dexterous long-horizon real-world tasks (the laundry / box assembly tier β the paper's real-world tasks are shorter-horizon).
- It does not displace GR00T on humanoid breadth (no Atlas, no Optimus, no Figure).
- It does not address the discrete-diffusion-vs-flow-matching debate; it is firmly in the flow-matching camp.
- Does T2A scale beyond LBM/Ο/GR00T scale? The 2k-step convergence finding is on a small Simpler-WidowX eval. Whether the language-only prior helps at 100k-step pretraining is unverified.
- Does the analytic flow-matching log-prob trick scale to long-horizon dexterous tasks beyond SimplerEnv? The RL stage was deliberately narrow ("a single simulation environment"); it has not yet been validated on a Ο*0.6-style laundry or box-assembly task.
- Does the concatenation-routing scale to 50-step chunks (Ο-series scale) or beyond 16 DiT blocks?
- Does increasing the VL co-training share beyond 8.5% β toward LBM's much larger share β help further?
- arXiv (main paper): https://arxiv.org/abs/2605.30280
- PDF: https://arxiv.org/pdf/2605.30280
- Blog (Qwen team): https://qwen.ai/blog?id=qwenvla
- Code: https://github.com/QwenLM/Qwen-VLA
- Qwen3-VL Technical Report (companion backbone documentation): https://arxiv.org/abs/2511.21631
- Qwen Team's VLA Program β cross-paper review placing this paper in the VLM4VLA β Qwen-VLA β Qwen-Robot Suite arc
- Qwen-RobotManip β the Qwen team's second VLA (Jun 2026): alignment-first thesis, camera-frame delta EEF, cross-attention expert β its Table 19 ablation finds last-layer cross-attention beats this paper's concatenation wiring
- Ο series evolution
- Ο0.6 long-form Β· Ο0.7 long-form
- LBM Co-training Study
- GR00T series
- Knowledge Insulation
- VLA Architectures
- VLMβAction Connection
- VLA Attention Architectures
- WAM vs VLA Robustness
- Cross-Embodiment Training
- ML Foundations Β· Attention (for the Qwen3.5 gated-linear + GQA discussion)
- ReinFlow (alternate flow-matching RL recipe)
- Ο*0.6 + RECAP (alternate flow-matching RL recipe β advantage-conditioned)
β Back to Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)