-
Notifications
You must be signed in to change notification settings - Fork 0
Review VLM Action Connection
Compiled April 2026 Β· Focus: the specific architectural question "how does the VLM condition the action generator?" β KV sharing, cross-attention, latent tokens, FiLM, shared parameters, async scheduling, and more.
This is the sibling review to Review-VLA-Architecture (which groups VLAs by action-decoder family) β this page slices the same papers along a different axis: where and how information flows from the VLM into the action generator. Companion reviews: RL for VLA Β· VLA Memory Β· Cross-Embodiment.
Everyone talks about "VLM + action expert," but the coupling mechanism varies wildly across 2024β2026 VLAs. This review isolates seven distinct interface patterns, with concrete mechanism details:
- Unified token stream (OpenVLA, VLA-0, discrete-diffusion VLAs) β no interface; actions are tokens in the VLM's own vocabulary
- Same-stack MoE + prefix-KV attention (Ο-series) β action expert is separate weights in the same transformer stack; action tokens attend into the VLM's cached prefix KV
- Cross-attention into VLM hidden states (GR00T N1, RDT-1B, ST4VLA) β separate action transformer cross-attends to VLM features, possibly from a specific intermediate layer
- Latent condition token (ThinkAct, RoboDual, WholeBodyVLA) β VLM emits a small summary; action head consumes only that
- FiLM / prefix conditioning (CogVLA, classic Diffusion Policy) β features modulate the action stream via Ξ³/Ξ² affine, without cross-attention
- Parameter sharing at layer boundary (Fast-in-Slow) β S1 is the last 2 transformer blocks of the VLM, re-run at higher frequency
- Outside the VLM (RFS residual, VITA-VLA distillation, RTC async scheduling) β interface unchanged; a wrapper at the action level adapts behavior
Key finding: no one has published a matched-compute head-to-head of these mechanisms under the same backbone and data. The field is bifurcating β Ο-series on same-stack MoE, ICLR 2026 on unified streams, NeurIPS 2025 dual-system on latent/embedded β without a clean comparison. That ablation is the most valuable unpublished paper in the space.
Independent of the action generator's objective (flow matching vs. diffusion vs. cross-entropy), the interface choice affects:
| Axis | What the interface controls |
|---|---|
| Latency | KV caching + layer reuse (same-stack / embedded) β fastest; cross-attn β extra compute; unified streams β depends on decoding strategy (AR slow, masked diffusion fast) |
| Gradient hygiene | Does the action-head gradient corrupt VLM features? Knowledge Insulation says "stop the backflow" |
| Pretraining preservation | Unified streams (VLA-0 extreme) and embedded sharing (FiS) inherit VLM pretraining directly; cascaded latent (ThinkAct) bottlenecks it; distillation (VITA-VLA) explicitly tries to transfer it |
| Backbone compatibility | Same-stack MoE needs matched head-dim + layer count; embedded needs knowable block structure; cross-attn is most portable |
| Bandwidth | Unified stream = high; cross-attn = medium (features, not just a vector); latent = low (one bottleneck) |
| Which VLM layer? | Unified: every layer, inherently. Cross-attn: you must choose β GR00T picks layer 12 of Eagle-2 empirically, ST4VLA queries k intermediate layers. |
| Frequency decoupling | Orthogonal to most choices. Embedded (FiS 1:4), cascaded latent (GR00T ~10:100 Hz), and async scheduling (RTC) each achieve it differently. |
flowchart TB
Q{Where does VLM information enter the action path?}
Q --> I1[1. Unified token stream<br/>same transformer, same vocab]
Q --> I2[2. Same-stack MoE + prefix KV<br/>shared layers, separate weights]
Q --> I3[3. Cross-attention into VLM hidden states<br/>separate action transformer]
Q --> I4[4. Latent condition token<br/>bottlenecked summary vector]
Q --> I5[5. FiLM / prefix conditioning<br/>Ξ³Β·Ξ² modulation]
Q --> I6[6. Parameter sharing at layer boundary<br/>embedded S1-in-S2]
Q --> I7[7. Outside the VLM<br/>residual / distilled / async]
Core idea. Actions are tokens in the VLM's own vocabulary or token space. One transformer, one loss, no "interface."
Exemplars and sub-patterns:
| Paper | Sub-pattern | How actions enter the stream |
|---|---|---|
| OpenVLA (CoRL 2024) | AR discrete tokens | 256 bins in extra vocab, AR next-token prediction via LLaMA-2 decoder |
| RT-2 / RT-X / Ο0-FAST | AR with custom tokenizer | Same stream, FAST/DCT tokenizer compresses actions to fewer tokens |
| VLA-0 (2510.13054) | Actions as plain text | No new tokens at all β actions encoded as decimal numerals; Qwen2.5-VL-3B backbone untouched (no action head/expert). 94.7% avg LIBERO β beats all same-data methods (Ο0.5-KI, OpenVLA-OFT, SmolVLA) and, without large-scale robot pretraining, also beats large-data methods (Ο0, GR00T-N1, MolmoAct) |
| Discrete Diffusion VLA (2508.20072) | Masked discrete diffusion | Action tokens in VLM vocab; parallel unmasking with adaptive order. 96.3% LIBERO |
| Unified Diffusion VLA | Joint frame+action diffusion | Block-wise causal mask; action tokens attend to still-being-denoised future-image tokens |
| dVLA | Multimodal CoT in one stream | Discrete diffusion over future frames + text CoT + actions, all in parallel |
| HybridVLA (2503.10631) | Interleaved AR + diffusion | Single LLM; diffusion denoising interleaved into next-token prediction |
- β Zero architectural surgery Β· inherits LLM training + serving tooling Β· preserves VLM knowledge most directly Β· no inter-module interface to hand-design
- β AR decoding is sequential (slow) unless you use parallel decoding (OFT) or masked diffusion Β· discretization caps precision on continuous control Β· mixing text CoT + action in one stream can compete for attention capacity
- Pick when: simplest possible recipe, need interpretable unified stream, or want VLA-0-style minimal modification
Core idea. The action expert is a separate set of weights, but lives in the same transformer stack as the VLM. Action tokens attend to the VLM's prefix using shared attention-head geometry, and the VLM's prefix KV is cached at inference β only action tokens get recomputed per step.
Key constraints (why Ο0.6 and Ο0.7 have "860M expert, same layer count as backbone"):
- Matched attention-head dim and layer count between VLM and expert β required so the attention operation is compatible
- KV caching of the VLM prefix β production latency (Ο0.6: 63 ms / chunk on H100 with 5 Euler steps)
- Bidirectional attention among action tokens, but action tokens attend the VLM prefix through the standard KV interface
Exemplars:
| Paper | Expert size | Details |
|---|---|---|
| Ο0 (2410.24164) | Smaller hidden/MLP, matched heads | Canonical Ο-style MoE action expert; the template every Ο-series release inherits |
| Ο0.5 | Same | + autoregressive subtask text as prompt prefix; VLM re-forwards after subtask emission |
| Ο0.6 | 860M, same layer count | Backbone upgrade to Gemma3-4B; interface identical |
| Ο0.7 | Same 860M | Interface unchanged; all deltas are prompt-side (subgoal images, metadata, CFG) |
| FLOWER (2509.04996) | ~950M | Efficient variant; prunes 50% LLM layers + Global-AdaLN; 200 H100-hr pretraining |
| Knowledge Insulation | Same interface | Gradient modifier, not a new interface. Blocks flow-matching gradient from corrupting VLM weights |
- β Smooth continuous actions Β· 5-step inference Β· cached prefix KV β fast Β· clean gradient hygiene with KI Β· production-deployed at PI scale
- β Requires backbone-compatible architecture (head-dim, layer count) Β· action-head gradient corrupts VLM without Knowledge Insulation Β· more parameters than pure unified stream
- Pick when: production latency matters, continuous control is primary, you can design expert + backbone together
Core idea. Action transformer is a separate network. Cross-attention layers attend to VLM token embeddings, often from a specific intermediate VLM layer rather than the last.
Exemplars with the non-obvious details:
| Paper | Where cross-attn attends | Non-obvious detail |
|---|---|---|
| GR00T N1 / N1.5 / N1.6 (2503.14734) | Layer 12 of Eagle-2 (intermediate, not last) | Chosen empirically for speed + success; alternating self-attn + cross-attn blocks Flamingo/VIMA-style |
| RDT-1B (2410.07864) | Frozen SigLIP + frozen T5-XXL | "Alternating Condition Injection" β image tokens and text tokens are cross-attended in alternating layers, not both every layer, because image would drown text. Explicitly rejects AdaLN because conditions are "high-dim and variable length" |
| ST4VLA (ICLR 2026) | k intermediate VLM layers | Query transformer cross-attends to k intermediate layers (not just the last) to stabilize expert learning |
| RetoVLA | + register-token KV injection | Injects discarded register tokens as auxiliary KV pairs β cheap global spatial context |
| DexVLA (2502.05855) | Plug-in ~1B diffusion expert | Conditions on VLM features; exact pathway less crisply specified in the paper |
| Cosmos Policy | Cosmos video backbone + control tokens | VLM (video backbone) features feed control-token decoders |
- β Flexible β VLM and expert can have different sizes Β· can freeze VLM, train expert Β· can pick which layer's features matter (ST4VLA ablation)
- β Extra cross-attention compute Β· interface layer choice is ad-hoc (GR00T's layer-12 is empirical) Β· image tokens can drown text (RDT's alternating trick exists to fix this)
- Pick when: mixing vendor VLMs with custom experts, frozen-VLM setups, or you want to pluggably swap action transformers
Core idea. VLM emits a small fixed-size summary token/vector; the action head consumes only that.
Exemplars:
| Paper | What the latent is |
|---|---|
| ThinkAct (NVIDIA, 2507.16815) | RL-rewarded MLLM plan compressed into a visual latent that conditions a separate action head |
| RoboDual (2410.08001) | Generalist VLA emits latent; specialist DiT conditions on it Β· +26.7% real vs OpenVLA Β· specialist only 20M params |
| WholeBodyVLA | Unified latent decodes to coordinated base / arms / hands |
| Hi-Robot | Hierarchical planner latent feeds low-level controller |
- β Sharp frequency decoupling (S2 ~10 Hz, S1 ~100 Hz) Β· minimal bandwidth between systems Β· interface is small and easily cached Β· training curricula can be separated
- β Bottleneck loses information Β· hand-designed latent shape Β· poor fine-grained visual grounding for contact tasks Β· separate training cadence
- Pick when: long-horizon planning with slow S2 + fast S1; plan caching matters; modular development with separate teams
Core idea. Instruction / plan features modulate (Ξ³, Ξ² affine) the action stream at multiple layers. No explicit cross-attention.
Exemplars:
| Paper | Where FiLM is applied |
|---|---|
| CogVLA (2508.21046) | Twice β EFA-Routing applies FiLM at the vision encoder for token aggregation; LFP-Routing applies FiLM at the LLM for token pruning. Plus V-L-A Coupled Attention (causal V-L + bidirectional action parallel decoding). 97.4% LIBERO, 2.5Γ training / 2.8Γ inference speedup over OpenVLA |
| Classic Diffusion Policy | FiLM conditions the U-Net at every block |
| RDT-1B | Considers and rejects AdaLN β "lossy for high-dim variable-length conditions" |
- β Parameter-efficient Β· no explicit cross-attn layers Β· composes beautifully with token pruning (CogVLA's 2.8Γ speedup) Β· classical robustness
- β Lossy for long / variable-length conditions Β· weaker than cross-attention on complex prompts (empirical: RDT chose cross-attn over AdaLN for this reason)
- Pick when: small efficient VLAs, instruction-conditioned vision pruning, or the classical diffusion-policy recipe
Core idea. Action head and VLM share some transformer blocks; S1 is a subset of S2's layers re-run at higher frequency.
Exemplar:
-
Fast-in-Slow (2506.01953) β the canonical example, and currently the only one.
- Last 2 of 32 LLM blocks = S1 (Prismatic-VLM backbone: SigLIP+DINOv2 vision + LLaMA-2-7B LLM; ablated optimum β performance saturates at 2 of 32 shared blocks)
- S2 = full 32 blocks at 1/4 the frequency of S1 (1:4 ratio, ablated optimum)
- S1 sees extra modalities S2 doesn't: 3D point clouds (lightweight tokenizer + shared encoder), robot state, noised actions
- 117.7 Hz control on NVIDIA 4090 with chunk=8
- See Review-Fast-in-Slow for the full deep-dive
-
β S1 inherits VLM pretraining for free (shared weights) Β· single set of weights to maintain Β· naturally handles frequency decoupling Β· production-rate control
-
β Block count is a hyperparameter (2 optimal for LLaVA; untested on other backbones) Β· still passes a latent forward at S2βS1 boundary Β· backbone-specific
-
Pick when: you want dual-system benefits without maintaining two networks; dense (non-MoE) backbones; your backbone has a consistent block structure
Core idea. Don't modify the VLMβaction link. Add a residual adapter, distill a teacher into the VLM, or change the scheduling around inference.
Exemplars (three different patterns):
| Paper | What it adds | Mechanism |
|---|---|---|
| RFS (Residual Flow Steering) | Residual adapter | Base flow-matching policy frozen; small residual steering policy trained with RL; outputs summed with base flow field at the action-vector level |
| VITA-VLA (2510.09607) | Teacher-student distillation (reverse direction) | Distills a small pretrained action model INTO a 7B VLM via hidden-state alignment. Two stages: (1) alignment β map VLM hiddens to teacher's action space; (2) fine-tune. 97.3% LIBERO, 82.0% real. Teacher's decoder is reused |
| Real-Time Chunking (2506.07339) | Async scheduling | Plug-and-play on any diffusion/flow VLA, no retraining. Async chunk inpainting: freeze committed actions, inpaint the rest while previous chunk executes |
| PLD | Residual RL + distill | Frozen base, residual policy trained with RL, then distilled back into the base |
- β Plug-and-play on frozen base Β· preserves base VLA knowledge Β· sidesteps the interface question entirely Β· scales to new behaviors without touching the backbone
- β Inherits the base's ceiling Β· doesn't fix a poor VLMβaction coupling Β· works only if the base already solves the core problem
- Pick when: fine-tuning to new embodiments / tasks; contact-rich residual correction; serving an existing production VLA without retraining
The ICRA 2026 VLA cohort is deployment- and sensor-centric, and that shifts where it pushes on the interface. It contributes no new coupling primitive, but it stress-tests three of the seven along axes the ML venues ignore β chiefly what extra modality enters the VLM, and how to upgrade a frozen interface from reward.
A "sensor-token into the VLM prefix" variant of mechanism 1/5. FD-VLA is the cleanest example: a Force Distillation Module fuses a learnable query token over vision + robot state into a predicted force token that is injected into the pretrained VLM, distilled at training time against the latent of a real force/torque signal so the sensor is needed only during training. The injected token rides the VLM's own stream (mechanism 1) rather than cross-attending from outside β but unlike VLA-0's text tokens it carries a non-linguistic sensor channel, and the design is explicitly engineered to preserve the VLM's vision-language semantics (the force token augments, does not disrupt, the prefix). The same theme appears as FiLM in the cohort β Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration modulates the visual stream with F/T features (mechanism 5) instead of adding a prefix token. Together they show the interface question now includes which sensor modality gets wired in, and via which of the seven slots.
Mechanism 4, in production dual-system form. Galaxea / G0 is a textbook latent/subtask cascade: a System-2 G0-VLM (Qwen2.5-VL) emits a high-level subtask goal that conditions a separate System-1 G0-VLA (PaliGemma-3B + SigLIP, FAST tokenizer + flow matching). The hand-off is a low-bandwidth subtask interface between two independently-trained networks β the canonical mechanism-4 frequency-decoupling trade-off β and G0 is a real-robot, 23-DoF mobile-bimanual instantiation of it rather than a benchmark study.
Mechanism 7, sharpened for flow heads. ICRA's RL cluster targets the "outside-the-VLM" slot. FPO is a drop-in online-RL recipe that needs no architectural change: it builds a likelihood-free policy ratio from per-sample changes in the conditional-flow-matching loss the policy is already trained with, leaving the VLMβaction coupling (here Ο0's same-stack MoE) untouched. Like RFS and RTC, it adapts behavior around a frozen interface β extending the mechanism-7 pattern to reward-driven fine-tuning of flow-matching experts specifically. See the ICRA VLA topic page for the full cohort.
These don't define new interface patterns β they optimize existing ones:
| Paper | Contribution |
|---|---|
| VLA-Cache (2502.02175) | Reuse KV cache for visual tokens that don't change step-to-step |
| KV-Efficient VLA | Compresses historical KV cache via recurrent gating |
| VLA-Adapter (2509.09372) | "Bridge Attention" β learned selector that autonomously picks which VLM conditions to inject. 0.5B backbone, trains in 8 hr on 1 consumer GPU. Closest paper to an interface-ablation |
| VLA-OS (2506.17561) | Controlled paradigm study: Hierarchical > Integrated > Action-Only Β· visual-grounded > language-grounded planning |
| ST4VLA (ICLR 2026) | Querying transformer over k intermediate VLM layers (layer-selection ablation) |
| Axis | Winner | Runner-up |
|---|---|---|
| Lowest latency at production scale | 2. Same-stack MoE (Ο0.6 @ 63 ms / chunk) | 6. Embedded (FiS @ 117.7 Hz control) |
| Best pretraining preservation | 1. Unified stream VLA-0 (zero modification, 94.7% LIBERO) | 2. Same-stack + KI |
| Highest bandwidth VLMβaction | 3. Cross-attention | 1. Unified stream |
| Sharpest frequency decoupling | 4. Latent condition (10 Hz S2 / 100 Hz S1) | 6. Embedded (1:4 ratio) |
| Most portable across VLM sizes | 3. Cross-attention (VLM + expert can differ) | 7. Outside-the-VLM (plug-and-play) |
| Most parameter-efficient | 5. FiLM (CogVLA 2.8Γ inference speedup) | 6. Embedded |
| Best interpretability | 4. Latent condition (inspectable summary) | 1. Unified stream (text CoT visible) |
| Cheapest to adapt to new task | 7. Outside-the-VLM (RFS/PLD residual) | 5. FiLM adapters |
| Single unifying objective | 1. Unified stream | β |
| Empirical best on LIBERO | 1. Discrete Diffusion VLA (96.3%) | 1. VLA-0 (94.7%, actions as text) beats Ο0.5-KI/OpenVLA-OFT/SmolVLA (same data) and Ο0/GR00T-N1/MolmoAct (large data) |
-
The Ο-style same-stack MoE + prefix-KV has become the production default for continuous control (Ο0βΟ0.5βΟ0.6βΟ0.7, FLOWER). Knowledge Insulation formalized the gradient hygiene that makes it stable. But it requires backbone-compatible architecture (head-dim, layer count), which locks you into specific VLM sizes.
-
ICLR 2026 pulled the interface back into the VLM via discrete diffusion (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA). This collapses mechanisms 2 + 3 back to mechanism 1 by making the VLM itself generate actions through masked parallel denoising.
-
VLA-0's counter-swing: the simplest mechanism (actions as text, zero architectural change) hits 94.7% avg LIBERO β beating same-data methods (Ο0.5-KI, OpenVLA-OFT, SmolVLA) and, with no large-scale robot pretraining, large-data methods (Ο0, GR00T-N1, MolmoAct). Suggests the "right" interface may be to not have one.
-
NeurIPS 2025 diversified dual-system interfaces into 3 distinct patterns at one conference: cascaded latent (ThinkAct), MoE-routed (ChatVLA-2), embedded-shared (Fast-in-Slow). No consensus.
-
Frequency decoupling is now orthogonal to the interface choice. Real-Time Chunking works on any diffusion/flow VLA, no retraining. Fast-in-Slow's 1:4 ratio is a separate axis. You can pick any interface and layer async scheduling on top.
-
The field is NOT converging β it's bifurcating:
- Production continuous control β mechanism 2 (Ο-series) + gradient insulation
- Research / unified objective / interpretability β mechanism 1 (discrete diffusion, VLA-0)
- Dual-system / long-horizon β mechanisms 4 or 6
-
By ICRA 2026 the mechanism set has stabilized β the frontier moved from inventing wirings to deploying them. The Vienna cohort adds no eighth coupling primitive: its contributions slot into the existing seven and stress-test them on deployment/sensor/RL axes the ML venues under-explore β a sensor token distilled into the VLM prefix (FD-VLA, a variant of mechanism 1/5), a production S2βS1 subtask hand-off (Galaxea G0, mechanism 4 at 23-DoF), and architecture-free RL around a frozen flow interface (FPO, mechanism 7). The interesting question is no longer "which wiring?" but "how cheaply can I adapt a frozen one at deploy time?"
-
Matched-compute head-to-head of KV-sharing vs. cross-attention vs. latent-token vs. FiLM on the same backbone, same data, same action decoder. VLA-OS compared paradigms (Hierarchical vs. Integrated vs. Action-Only) but not interfaces. VLA-Adapter ablates which conditions matter, not how to inject them. This is the single most useful unpublished study.
-
Which VLM layer should cross-attention target? GR00T's layer 12 of Eagle-2 was chosen empirically for speed + success. ST4VLA uses k intermediate layers. No published layer-sweep on a standard benchmark.
-
Does KV-prefix sharing (Ο-style) actually preserve VLM knowledge better than cross-attention? The motivating claim for KI is "action gradients corrupt VLM features" β but that's about gradient flow, not attention flow. Does the same failure mode occur in GR00T-style cross-attention? Unknown.
-
Is the embedded pattern (Fast-in-Slow) an artifact of LLaVA's 32-block depth? "2 blocks optimal" on LLaVA. Completely untested on Gemma3-4B (Ο-series), Qwen-VL, Eagle-2. If the optimum varies, the design isn't transferable.
-
Does teacherβstudent distillation (VITA-VLA) preserve VLM reasoning better than joint training + KI? No head-to-head. VITA-VLA reports 97.3% LIBERO but doesn't run open-world reasoning preservation metrics Γ la ChatVLA-2 / VLM4VLA.
-
Is the interface question obsolete under unified token streams? If Discrete Diffusion VLA / VLA-0 close the performance gap on continuous control, mechanisms 2β5 may be legacy. Conversely, if flow matching remains Pareto-optimal on latency, unified-stream approaches need to match 63 ms on H100.
-
How does the interface interact with cross-embodiment transfer? Soft prompts (X-VLA), MoE (HiMoE-VLA), shared codebooks (XR-1) place heterogeneity at different interface sites. No paper ablates interface Γ embodiment heterogeneity jointly.
If you're designing a VLA and wondering which interface to pick:
- Shipping continuous control at <100 ms/chunk? β Mechanism 2 (same-stack MoE + prefix KV) with Knowledge Insulation. Matched head-dim and layer count between VLM and expert. This is Ο0.6 / Ο0.7.
- Want maximum VLM-pretraining preservation + simplest recipe? β Mechanism 1 (unified stream). Try VLA-0 (actions as text) before anything else. If you need parallelism, Discrete Diffusion VLA.
- Mixing a frozen vendor VLM with a custom action network? β Mechanism 3 (cross-attention). Pick an intermediate layer to attend to (following GR00T / ST4VLA).
- Long-horizon + sharp frequency decoupling? β Mechanism 4 (latent cascade) or mechanism 6 (embedded sharing). Embedded saves parameters; cascaded caches plans.
- Small / efficient / CPU edge? β Mechanism 5 (FiLM, CogVLA-style). Or mechanism 1 with a small backbone (SmolVLA).
- Adapting an existing production VLA without retraining? β Mechanism 7. RFS for RL-residuals, RTC for latency, VITA-VLA for distillation, MAP-VLA-style prompt-library retrieval for new tasks.
- Don't stack mechanisms thoughtlessly. Async scheduling (RTC) composes with any interface. FiLM and cross-attention can coexist. But same-stack MoE + cross-attention in the same model doesn't make sense β you'd have two interfaces to the same VLM.
Papers with primary interface innovation:
- Ο0 β https://arxiv.org/abs/2410.24164
- Ο0.5 β https://arxiv.org/abs/2504.16054
- Ο0.6 model card β https://website.pi-asset.com/pi06star/PI06_model_card.pdf
- Ο0.7 β https://www.pi.website/blog/pi07 Β· https://www.pi.website/download/pi07.pdf
- Knowledge Insulation β https://arxiv.org/abs/2505.23705
- RDT-1B β https://arxiv.org/abs/2410.07864
- DexVLA β https://arxiv.org/abs/2502.05855
- GR00T N1 β https://arxiv.org/abs/2503.14734
- Fast-in-Slow β https://arxiv.org/abs/2506.01953 Β· Review-Fast-in-Slow
- ChatVLA-2 β https://arxiv.org/abs/2505.21906
- ThinkAct β https://arxiv.org/abs/2507.16815
- Real-Time Chunking β https://arxiv.org/abs/2506.07339
- CogVLA β https://arxiv.org/abs/2508.21046
- VITA-VLA β https://arxiv.org/abs/2510.09607
- VLA-0 β https://arxiv.org/abs/2510.13054
- OpenVLA β https://arxiv.org/abs/2406.09246
- Discrete Diffusion VLA β https://arxiv.org/abs/2508.20072
- Unified Diffusion VLA β https://openreview.net/forum?id=UvQOcw2oCD
- VLA-Adapter (Bridge Attention) β https://arxiv.org/abs/2509.09372
- VLA-OS β https://arxiv.org/abs/2506.17561
- RoboDual β https://arxiv.org/abs/2410.08001
- HybridVLA β https://arxiv.org/abs/2503.10631
- VLA-Cache β https://arxiv.org/abs/2502.02175
Companion reviews:
- Review: VLA Architectures β same papers sliced by action-decoder family
- Review: VLA Memory Β· Cross-Embodiment Β· RL for VLA
- Per-paper: Review-pi07 Β· Review-pi06 Β· Review-VLM4VLA Β· Review-Fast-in-Slow
Verdict: first direct ablations favor last-layer cross-attention, but margins are ~0.5 pp β the wiring axis is live, not settled.
2025 ended in a four-way standoff, one pattern per lab (Ο's same-stack MoE + prefix-KV Β· GR00T's cross-attention Β· LBM's adaLN token Β· Qwen-VLA's concatenation). H1 2026 produced the first direct evidence: Qwen-RobotManip Table 19 (cross-attention 87.5 > concatenation 87.0 > layer-wise fusion 86.4 on LIBERO-Plus, at lowest cost) and Ξ¨β's hardware ablation (SD3-style MM-DiT dual modulation + joint attention > naive DiT head).
| Pattern | Champion | Pros | Cons | 2026 evidence |
|---|---|---|---|---|
| Same-stack MoE + prefix-KV | Ο0.6/0.7 | KV reuse, tight coupling | Backbone surgery, hard to retrofit | Production-proven; never directly ablated vs others |
| Cross-attention (last layer) | GR00T, RobotManip | Decoupled; ~500M experts suffice | One-way information flow | Wins both 2026 ablations |
| Concatenation + joint self-attn | Qwen-VLA | Simplest | Re-attends full VLM state per denoise step | Loses narrowly in its sibling's ablation |
| adaLN single token | TRI LBM | Cheapest | Information bottleneck | No 2026 head-to-head |
| MM-DiT dual-stream | Ξ¨β | Timestep modulates VL & action branches separately | Newest, least replicated | Beats naive DiT on a real humanoid |
| Mixture-of-Transformers dual-system | HALO, LaSTβ (ICML 2026) | Separates low-frequency reasoning from high-frequency action experts in one model | Two-rate coherence | HALO +34.1% over Ο0 on RoboTwin |
| Dual-expert phase routing | Move-Then-Operate (ICML 2026) | Coarse "move" vs contact-critical "operate" experts, learned selector | Phase-label supervision needed | +24% over monolithic Ο0 |
- Margins ~0.5 pp on single benchmarks; no matched-scale cross-lab comparison; the two Qwen flagships disagree internally with no reconciliation.
- Latency implications (prefix-KV reuse vs full re-attention) are argued, never measured.
- Ablations are manipulation-only; the ranking for whole-body/humanoid conditioning rests on Ξ¨β's single comparison.
- A second axis opened at ICML 2026 β what flows through the connection: LangForce's Bayesian decomposition shows naive wiring lets policies shortcut past language entirely (+11.3% OOD when countered) β wiring and grounding are not independent choices.
β Back to Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)