-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 XR 1
Venue: ICLR 2026 Category: VLA Architecture β Cross-Embodiment Trend tag: Trend 4
flowchart LR
V[Visual dynamics] --> Vbr[Vision branch VQ-VAE]
M[Motion trajectory] --> Mbr[Motion branch VQ-VAE]
Vbr --> CB((SHARED codebook<br/>UVMC))
Mbr --> CB
CB --> Lat[Common embodiment-free latent]
Lat --> Df[Decode for Franka]
Lat --> Du[Decode for UR5]
Lat --> Dh[Decode for hand]
Vision and motion are typically learned as separate modalities in VLAs β a vision encoder for pixels, an action head for motor output. This separation wastes the strong correlation between visual dynamics and the motions that caused them β a correlation that is embodiment-independent.
Unified Vision-Motion Codes (UVMC): a dual-branch VQ-VAE in which a vision branch (visual dynamics) and a motion branch (robot motion) quantize into one shared discrete codebook, forcing the two modalities into a common latent space. UVMC acts as an intermediate representation between observations and actions. Crucially, for human videos (no action labels), the objective reduces to the vision-reconstruction loss only (β_total^human = β_vis), so Internet-scale human video (e.g. Ego4D) trains the vision branch alongside robot data.
The VLM backbone is PaliGemma (SigLIP visual encoder + Gemma transformer, ~2.6B params); a lightweight variant XR-1-Light uses Florence-2 (~230M trainable params). The "X" denotes XR-1's three "cross" capacities: cross-data (web human video + robot data), cross-modality (visionβmotion alignment), and cross-embodiment control.
XR-1 uses a three-stage paradigm:
- Self-supervised UVMC learning β train the dual-branch VQ-VAE on large-scale robot + human video.
- UVMC-guided generalist pre-training β inject UVMC knowledge into a VLM via learnable tokens on the cross-embodiment XR-D dataset (~158k trajectories, ~69.1M frames).
- Task-specific post-training β refine for deployment.
Validated with >14,000 real-world rollouts across six robot embodiments and 120+ manipulation tasks, consistently beating Ο0.5, Ο0, RDT, UniVLA, and GR00T-N1.5 with strong generalization to novel objects, backgrounds, distractors, and illumination. Reported gains include Dual-Arm UR-5e ~72% vs Ο0.5 ~62% / Ο0 ~43%, Tien Kung 2.0 72% vs Ο0.5 ~41%, and CALVIN 4.256 vs Ο0.5 3.885.
The strongest version of the "learn in a common embodiment-free latent" idea in 2026. The shared codebook is a commitment: it forces vision and motion to negotiate a joint vocabulary, and the quality of that vocabulary determines downstream cross-embodiment transfer.
- ICLR 2026 listing
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)