-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 UniCoD
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Jianke Zhang, Yucheng Hu, Yanjiang Guo, Xiaoyu Chen, Yichen Liu, Wenna Chen, Chaochao Lu, Jianyu Chen Traction (2026-06): 6 citations (arXiv)
Building generalist robot policies that handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage large-scale pretraining, prior Vision-Language-Action (VLA) work typically builds generalist policies either on top of vision-language models (VLMs) or on generative models. However, the authors argue that both semantic understanding (from vision-language pretraining) and visual dynamics modeling (from visual-generation pretraining) are crucial for embodied robots, and neither family alone captures both.
UniCoD builds on recent unified models of generation and understanding, which have shown strong capabilities in both comprehension and generation through large-scale pretraining. The premise is that robotic policy learning can likewise benefit from the combined strengths of understanding, planning, and continuous future-representation learning. Concretely:
- UniCoD acquires the ability to dynamically model high-dimensional visual features by pretraining on over 1M internet-scale instructional manipulation videos, giving it predictive/generative representations of future visual states.
- It is then fine-tuned on data collected from the robot embodiment, learning the mapping from these predictive representations to action tokens.
- The "unified continuous and discrete" framing combines continuous future visual representation learning with discrete action-token prediction within one model.
flowchart LR
A[1M+ internet manipulation videos] --> B[Unified generation + understanding pretraining]
B --> C[Continuous predictive visual representations]
C --> D[Fine-tune on robot embodiment data]
D --> E[Map predictive reps -> discrete action tokens]
E --> F[Generalist robot policy]
"9% and 12% over baselines in sim and real-world OOD tasks". The paper reports that UniCoD consistently outperforms baseline methods by 9% in simulation environments and by 12% on real-world out-of-distribution tasks. (The verbatim ICML abstract does not break these numbers down further per benchmark.)
UniCoD argues that the two dominant VLA pretraining paradigmsβVLM-based semantic grounding and generative visual-dynamics modelingβare complementary rather than competing, and that a single unified generation-and-understanding backbone can inherit both. Leveraging 1M+ internet manipulation videos for predictive representation learning, then grounding to action tokens, points toward more data-efficient and OOD-robust generalist policies in the 2026 VLA landscape.
- ICML 2026: https://icml.cc/virtual/2026/poster/63426
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)