-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2025 pi05
Venue: CoRL 2025 (Oral) Β· Author: Physical Intelligence Β· arXiv: 2504.16054 Category: VLA Architecture (Baseline) Trend tag: Hierarchy wins in the open world
flowchart LR
V[Vision] --> VLM[VLM Backbone<br/>web-pretrained]
L[Task: 'clean the kitchen'] --> VLM
VLM --> HL[High-level<br/>subtask predictor<br/>'open the drawer']
VLM --> AE[Flow-Matching<br/>Action Expert<br/>continuous chunks]
HL --> AE
noise --> AE
subgraph Co-training sources
WD[Web VQA + captions]
DET[Object detection]
SUB[Subtask semantic prediction]
TELE[Multi-robot teleop]
end
WD & DET & SUB & TELE -. mixed batches .-> VLM
Ο0 was a cross-embodiment flow-matching VLA that worked on in-distribution tasks. But dropping it into an entirely new home β cleaning an unfamiliar kitchen or bedroom β broke generalization: the model couldn't parse novel scene layouts and couldn't decompose long-horizon housework into executable subtasks.
Two additions on top of Ο0:
- Hierarchical split. A high-level head predicts the next semantic subtask (βΜ, e.g., "open the fridge door") via autoregressive token decoding; a low-level ~300M-param flow-matching expert generates continuous action chunks (50-step / 1-second) conditioned on both the overall task β and the subtask βΜ.
- Heterogeneous co-training. Mixed batches include multi-robot teleop + web VQA/captions + object-detection tasks + subtask-prediction tasks. Pretraining happens with FAST action tokens (discrete, for gradient stability), then flow-matching for the action expert.
Data scale: ~400 hrs mobile-manipulator data Γ ~100 homes, plus cross-embodiment lab data and OXE; 97.6% of pretraining examples are from non-MM sources.
First end-to-end learning-enabled system to do long-horizon, dexterous manipulation in entirely unseen homes β cleaning kitchens, tidying bedrooms. In out-of-distribution homes Ο0.5 reaches ~94% task success and ~94% language-following, approaching baselines trained directly on the test environments, with quantitatively significant jumps over Ο0 on open-world generalization suites; task-specific post-training still often needed for polish (relaxed later by Ο0.6 and Ο0.7).
Ο0.5 is the ancestor of the entire current Ο series. Every subsequent release builds on it:
- Ο0.6 upgrades backbone to Gemma3-4B + Knowledge Insulation.
- Ο*0.6 + RECAP adds advantage-conditioned RL for flow-matching.
- Ο0.7 adds MEM video history, subgoal-image world-model conditioning, and metadata prompting.
See Ο series evolution for the full side-by-side. Ο0.5 also establishes the "hierarchical generalist" template that OneTwoVLA, Long-VLA, and ICLR 2026's reasoning-augmented VLAs follow.
- arXiv: https://arxiv.org/abs/2504.16054
- Project PDF: https://www.pi.website/download/pi05.pdf
- Blog: https://www.pi.website/blog/pi05
- OpenPi: https://github.com/Physical-Intelligence/openpi
β Back to CoRL-2025
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)