Skip to content

CoRL 2025 pi05

Heungwoo edited this page Jun 1, 2026 · 2 revisions

Ο€0.5 β€” A VLA Model with Open-World Generalization

Venue: CoRL 2025 (Oral) Β· Author: Physical Intelligence Β· arXiv: 2504.16054 Category: VLA Architecture (Baseline) Trend tag: Hierarchy wins in the open world

Approach diagram

flowchart LR
  V[Vision] --> VLM[VLM Backbone<br/>web-pretrained]
  L[Task: 'clean the kitchen'] --> VLM
  VLM --> HL[High-level<br/>subtask predictor<br/>'open the drawer']
  VLM --> AE[Flow-Matching<br/>Action Expert<br/>continuous chunks]
  HL --> AE
  noise --> AE
  subgraph Co-training sources
    WD[Web VQA + captions]
    DET[Object detection]
    SUB[Subtask semantic prediction]
    TELE[Multi-robot teleop]
  end
  WD & DET & SUB & TELE -. mixed batches .-> VLM
Loading

Problem

Ο€0 was a cross-embodiment flow-matching VLA that worked on in-distribution tasks. But dropping it into an entirely new home β€” cleaning an unfamiliar kitchen or bedroom β€” broke generalization: the model couldn't parse novel scene layouts and couldn't decompose long-horizon housework into executable subtasks.

Method

Two additions on top of Ο€0:

  1. Hierarchical split. A high-level head predicts the next semantic subtask (β„“Μ‚, e.g., "open the fridge door") via autoregressive token decoding; a low-level ~300M-param flow-matching expert generates continuous action chunks (50-step / 1-second) conditioned on both the overall task β„“ and the subtask β„“Μ‚.
  2. Heterogeneous co-training. Mixed batches include multi-robot teleop + web VQA/captions + object-detection tasks + subtask-prediction tasks. Pretraining happens with FAST action tokens (discrete, for gradient stability), then flow-matching for the action expert.

Data scale: ~400 hrs mobile-manipulator data Γ— ~100 homes, plus cross-embodiment lab data and OXE; 97.6% of pretraining examples are from non-MM sources.

Results

First end-to-end learning-enabled system to do long-horizon, dexterous manipulation in entirely unseen homes β€” cleaning kitchens, tidying bedrooms. In out-of-distribution homes Ο€0.5 reaches ~94% task success and ~94% language-following, approaching baselines trained directly on the test environments, with quantitatively significant jumps over Ο€0 on open-world generalization suites; task-specific post-training still often needed for polish (relaxed later by Ο€0.6 and Ο€0.7).

Significance

Ο€0.5 is the ancestor of the entire current Ο€ series. Every subsequent release builds on it:

  • Ο€0.6 upgrades backbone to Gemma3-4B + Knowledge Insulation.
  • Ο€*0.6 + RECAP adds advantage-conditioned RL for flow-matching.
  • Ο€0.7 adds MEM video history, subgoal-image world-model conditioning, and metadata prompting.

See Ο€ series evolution for the full side-by-side. Ο€0.5 also establishes the "hierarchical generalist" template that OneTwoVLA, Long-VLA, and ICLR 2026's reasoning-augmented VLAs follow.

Links

Related pages

← Back to CoRL-2025

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally