Skip to content

PI pi06

Heungwoo edited this page Jun 1, 2026 · 1 revision

Ο€0.6 β€” A Vision-Language-Action Flow Model for General Robot Control

Venue: Physical Intelligence (technical report) Β· Date: Nov 2025 Category: Baseline (VLA Architecture) Trend tag: Baseline for Trend 1

Approach diagram

flowchart LR
  V[Camera frames<br/>3 views] --> B[Gemma3-4B<br/>VLM backbone]
  L[Language instruction] --> B
  B --> HL[High-level subtask predictor]
  B --> AE[~860M Flow-Matching<br/>Action Expert]
  HL --> AE
  noise[Noise prior] --> AE
  AE -- 5 Euler steps --> A[Continuous action chunk]
  A -- 63 ms on H100 --> R[Real robot]
Loading

See the Ο€0.6 model card (link below) for the authors' own architecture diagram.

Problem

A generalist robot policy needs strong visual-linguistic grounding and precise low-level control and sub-100 ms inference. Earlier generalist models either decoded actions autoregressively (too slow at billion-scale) or used ad-hoc MLP heads (no generative modeling of action distributions).

Method

Hierarchical VLA (preserving the Ο€0.5 design): a high-level component predicts language-level subtasks; a low-level component generates continuous action chunks. The backbone is initialized from Gemma3-4B (with a SigLIP-400M vision encoder). The ~860M-parameter "action expert" is a parallel transformer expert with the same number of layers as the backbone (a mixture-of-experts arrangement, as in Ο€0/Ο€0.5 β€” not an external MLP head), and is trained with flow matching β€” it learns a continuous vector field that transports a noise distribution to action chunks, conditioned on backbone features; gradients from the action expert do not flow back into the VLM backbone (Knowledge Insulation). The backbone additionally predicts FAST discrete action tokens and co-training web data alongside the flow-matched continuous actions. At inference, 5 denoising (Euler) steps suffice.

Results

On a single H100 GPU with 3 cameras: ~63 ms per action chunk. Generalizes across multiple embodiments and real-world manipulation tasks. Deployed on Physical Intelligence's commercial and academic platforms.

Significance

The de facto production baseline for 2026 VLA research. Almost every ICLR 2026 VLA paper compares against Ο€0, Ο€0.5, or Ο€0.6. Architectural legacy: established "VLM backbone + separate flow-matching action expert" as the dominant pattern β€” the pattern that ICLR 2026's discrete-diffusion VLAs directly challenge (Discrete Diffusion VLA, Unified Diffusion VLA).

Now superseded by Ο€0.7 (Apr 2026), which keeps the same backbone+expert and adds MEM history, subgoal-image world-model conditioning, expanded episode-metadata prompting (Ο€0.6 already supports optional metadata conditioning), and suboptimal-data distillation.

Links

πŸ“– In-depth review

For a long-form review with model-card figures, accuracy tables, Knowledge Insulation training details, limitations, and the Ο€0.6 β†’ Ο€*0.6 β†’ Ο€0.7 delta: In-Depth Review of Ο€0.6.

πŸ“– In-depth architecture review

Compare the Ο€-series' flow-matching action expert to other VLA architecture families: VLA Architectures Review.

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally