-
Notifications
You must be signed in to change notification settings - Fork 0
pi series evolution
A side-by-side comparison of Physical Intelligence's VLA releases through model, data, and training lenses. Each version is an additive delta on the previous one β the table makes the deltas legible.
| Model | Date | arXiv / Ref | Headline claim |
|---|---|---|---|
| Ο0 | Oct 2024 (RSS 2025) | 2410.24164 | First cross-embodiment flow-matching VLA |
| Ο0-FAST | Jan 2025 | FAST tokenizer (RSS 2025) | 5Γ faster autoregressive variant via action tokenization |
| Ο0.5 | Apr 2025 (CoRL 2025) | 2504.16054 | Open-world generalization in unseen homes via co-training |
| Ο0.6 | Nov 2025 | Model card | Out-of-the-box dexterity β no post-training needed |
| Ο*0.6 + RECAP | Nov 2025 | 2511.14759 ("Ο0.6: a VLA That Learns From Experience"*) | RL-from-experience for flow-matching VLAs (RECAP = RL with Experience & Corrections via Advantage-conditioned Policies) |
| Ο0.7 | Apr 2026 | pi.website/pi07 | Steerable generalist with compositional generalization |
| Ο0 | Ο0.5 | Ο0.6 | Ο0.7 | |
|---|---|---|---|---|
| VLM backbone | PaliGemma-3B | Standard VLM (web-pretrained) | Gemma3-4B (switch) | Gemma3-4B (same) |
| Action expert | Flow matching, ~300M | Flow matching (smaller than backbone) | 860M flow matching | 860M flow matching (same) |
| Total params | ~3B | ~3β4B | ~5B | ~5B + 14B world model (BAGEL) |
| Hierarchy | Flat | High-level subtask prediction added | Same hierarchical (preserved) | + world-model subgoal images in prompt |
| Memory / history | Single-frame | Single-frame | Single-frame | MEM video history encoder (6 frames, temporal+spatial compression) |
| Proprioception | Discretized tokens | Discretized tokens | Discretized tokens | Linear projection to backbone dim |
| Input cameras | 3 | up to 3 | up to 4 (base + 2 wrist + rear) @ 448Β² | 4 cameras + up to 3 subgoal images |
| Action chunk / denoise | 50 / 10 | 50 / fewer | 50 / 5 (63ms on H100) | 50 / 5 (+ RTC training for 0β240ms latency) |
| Attention mask | Causal | Global bidirectional (image + text λͺ¨λ bidir, per Ο0.7 Appendix B) | (paper λͺ μ μμ β Ο0.7 Appendix B "no image goals β Ο0.5 λμΌ"μμ μμΆλ‘ νλ©΄ Ο0.5μ κ°μ global bidirectionalμΌ κ°λ₯μ± λμ) | Conditional block-causal (Appendix B): subgoal images μμ λ β obs & subgoal tokens bidir within themselves, subgoalμ΄ obs attend, νμ text tokens (subtask/metadata) causal, 50 action tokens bidir + VLM activations attend; μ΄λ―Έμ§ goals μμ λ β Ο0.5-style global bidirectionalλ‘ fallback |
| Ο0 | Ο0.5 | Ο0.6 | Ο0.7 | |
|---|---|---|---|---|
| Main loss | Flow-matching (continuous) | Two-stage: FAST-token pretrain β flow-matching post-train | Knowledge Insulation (KI): VLM learns FAST + web, action expert learns continuous; gradients don't flow back | KI (inherited) |
| Prompt Ct | Task string only | Task + semantic subtask βΜ | Task + subtask + optional metadata | Task + subtask + subgoal images + metadata + control-mode, each with independent dropout |
| Post-training | Required per task | Required for many tasks | Out-of-the-box for many tasks | Out-of-the-box matches RL specialists |
| RL | β | β | β | Ο*0.6 RECAP is a separate branch |
| CFG | β | β | β | Classifier-free guidance on metadata (Ξ² β {1.3, 1.7, 2.2}) |
| Inference-latency handling | β | β | β | Real-Time Chunking (RTC) trained with simulated 0β240ms delays |
| Ο0 | Ο0.5 | Ο0.6 | Ο0.7 | |
|---|---|---|---|---|
| Demonstrations | Multi-robot (single-arm, bimanual, mobile) | ~400 hrs mobile manipulator Γ ~100 homes + cross-embodiment lab + OXE | Ο0.5 mix + expanded in-home | Ο0.5/0.6 mix further expanded |
| Cross-embodiment | Yes | Yes (97.6% pretraining = non-MM) | Yes | Yes + UR5e zero-shot transfer demonstrated |
| Web / multimodal data | Inherited via VLM | Added (detection, captioning, VQA, semantic prediction co-training) | Same + bbox/keypoint prediction | Same + video captioning of robot data and web videos |
| Human video | β | β | β | Egocentric human video as first-class data source |
| Suboptimal / failure data | Excluded | Excluded | Excluded | Heavily used: failures, RL rollouts from Ο*0.6, autonomous eval data, interventions |
| Open-source robot data | OXE | OXE + community sets | Same | + DROID (Franka) |
flowchart TB
pi0["Ο0 (2024)<br/>Cross-embodiment<br/>flow-matching VLA"]
pi05["Ο0.5 (2025)<br/>+ co-training<br/>+ subtask hierarchy<br/>+ FAST-token pretrain<br/>β open-world generalization"]
pi06["Ο0.6 (Nov 2025)<br/>+ Gemma3-4B + KI<br/>+ optional metadata<br/>β out-of-the-box dexterity"]
pistar["Ο*0.6 (Nov 2025)<br/>+ RECAP<br/>(advantage-conditioned RL)<br/>β specialist-grade policies"]
pi07["Ο0.7 (Apr 2026)<br/>+ MEM history + subgoal<br/> images + metadata<br/>+ suboptimal + RL data<br/>β compositional generalization"]
pi0 --> pi05 --> pi06
pi06 --> pistar
pi06 --> pi07
pistar -. rollouts distilled .-> pi07
- Ο0 β Ο0.5: Hierarchical split (high-level subtask head + low-level action expert). Discrete FAST-token stage for pretraining.
- Ο0.5 β Ο0.6: Backbone upgraded to Gemma3-4B, action expert grown to 860M, Knowledge Insulation stabilizes joint training.
- Ο0.6 β Ο0.7: Keeps the same backbone/expert; adds MEM video-history encoder, subgoal-image conditioning, linear-projected proprioception, RTC, and CFG on metadata. Worldmodel (BAGEL-14B) is a separate module that feeds the prompt, not an architectural change to the policy itself.
- Ο0 β Ο0.5: Large-scale home data (~400 hrs MM Γ 100 homes), web multimodal co-training.
- Ο0.5 β Ο0.6: More of the same, richer metadata annotation.
- Ο0.6 β Ο0.7: Doubles down on heterogeneity β failures, autonomous rollouts, RL-specialist traces, egocentric human video, DROID. The novelty is not the sources but that metadata conditioning lets them be used productively instead of averaged.
- Ο0 β Ο0.5: Two-stage FASTβflow-matching; co-training on VQA/detection/subtask text.
- Ο0.5 β Ο0.6: Knowledge Insulation (one-stage joint training with gradient-insulated VLM).
-
Ο0.6 β Ο*0.6: Advantage-conditioned RL (RECAP) β no log-prob RL needed (flow-matching policies give no action log-probs, so PPO/SAC are inapplicable). A multi-task distributional value function
$V^{\pi_{ref}}(o_t,\ell)$ is trained on all data (demos + autonomous rollouts + teleoperated interventions); the policy then conditions on a binarized advantage indicator as a text token ("Advantage: positive/negative"), and at inference is run on the positive branch. Trained end-to-end with Knowledge Insulation. On the hardest tasks (laundry folding, box assembly, espresso) RECAP more than doubles task throughput and roughly halves the failure rate vs the base policy. - Ο0.6 β Ο0.7: Rich prompt conditioning + per-component dropout + CFG, distilling Ο*0.6 RL behavior into a generalist without RL at post-train time.
The Ο series is evolving along a consistent thesis: push more capability into pre-training so post-training becomes optional. Ο0 needed per-task finetuning; Ο0.5 reduced it; Ο0.6 removed it for many tasks; Ο0.7 removes it even for tasks previously only solvable by RL specialists β and shows the first signs of compositional (not just in-distribution) generalization in the series (the Ο0.7 blog phrases it as "the first signs of compositional generalization," recombining skills across tasks).
β Back to Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)