Skip to content

Review pi06

Heungwoo edited this page Jun 1, 2026 · 2 revisions

In-Depth Review β€” Ο€0.6: A VLA Model for General Robot Control

Paper: Ο€0.6 Model Card Β· Physical Intelligence Β· November 17, 2025 Related summary: Ο€0.6

This page is the long-form companion to the Ο€0.6 summary. Everything is sourced from the Ο€0.6 model card at website.pi-asset.com/pi06star/PI06_model_card.pdf.

πŸ“Ž Ο€ series context

Release Date Page Role
Ο€0 Oct 2024 (RSS 2025) β€” Flow-matching VLA foundation
Ο€0.5 Apr 2025 (CoRL 2025 Oral) Ο€0.5 Hierarchical + co-training
Ο€0.5-KI Dec 2025 Knowledge Insulation The training recipe Ο€0.6 adopts
Ο€0.6 Nov 2025 this review / summary Gemma3-4B backbone + KI + metadata prompting
Ο€*0.6 + RECAP Nov 2025 Ο€*0.6 + RECAP RL-from-experience on Ο€0.6 base
Ο€0.7 Apr 2026 Ο€0.7 / long-form MEM + world-model subgoals + metadata CFG
(meta) β€” Ο€ series evolution Side-by-side of all releases

Ο€0.6 is the production baseline for 2026. It's the model every ICLR 2026 VLA paper compares against; it's the base for Ο€*0.6's RL specialists; and it's the architectural template Ο€0.7 inherits verbatim (same Gemma3-4B + 860M action expert, same Knowledge Insulation).


1. TL;DR

Ο€0.6 is Physical Intelligence's November 2025 VLA that upgrades Ο€0.5 along three axes β€” backbone (PaliGemma β†’ Gemma3-4B), training (joint β†’ Knowledge-Insulation two-path), and prompt (task-only β†’ task + optional metadata) β€” while keeping the hierarchical design. Result: first Ο€-series release that achieves strong out-of-the-box performance without task-specific post-training. Laundry (T-shirts + shorts) goes from ~0% (Ο€0.5 needed fine-tuning) to "folds reliably" out-of-the-box (the card gives no laundry %); box assembly goes from 0% to 20% full assembly out-of-the-box (the one hard static-task number in the card); shirt folding, table bussing, and mobile tasks see consistent throughput/speed gains (exact multipliers are bar-chart estimates, not quoted). 63 ms per action chunk on a single H100 with 3 cameras. Ο€0.6 is the base model for Ο€*0.6 + RECAP (RL-from-experience specialists) and the architectural template for Ο€0.7 (which inherits the backbone + action expert verbatim).

2. Motivation β€” what Ο€0.5 left on the table

Ο€0.5 (April 2025, CoRL 2025 Oral) demonstrated open-world generalization in unseen homes via co-training on heterogeneous data + a hierarchical subtask head. But:

  • Post-training was still necessary. For dexterous tasks (laundry folding, box assembly), Ο€0.5 needed task-specific fine-tuning with curated high-quality data to reach non-zero success rates. This is expensive and brittle.
  • Action-expert gradients corrupted the VLM. Joint training mixed language-grounded representations with flow-matching gradients β€” a known stability problem that limited how hard you could push the action expert.
  • No handle to steer behavior at deployment. A trained policy was what it was β€” no metadata or prompt knob to bias toward high-quality modes at test time.
  • Backbone was showing its age. PaliGemma (Ο€0 / Ο€0.5) was fine but Gemma3's multimodal improvements were worth capturing.

Ο€0.6 addresses all four.


3. Representative diagrams

Figure 1 from the model card β€” Ο€0.6 architecture

Ο€0.6 architecture (Figure 1 from Physical Intelligence, Nov 2025)

Figure 1 of the Ο€0.6 model card (Physical Intelligence, Nov 17 2025). The VLA consists of a pre-trained VLM (SigLIP 400M + Gemma3 4B) that consumes up to four 448Γ—448 images, a language prompt, tokenized proprioceptive state, and optional episode metadata. The backbone produces both discretized actions (FAST tokens) β€” which supply the VLM's training signal via Knowledge Insulation β€” and features feeding a separate 860M action expert that generates continuous actions via flow matching. Included for scholarly review.

Figure 2 from the model card β€” static-task results

Ο€0.6 static-task results (Figure 2 from Physical Intelligence, Nov 2025)

Figure 2 of the model card. Out-of-the-box (no task-specific fine-tuning) comparison of Ο€0.5 vs. Ο€0.6 on four static-robot tasks. Top row: success rate. Bottom row: throughput (successes per hour). Ο€0.6's biggest wins are on laundry folding (now folds reliably out-of-the-box vs. ~0% before) and box assembly (0 β†’ 20% full assembly, the one number quoted in the card text) β€” previously both required fine-tuning with curated data.

Our reconstruction as mermaid

flowchart TB
  subgraph Prompt[Prompt Ct]
    L[Task: 'clean the bedroom']
    SL[Subtask: 'pick up the pillow']
    M[Optional metadata<br/>speed / quality / etc.]
  end

  V[Up to 4 cameras<br/>448Γ—448<br/>base + 2 wrist + rear] --> B[Ο€0.6 VLA<br/>SigLIP 400M + Gemma3-4B]
  PR[Proprioception<br/>tokenized] --> B
  Prompt --> B

  B -- FAST-tokenized actions --> CE[Discrete CE loss<br/>trains VLM representation]
  B -- Web co-training<br/>subtask prediction --> CE
  B -- conditioning features --> AE[860M Action Expert<br/>flow matching]
  noise[Noise] --> AE
  AE -- 5 denoising steps --> A[Continuous action chunk<br/>63 ms / chunk on H100]

  AE -. 🚫 NO gradient back to VLM<br/>(Knowledge Insulation) .-> B
Loading

4. Method β€” what changed vs. Ο€0.5

4.1 Backbone: PaliGemma β†’ Gemma3-4B

Ο€0.5 used PaliGemma (SigLIP vision + Gemma language). Ο€0.6 swaps in Gemma3-4B (with a 400M SigLIP vision encoder) β€” a later-generation VLM with stronger multimodal pretraining. This is a ~1-generation upgrade and is kept unchanged in Ο€0.7.

4.2 Action expert: ~860M parameters, same layer count as backbone

The action expert:

  • Has the same number of layers as the VLM backbone.
  • Consists of ~860M parameters.
  • Uses flow matching to produce continuous action chunks from noise.
  • 5 denoising (Euler) steps at inference.
  • 63 ms per action chunk on a single H100 with 3 camera inputs.
  • Bidirectional attention among action tokens.

Identical to Ο€0.7's action expert β€” Ο€0.7 keeps this subsystem unchanged and adds new prompt modalities + MEM history around it.

4.3 Input format

  • Up to 4 images at 448Γ—448: base camera, up to two wrist cameras, optional backward camera for mobile manipulators.
  • Image tokens + tokenized language prompt + tokenized proprioception concatenated.
  • Bidirectional attention among image tokens (inherited from Ο€0.5); causal attention among text tokens.

4.4 Knowledge Insulation β€” the key training change

The VLM and action expert are trained simultaneously but with no gradient flow from the action expert into the VLM:

  • VLM is supervised by discretized FAST action tokens (cross-entropy loss β€” the stable objective it was built for) + co-training tasks including multi-modal web data, subtask prediction, bounding-box / keypoint prediction.
  • Action expert learns continuous actions via flow matching, attending to VLM activations.
  • Gradients from the action-expert flow-matching loss are blocked from propagating into the VLM's parameters.

This is the Knowledge Insulation recipe (NeurIPS 2025 Spotlight β€” Driess et al.) β€” formalized at NeurIPS in December 2025 but already deployed in Ο€0.6 (November 2025). The recipe then propagates unchanged into Ο€0.7.

4.5 Optional metadata in the prompt

Ο€0.6 optionally accepts conditioning metadata alongside the language command β€” a small prompt knob that biases how the task is performed. The space of metadata is not fully documented in the model card but is expanded considerably in Ο€0.7 (speed / quality / mistake / control mode, each with independent dropout and CFG).

4.6 Data

Largely inherits Ο€0.5's mix:

  • Cross-embodiment data from PI robots (static + mobile + bimanual + single-arm).
  • External data sources (OXE, community datasets).
  • Diverse mobile and non-mobile home data collected in-house.
  • High-level subtask prediction examples.
  • Multi-modal web datasets including bounding-box and keypoint prediction co-training tasks.

Notable non-inclusion (vs. Ο€0.7): no autonomous rollouts / failures / RL traces are explicitly part of Ο€0.6's training data β€” that's a Ο€0.7 addition.


5. Results β€” every number the model card reports

All results are out-of-the-box (no task-specific fine-tuning for either model). Ο€0.6 is compared against the Knowledge-Insulation-trained Ο€0.5 (which is itself an improvement over CoRL 2025 Ο€0.5 β€” open-sourced at openpi).

5.1 Static tasks (Figure 2)

Caveat (verified against the model card): The model card's text reports only two hard quantitative statements for the static tasks β€” that Ο€0.6 can "fully assemble the box 20% of the time" out-of-the-box, and that it "can out-of-the-box fold laundry reliably/consistently" (no percentage given). All other figures below are approximate values read off the Figure 2 bar charts β€” they are not stated numerically anywhere in the card text, nor in the RECAP report. The reads were cross-checked against the actual Figure 2 image (each task panel has its own y-axis scale, with standard-error bars), and the throughput bars match the table values within reading precision: Shirt β‰ˆ21β†’β‰ˆ50, Laundry 0β†’β‰ˆ19, Box 0β†’β‰ˆ5 (very large error bar), Table bussing β‰ˆ28β†’β‰ˆ45 succ/hr. Treat all such values as approximate chart reads, not quoted numbers.

Task Ο€0.5 success Ο€0.6 success Ο€0.5 throughput (succ/hr) Ο€0.6 throughput
Shirt folding (flat start) ~85%* ~87%* ~20* ~50*
Laundry folding (T-shirts + shorts from basket) ~0% (needed fine-tuning) "folds reliably" (~65%*) 0 ~20*
Box assembly ~0% (needed fine-tuning) 20% full assembly (card text) 0 ~5*
Table bussing ~80%* ~100%* ~28* ~45*

* = approximate value read off the Figure 2 bar chart (cross-checked against the figure image), not a number stated in the card text or RECAP report.

Card text says: "Across tasks, Ο€0.6 shows significant improvement over Ο€0.5 in speed and often success rates. The biggest differences lie in laundry folding and box assembly β€” previously these two tasks require fine-tuning with high-quality data to achieve non-zero success rates." So: success on already-solved tasks is roughly flat; the consistently improved axis is throughput/speed, and the headline new capability is non-zero performance on laundry folding and box assembly without fine-tuning.

5.2 Mobile tasks (Figure 3) β€” four tasks from the Ο€0.5 paper

Picking up laundry β†’ basket Β· Tidying bed Β· Putting dishes in sink Β· Putting items in drawer.

Results: "Ο€0.6 improves throughput over Ο€0.5 when the average task progress is saturated, or otherwise improves both performance and throughput." No single headline number β€” the pattern is uniform improvement across all four.

5.3 Generalization tasks (Figure 4) β€” 4 task suites Γ— 3 difficulty levels Γ— 12–18 instructions

Mixed static + mobile tasks requiring either:

  • Language-following generalization β€” "pick up the third fruit from the left", "move to where the fresh milk is kept"
  • Novel-skill generalization β€” "wipe the spill with the bread", "hang the shorts into oven handle"

Most language instructions and objects are not seen in training. Ο€0.6 shows "healthy improvements over Ο€0.5 across all settings." Mobile settings are generally harder (longer horizon, more distractors).

5.4 Latency headline

63 ms per action chunk on a single H100 GPU with 3 camera inputs, using 5 denoising steps. This is the production-deployment number the Ο€-series carries forward β€” Ο€0.7 keeps the same 5-step inference budget.


6. Why Ο€0.6 works β€” three contributing factors

The model card attributes the gain to three changes, not one:

  1. Gemma3-4B backbone β€” better multimodal pretraining transfer.
  2. Knowledge Insulation β€” lets the VLM learn stable representations while the action expert is free to be aggressive on flow matching.
  3. Diverse training data + optional metadata prompting β€” the metadata gives a knob for test-time steering without additional training; the diverse data mix provides the coverage that prior models had to fine-tune in.

The model card does not separately ablate these three β€” a key scientific limitation. Ο€0.7 later ablates the metadata contribution heavily and shows it's load-bearing for scaling on suboptimal data; Knowledge Insulation's paper provides the independent ablation for the KI training change.


7. Limitations

7.1 Stated limitations (model card)

The Ο€0.6 model card is a production release document, not a research paper β€” it does not enumerate limitations directly. Implicit gaps from the text:

  • No separate ablation of backbone upgrade / KI / metadata.
  • Benchmarks are PI-internal; no LIBERO / RLBench / SimplerEnv numbers for external comparison.
  • Generalization results are qualitative ("healthy improvements") rather than fully tabulated per task.
  • Throughput comparisons use normalized axes in several plots; raw successes-per-hour for the hardest tasks are not always disclosed.

7.2 Reviewer's concerns

  • No head-to-head vs. discrete-diffusion VLAs. The ICLR 2026 alternative family (Discrete Diffusion VLA, Unified Diffusion VLA, dVLA) directly challenges the Ο€-series' separate-action-expert design β€” Ο€0.6 doesn't engage with this.
  • No open weights. Like every PI tech report, Ο€0.6 is not an independently reproducible artifact; the reference open implementation is Ο€0.5-KI (openpi), not Ο€0.6.
  • Dexterous tasks still fail ~35% of the time. Laundry at 65% is a huge jump from 0%, but it's not production-ready β€” hence the immediate follow-up with Ο€*0.6 + RECAP RL specialists.
  • Box assembly at 20% full assembly is more a "proof of feasibility" than a deployable capability.
  • The metadata space is under-documented. What metadata fields exist? What's their distribution at train time? How are they prompted at test time? Ο€0.7 fills most of this in retroactively; Ο€0.6 leaves it opaque.

8. Significance β€” the production baseline of 2026

Ο€0.6 is the single most-compared-against VLA of 2026. Concretely:

  • ~All ICLR 2026 VLA papers benchmark against Ο€0.6 or Ο€0.5-KI.
  • Ο€*0.6 + RECAP uses Ο€0.6 as the base model for RL-from-experience β€” specialist policies for laundry, box, espresso that are then distilled back into Ο€0.7 via metadata conditioning.
  • Ο€0.7 inherits the Gemma3-4B + 860M action expert + Knowledge Insulation stack unchanged. The Ο€0.7 delta is entirely new prompt modalities + MEM history.
  • Architectural legacy: Ο€0.6 cements the "VLM backbone + separate flow-matching action expert" pattern that every flow-matching VLA at CoRL 2025 and ICLR 2026 follows. It's the pattern that Discrete Diffusion VLA / Unified Diffusion VLA explicitly challenge by unifying action generation back into the VLM transformer.

Category placement (see Review-VLA-Architecture Β§5.B): Ο€0.6 is the canonical exemplar of Category B (Flow-matching action expert) β€” the one other VLA architectures are compared against.


9. What changed from Ο€0.6 β†’ Ο€*0.6 β†’ Ο€0.7

A one-screen summary for readers jumping through the lineage:

Axis Ο€0.6 (Nov 2025) Ο€*0.6 + RECAP (Nov 2025) Ο€0.7 (Apr 2026)
Backbone Gemma3-4B same same
Action expert 860M flow-matching same same
Training recipe Knowledge Insulation same + advantage-conditioned RL same + suboptimal data distillation
Prompt Task + optional metadata same + positive/negative conditioning Task + subtask + subgoal images + rich metadata + control mode (each dropout)
History Single frame same MEM video encoder β€” 4 cameras Γ— 6 history frames
World model None None BAGEL-14B generating subgoal images
Data Ο€0.5 mix Ο€0.6 mix + real-robot RL Ο€0.6 mix + failures + autonomous + RL rollouts + egocentric human + DROID
Post-training needed Sometimes Per-task RL Zero-shot matches specialists

Core insight: the architectural chassis from Ο€0.6 is stable across the entire Nov-2025-to-Apr-2026 window. Improvements come from data, prompt, and history β€” not from the VLM + action-expert core.


10. Links

← Back to PI-pi06 Β· ICLR-2026 Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally