Skip to content

History / Review pi07

Revisions

  • Source-verification audit: fix fabrications, rename PI slugs, link hygiene, add foundational pages - Re-verified all 195 pages vs original sources; removed 47 confirmed fabricated tables/numbers, restored 19 false-positive deletions (full-PDF re-check) - Renamed PI tech-report slugs ICLR-2026-pi07/pi06/RECAP -> PI-pi07/PI-pi06/PI-RECAP (these are PI technical reports, not ICLR 2026 papers); updated 171 wikilinks - Fixed 60 broken wikilinks -> 0 dangling across the wiki - Added foundational pages: OpenVLA, ReKep, AgiBot World Colosseo, RoboBrain 2.0; linked from Home + sidebar + venue indexes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Add VLA attention architectures cross-paper review + correct π series attention - New Review-VLA-Attention.md: per-family attention deep dive (π series, GR00T N1→N1.7, StarVLA, NORA/Qwen-VL) with mermaid topology diagrams and HTML colored mask matrices for block-causal / causal / bidirectional / prefix-LM. - Correct π0.5/π0.6/π0.7 attention story per π0.7 paper Appendix B: π0.5 used global bidirectional attention (image + text), π0.7 introduces conditional block-causal that falls back to π0.5-style global bidir when no subgoal images are in the prompt. - pi-series-evolution.md: rewrite Attention mask row to reflect Appendix B. - Review-pi07.md §4.1: replace single-line attention bullet with conditional-on-image-goals description quoting Appendix B directly. - Home.md: link the new Review-VLA-Attention page in cross-paper reviews.

    @Heungwoo Heungwoo committed May 1, 2026
  • Restructure Home with in-depth reviews as the 3rd browsing axis - Home.md: promote "In-depth reviews" from a bullet under What's-new to a first-class section 3 (peer of Conferences and Topics). Now three browsing axes: venue, topic, in-depth review. - Add a dedicated "Series evolution" section for the π-series page. - Reorder and expand "What's new" as a reverse-chronological list of the latest additions (cross-embodiment review, memory review, π0.7 and VLM4VLA long-form reviews, π0.7 summary, CoRL 2025 survey). - Update CoRL venue status from placeholder to "2025 survey live". - Refresh the "How this wiki is organized" tree. Review-pi07.md: add a prominent "π series context" table near the top listing π0, π0.5, π0.6, π*0.6 + RECAP, π0.7, and the π series evolution page, with one-line role descriptions. This makes the review readable standalone and surfaces the series lineage before the method deep-dive starts. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Apr 17, 2026
  • Embed paper architecture figures in in-depth reviews Extract and include the key architecture / prompt figures directly from the original papers so readers can see the authors' own diagrams alongside the reconstructions: - assets/pi07_fig2_architecture.png — Fig 2 of π0.7 (arch overview) - assets/pi07_fig3_prompts.png — Fig 3 of π0.7 (prompt composition) - assets/vlm4vla_fig1_framework.png — Fig 1 of VLM4VLA (eval pipeline) - assets/vlm4vla_fig2_network.png — Fig 2 of VLM4VLA (VLA network) VLM4VLA figures are licensed CC-BY-4.0 per the arXiv listing; π0.7 figures are from the Physical Intelligence technical report, included for scholarly review with attribution. Each embedded figure has a caption that (a) credits the source paper and figure number and (b) explains what the figure is showing, so the wiki page is self-contained. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Apr 17, 2026
  • Add in-depth reviews for π0.7 and VLM4VLA New long-form review pages with representative diagrams, full accuracy tables, ablations, and limitations — the short summary pages link out to these for readers who want the full numbers. Review-pi07.md covers: - Full multimodal prompt architecture + CFG + RTC - All out-of-the-box-vs-specialist task comparisons - Cross-embodiment UR5e laundry result (85.6% progress, 80% success, matching expert teleoperators on their first UR5e attempt) - Ablations: no metadata, no eval data, mixed-quality scaling, diversity scaling - Authors' stated limitations + reviewer concerns (no head-to-head with discrete-diffusion VLAs, BAGEL cost, distillation confound) Review-VLM4VLA.md covers: - 9 VLM backbones × 3 benchmarks (Calvin / SimplerEnv / Libero) with full tables - Correlation analysis: r=0.84 on Calvin, r=-0.36 on SimplerEnv, r=-0.19 on Libero - Vision-vs-language ablation (freezing vision: -42 points; freezing word embeds: -0.2) - 7 embodied auxiliary tasks all hurt downstream VLA performance - Kosmos-2 (1.7B) matches π0 (~3.1B) on SimplerEnv; Qwen3VL-2B tops Calvin - Action-supervised vision-encoder FT closes the gap by +18.1 points Navigation updates: - Short summary pages for π0.7 and VLM4VLA now link to the reviews - _Sidebar.md has a new "In-depth reviews" section - Home.md "What's new" surfaces both reviews Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Apr 17, 2026