VLM4VLA review: re-verify against arXiv v2, fix pi0 Calvin Task-3 value v1->v2 diff contains a single substantive change: pi0 Calvin Task-3 0.786 -> 0.686 (fixes internal sum; total 3.509 unchanged). Header now records version history and the Tsinghua x Qwen affiliation split. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add cross-paper review: Qwen Team's VLA Program VLM4VLA -> Qwen-VLA -> Qwen-Robot Suite (RobotManip/RobotNav/RobotWorld): shared doctrine (Qwen3.5-4B, flow matching, lambda=0.1 VL co-training, synthetic-data scaling, language-as-interface), diagnostic-to-flagship trace, internal contradictions between the two flagship VLAs, and lab-program positioning vs PI/GR00T/TRI/Gemini. Cross-links from the three constituent reviews; indexed in Home/Changelog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Source-verification audit: fix fabrications, rename PI slugs, link hygiene, add foundational pages - Re-verified all 195 pages vs original sources; removed 47 confirmed fabricated tables/numbers, restored 19 false-positive deletions (full-PDF re-check) - Renamed PI tech-report slugs ICLR-2026-pi07/pi06/RECAP -> PI-pi07/PI-pi06/PI-RECAP (these are PI technical reports, not ICLR 2026 papers); updated 171 wikilinks - Fixed 60 broken wikilinks -> 0 dangling across the wiki - Added foundational pages: OpenVLA, ReKep, AgiBot World Colosseo, RoboBrain 2.0; linked from Home + sidebar + venue indexes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Embed paper architecture figures in in-depth reviews Extract and include the key architecture / prompt figures directly from the original papers so readers can see the authors' own diagrams alongside the reconstructions: - assets/pi07_fig2_architecture.png — Fig 2 of π0.7 (arch overview) - assets/pi07_fig3_prompts.png — Fig 3 of π0.7 (prompt composition) - assets/vlm4vla_fig1_framework.png — Fig 1 of VLM4VLA (eval pipeline) - assets/vlm4vla_fig2_network.png — Fig 2 of VLM4VLA (VLA network) VLM4VLA figures are licensed CC-BY-4.0 per the arXiv listing; π0.7 figures are from the Physical Intelligence technical report, included for scholarly review with attribution. Each embedded figure has a caption that (a) credits the source paper and figure number and (b) explains what the figure is showing, so the wiki page is self-contained. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add in-depth reviews for π0.7 and VLM4VLA New long-form review pages with representative diagrams, full accuracy tables, ablations, and limitations — the short summary pages link out to these for readers who want the full numbers. Review-pi07.md covers: - Full multimodal prompt architecture + CFG + RTC - All out-of-the-box-vs-specialist task comparisons - Cross-embodiment UR5e laundry result (85.6% progress, 80% success, matching expert teleoperators on their first UR5e attempt) - Ablations: no metadata, no eval data, mixed-quality scaling, diversity scaling - Authors' stated limitations + reviewer concerns (no head-to-head with discrete-diffusion VLAs, BAGEL cost, distillation confound) Review-VLM4VLA.md covers: - 9 VLM backbones × 3 benchmarks (Calvin / SimplerEnv / Libero) with full tables - Correlation analysis: r=0.84 on Calvin, r=-0.36 on SimplerEnv, r=-0.19 on Libero - Vision-vs-language ablation (freezing vision: -42 points; freezing word embeds: -0.2) - 7 embodied auxiliary tasks all hurt downstream VLA performance - Kosmos-2 (1.7B) matches π0 (~3.1B) on SimplerEnv; Qwen3VL-2B tops Calvin - Action-supervised vision-encoder FT closes the gap by +18.1 points Navigation updates: - Short summary pages for π0.7 and VLM4VLA now link to the reviews - _Sidebar.md has a new "In-depth reviews" section - Home.md "What's new" surfaces both reviews Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>