Skip to content

History / Review VLM Action Connection

Revisions

  • Escape pipe in all in-table wikilinks wiki-wide (202 links, 17 files) GitHub-wiki table cells read a wikilink's separator | as a column delimiter, splitting the cell and breaking the link. Escape to \| in every table-row wikilink (Home nav, RSS-2026-Papers, topic surveys). Prose wikilinks left as plain | (render correctly outside tables). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Weave ICML 2026 evidence into deep-dive surveys; revise three verdicts 14 State-of-the-Field sections gain ICML 2026 findings from the 99-paper index: recipes and latent-action supervision (VLANeXt, From-Pixels-to-Tokens, XR-1), MoT dual-systems and shortcut counters, the 9-paper efficiency cluster (Reflex 50Hz, GridS -76% FLOPs, XPU profile, latent reasoning -90%), reward/critic and model-based RL (VLAC, VLAW +39.2%), memory (HiMe/SOMA/CAPS), world models (DreamDojo 44kh, LAC-WM, dWorldEval), dexterous (DexMachina/DECO/Tabero/CTSRL), cross-embodiment (OXE-AugE, latent motion codes), evaluation (LIBERO-Gen, VLA-Arena, FixBench, TRAP). Verdicts revised: forgetting milder than assumed; discrete-token verdict scoped to robot-action auxiliaries; WM-evaluator action gap first crack. Home synced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Decision map: prominent deep-dive links + uniform detail-page template Home fold-outs now lead with a heading-level "Deep dive ->" link and compress trend/approaches/limitations into a labeled 3-row table. The 12 detail-page State-of-the-Field sections are rewritten to one template (Verdict quote + Trend + Approaches-and-trade-offs + optional Established-findings + Limitations, dated Aug 2026); the three standalone surveys get matching headers with structure legends. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Survey-depth topic pages: 12 State-of-the-Field updates + 3 new surveys Each decision-map topic's detail page now carries a dated July-2026 survey section: trend arc through the latest venues, approach taxonomy with definitions and trade-offs, and current limitations. Three previously page-less topics get dedicated surveys: Human-Video Transfer (emergence/decoupling/synthesis fork + decision guide), VLA Evaluation (indictment + 2026 toolkit + emerging norms), Real-Time Execution (RTC->Legato arc + approach comparison). Home fold-outs link the full surveys; Reviews catalog updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 27, 2026
  • Fix broken tables: escape unescaped pipes inside wikilinks in table rows ICLR.md (and 12 other pages) had [[label|Page]] wikilinks with raw pipes inside GFM table cells, which the GitHub-wiki renderer reads as column separators — mangling the table. Escaped the wikilink-internal pipes to \| (matching the convention already used by CVPR/NeurIPS/CoRL pages); table column separators left intact. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 11, 2026
  • Update review theses where ICRA 2026 shifts the message (not just content) - Cross-Embodiment Q1: fold in Galaxea G0 counter-evidence — single-embodiment data consistency is a third lever alongside scale-for-coverage / architecture- for-extrapolation (raw cross-embodiment scale is not automatically better) - VLM-Action trends: add trend #7 — by ICRA 2026 the mechanism set has stabilized (no 8th coupling primitive); frontier moved from inventing wirings to deploying frozen ones (FD-VLA, G0, FPO) - System-0/1/2 TL;DR: dual-system is now the label-free systems-community default (Galaxea G0, DualVLN) Other reviews' theses are reinforced (not changed) by ICRA — left as-is. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Integrate ICRA 2026 into cross-paper in-depth reviews Wove ICRA 2026 developments into 9 cross-paper reviews (drawing only from the already-verified ICRA topic/per-paper pages — no new web claims), mapping ICRA papers onto each review's existing taxonomy: - VLA-Architecture, Dexterous-Manipulation, RL, VLA-Memory, Goal-Image-Conditioning, System-0-1-2, Cross-Embodiment, VLM-Action-Connection, WAM-vs-VLA-Robustness Six of these had zero ICRA content before. 0 dangling wikilinks. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Source-verification audit: fix fabrications, rename PI slugs, link hygiene, add foundational pages - Re-verified all 195 pages vs original sources; removed 47 confirmed fabricated tables/numbers, restored 19 false-positive deletions (full-PDF re-check) - Renamed PI tech-report slugs ICLR-2026-pi07/pi06/RECAP -> PI-pi07/PI-pi06/PI-RECAP (these are PI technical reports, not ICLR 2026 papers); updated 171 wikilinks - Fixed 60 broken wikilinks -> 0 dangling across the wiki - Added foundational pages: OpenVLA, ReKep, AgiBot World Colosseo, RoboBrain 2.0; linked from Home + sidebar + venue indexes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Add VLM-Action Connection cross-paper review Review-VLM-Action-Connection.md re-slices the VLA architecture space along a different axis from Review-VLA-Architecture: not "what kind of action decoder" but "how does the VLM information flow into the action path." 7 mechanisms with concrete technical details: 1. Unified token stream (OpenVLA, VLA-0, DDVLA, UDVLA, dVLA, HybridVLA) — no interface 2. Same-stack MoE + prefix-KV attention (pi0/pi0.5/pi0.6/pi0.7, FLOWER) — action expert is separate weights in the same transformer stack; matched head-dim + layer count; VLM prefix KV is cached; Knowledge Insulation is a gradient-direction modifier, NOT a new interface 3. Cross-attention into VLM hidden states (GR00T N1 at LAYER 12 of Eagle-2, RDT-1B with Alternating Condition Injection, ST4VLA over k intermediate layers, RetoVLA with register-token KV) 4. Latent condition token (ThinkAct, RoboDual, WholeBodyVLA) 5. FiLM / prefix conditioning (CogVLA applies FiLM TWICE — once at the vision encoder, once at the LLM — plus V-L-A Coupled Attn) 6. Parameter sharing at layer boundary (Fast-in-Slow: last 2 of 32 LLaVA blocks, ablated optimum; 1:4 S2:S1 frequency ratio) 7. Outside-the-VLM (RFS residual flow, VITA-VLA reverse distillation via hidden-state alignment, RTC async scheduling) Plus interface-efficiency work: VLA-Cache, VLA-Adapter's Bridge Attention (closest paper to an interface-ablation), VLA-OS paradigm study. Key findings: - The field is bifurcating, not converging: production (same-stack MoE) vs research (unified stream) vs dual-system (latent/embedded). - Nobody has published a matched-compute head-to-head ablation of interfaces on the same backbone and data. This is the single most valuable unpublished study. - VLA-0 (actions as text, zero modification) beats pi0.5-KI, OpenVLA-OFT, GR00T-N1 on LIBERO, suggesting much of the field's interface complexity may be over-engineered. - Which VLM layer cross-attention targets is ad-hoc: GR00T picked layer 12 empirically, ST4VLA picks k layers. Layer-sweep ablation missing. Navigation: - _Sidebar.md Cross-paper reviews: add VLM-Action Connection - Home.md Cross-paper reviews table + What's-new surface the addition Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Apr 18, 2026