Skip to content

History / Review World Models

Revisions

  • Fix 'too long to render' on Review-World-Models and Reviews: split link-dense mega-lines Both pages hit GitHub-wiki's render budget after recent growth. Root cause was single lines packed with 15-17 wikilinks (O(n^2) emphasis/link parsing) — the same trigger as the earlier RSS survey. Split the pure link-list lines into shorter lines (content and all links preserved; max wikilinks-per-line 17->7 in World-Models, 15->6 in Reviews). No links removed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 7, 2026
  • IROS 2026: add 5 full-paper analyses + reflect into existing in-depth reviews Full-paper pages (abstract-verified from official program): - IROS-2026-AtomVLA: subtask-aware VLA + latent-WM scoring for offline GRPO (WAM lens) - IROS-2026-3D-FlowMatch-Actor: CMU/NVIDIA unified single/dual-arm 3D policy, +41.4% PerAct2, ~30x faster (bimanual SOTA) - IROS-2026-EquiBim: symmetry-equivariant bimanual policy - IROS-2026-IMLE-VLA: single-step cIMLE action head, 55Hz, LIBERO 98.0% (efficiency) - IROS-2026-ICLR-Visual-Reasoning: in-context imitation with image-space reasoning traces Reflected IROS 2026 into existing reviews: - Review-World-Models: WAM-as-critic row (AtomVLA offline GRPO) - Review-In-Context-Imitation: ICLR-visual-reasoning + RoboSSM - Review-VLA-Memory: structured-vs-parametric memory row (GaussMemory/PROMPT vs TempoFit/RoboSSM) - Review-Multitask-VLA: VLA-RL / LAR-MoE / AtomVLA / MoE-humanoid - Review-Realtime-Execution: single-step head row (IMLE-VLA) - Review-Humanoid-VLA: IROS bimanual/whole-body trend (3DFA/EquiBim/ULTRA/CEER/MoE-VLA) Linked all from the IROS 2026 survey. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 5, 2026
  • DYNA-2: add detailed architecture, WAM comparison, and insights sections - Detailed model structure (mixture-of-transformers, hand-pose pseudo- actions, flow-matching co-training, mermaid diagram) with the key structural fact: video prediction is a co-training objective DROPPED at inference (action head never sees z_t) -> reactive real-time policy. - New 'How DYNA-2 differs from other WAMs' section: comparison table + three axes (data purity, world-model-at-inference reactive vs co-generate, fitted transfer law vs ablation) vs DreamZero/omega-0/ Cosmos-Policy/DreamGen/DreamDojo/EgoScale. - New 'Key insights' section (39/39 future-pred ablation, human-video vs teleop, free-at-inference world-modeling, threshold emergence, authors' own lower-bound/compute caveats). - Refined DYNA-2 entry in Review-World-Models to note reactive decoupling. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Escape pipe in all in-table wikilinks wiki-wide (202 links, 17 files) GitHub-wiki table cells read a wikilink's separator | as a column delimiter, splitting the cell and breaking the link. Escape to \| in every table-row wikilink (Home nav, RSS-2026-Papers, topic surveys). Prose wikilinks left as plain | (render correctly outside tables). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add DYNA-2 in-depth review (Dyna Robotics World-Action Model launch) Company announcement (Aug 10, 2026), not a paper: WAM on ~1M h human egocentric video with no robot data in pre-training, joint next-frame+ next-action, claimed first human-to-robot scaling law smooth over 1k->1M h (~50x EgoScale), 87% vs 46% zero-shot over DYNA-1. Reviewed with an explicit vendor-claim caveat (no technical paper/benchmark/ weights). Filed under Latest Papers; cross-linked from World-Models, Human-Video-Transfer, Reviews. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add Latest Papers tracker + omega-0 and Stellar VLA in-depth reviews New Latest-Papers.md preprint tracker (pre-publication reviews) and two figure-illustrated in-depth reviews: omega-0 (arXiv 2608.06375, whole- body humanoid latent-predictive World Action Model; 81.8% on 11 household tasks vs 44.5% psi-0; ships 40h omega-HOME dataset) and Stellar VLA (arXiv 2511.18085, continual imitation learning with a Dirichlet-Process knowledge space + knowledge-routed MoE, 1% replay). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA, World-Models. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add DreamZero in-depth review (World Action Models are Zero-shot Policies) NVIDIA's 14B video-diffusion World Action Model (arXiv 2602.15922): jointly predicts video+action, >2x over SOTA VLAs on unseen-env/ unseen-object real-robot evals, 38x inference stack (DreamZero-Flash) for 7 Hz closed-loop control, video-only cross-embodiment transfer. Fig. 4 architecture embedded. Cross-linked from World Models review, Home lab-programs, Reviews catalog, and sidebar. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 10, 2026
  • Weave ICML 2026 evidence into deep-dive surveys; revise three verdicts 14 State-of-the-Field sections gain ICML 2026 findings from the 99-paper index: recipes and latent-action supervision (VLANeXt, From-Pixels-to-Tokens, XR-1), MoT dual-systems and shortcut counters, the 9-paper efficiency cluster (Reflex 50Hz, GridS -76% FLOPs, XPU profile, latent reasoning -90%), reward/critic and model-based RL (VLAC, VLAW +39.2%), memory (HiMe/SOMA/CAPS), world models (DreamDojo 44kh, LAC-WM, dWorldEval), dexterous (DexMachina/DECO/Tabero/CTSRL), cross-embodiment (OXE-AugE, latent motion codes), evaluation (LIBERO-Gen, VLA-Arena, FixBench, TRAP). Verdicts revised: forgetting milder than assumed; discrete-token verdict scoped to robot-action auxiliaries; WM-evaluator action gap first crack. Home synced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Decision map: prominent deep-dive links + uniform detail-page template Home fold-outs now lead with a heading-level "Deep dive ->" link and compress trend/approaches/limitations into a labeled 3-row table. The 12 detail-page State-of-the-Field sections are rewritten to one template (Verdict quote + Trend + Approaches-and-trade-offs + optional Established-findings + Limitations, dated Aug 2026); the three standalone surveys get matching headers with structure legends. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Survey-depth topic pages: 12 State-of-the-Field updates + 3 new surveys Each decision-map topic's detail page now carries a dated July-2026 survey section: trend arc through the latest venues, approach taxonomy with definitions and trade-offs, and current limitations. Three previously page-less topics get dedicated surveys: Human-Video Transfer (emergence/decoupling/synthesis fork + decision guide), VLA Evaluation (indictment + 2026 toolkit + emerging norms), Real-Time Execution (RTC->Legato arc + approach comparison). Home fold-outs link the full surveys; Reviews catalog updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 27, 2026
  • Fix broken tables: escape unescaped pipes inside wikilinks in table rows ICLR.md (and 12 other pages) had [[label|Page]] wikilinks with raw pipes inside GFM table cells, which the GitHub-wiki renderer reads as column separators — mangling the table. Escaped the wikilink-internal pipes to \| (matching the convention already used by CVPR/NeurIPS/CoRL pages); table column separators left intact. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 11, 2026
  • Add in-depth review: OmniVTA (visuo-tactile world modeling) + 2026 tactile-VLA analysis New Review-OmniVTA page (arXiv 2603.19201): 4-module visuo-tactile world model (TactileVAE + two-stream contact-evolution predictor + contact-aware fusion policy + 60Hz Reflexive Latent Tactile Controller), OmniViTac dataset (21k+ traj / 86 tasks / 100+ objects), GelSight Mini; beats Diffusion Policy / FoAR, closed-loop >> open-loop. Includes a 4-camp analysis of 2026 tactile-VLA research (ICRA 2026: FD-VLA, SaTA, ManipForce, TranTac, FreeTacMan, Multi-Modal Consensus, DOT-Sim; CVPR 2026: HapticVLA), framing the sensor-in-loop (OmniVTA) vs sensor-free (FD-VLA/HapticVLA) fault line. Adds a new "predictive reference for reflexive control" world-model role. Linked from Dexterous/Architecture/World-Models reviews + Home + sidebar. 0 dangling links. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 6, 2026
  • World Models review: add Family F — WM + inverse-dynamics action decoding New §3 model family (compositional, cuts across A–D): world model predicts the future, an inverse-dynamics / latent-action model decodes the action. - generate-then-decode (UniPi, HiP, VLP, RoboDreamer; DreamGen data-factory use) - disentangled forward+inverse pretraining (DeFI) - latent-action IDM from action-free video (villa-X, UniVLA, ViPRA, Human-Video-Pretraining) Pros (decouple what-to-do from how-to-act; action-free/cross-embodiment pretraining) + cons (inverse-dynamics drift → shift to co-generation / goal-pose / π0.7 token-direct). Added matching row to the §7 comparison table. 0 dangling links. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Add in-depth review: World Models for Robot Learning New cross-paper, model-centric review (sibling to Review-VLA-Architecture's Category E usage taxonomy). Two axes — what the WM predicts (pixel video-diffusion / AR-token / latent-JEPA / 3D-4D-geometry / structured-cue) x how robotics uses it (backbone / RL-env / data-factory / planner / evaluator / aux-loss). Synthesizes in-wiki pages (Cosmos-Policy, Genie-Envisioner, WMPO, Ctrl-World, WorldGym, DreamGen, VLA-RFT, DreamVLA, Geometry-4D, ...) + external landscape (Cosmos, Genie 3, V-JEPA 2, DINO-WM, NWM). Latest trends, pros/cons per family, six tensions, decision guide. Linked from Home + sidebar. 0 dangling links. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026