Skip to content

History / Reviews

Revisions

  • Split Reviews into topic-reviews (Reviews) + per-paper long-forms (Reviews-Per-Paper) Durable fix for the recurring 'too long to render' on the growing catalog: - Reviews.md keeps cross-paper topic reviews, lab/series programs, latest-paper reviews (78 links). - New Reviews-Per-Paper.md holds all single-paper long-forms — architecture/ runtime, data/training, world-models/tactile, hybrid-MoT, dexterous-hand data, multi-task/in-context, IROS 2026 full-paper analyses, RSS pointer (56 links). - Both now safely under GitHub-wiki's ~100-link render limit. - Linked Per-Paper from Home and the sidebar; added the IROS 2026 + hybrid + dexterous-data long-forms that weren't catalogued before. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 12, 2026
  • Fix Reviews (In-depth reviews) 'too long to render': cut duplicate RSS per-paper link block Reviews.md crept to 103 wikilinks (over GitHub-wiki's ~100 catalog render limit) as IROS/in-context sub-entries were added. The 'RSS 2026 per-paper pages' block (16 links) duplicated the RSS survey's own catalog, so replaced it with a one- line pointer to the RSS survey + RSS-2026-Papers index. Total links 103 -> ~88, safely under the limit. No unique navigation lost (RSS pages reachable via the survey). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 7, 2026
  • Fix 'too long to render' on Review-World-Models and Reviews: split link-dense mega-lines Both pages hit GitHub-wiki's render budget after recent growth. Root cause was single lines packed with 15-17 wikilinks (O(n^2) emphasis/link parsing) — the same trigger as the earlier RSS survey. Split the pure link-list lines into shorter lines (content and all links preserved; max wikilinks-per-line 17->7 in World-Models, 15->6 in Reviews). No links removed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 7, 2026
  • Add in-depth survey: In-Context Imitation & Demo-Following Review-In-Context-Imitation: watch-a-demo-and-reproduce-it (no per-task FT) as a memory-conditioning problem. Taxonomy by how the demo is conditioned — (A) cross-attention/video-conditioned (Vid2Robot, VLBiMan, See-Once-Then-Act), (B) recurrent/query memory (HAMLET, MemoryVLA, RememVLA, ContextVLA), (C) fast-weight/TTT (RoboTTT), (D) retrieval (MemER, MAP-VLA, Memory-Retrieval, KEMO, Long-Context-IL), (E) token-sequence ICL (ICRT, Behavior Prompting), (F) play-video ICL (MimicDroid). Mapped to RoboMME's Imitation (procedural memory) suite; comparison table + design axes + open challenges. Cross-linked from Reviews and Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 31, 2026
  • Add in-depth survey: Egocentric Video for VLA Pre-Training Review-Egocentric-Video-Pretraining: how label-free first-person human video becomes a pretraining signal. Taxonomy of methods (A pseudo-action extraction: hand-pose/keypoints/optical-flow/part-motion; B latent-action models; C world-model/video-prediction; D reconstruct-then-retarget; E auxiliary-modality recovery), the dataset landscape (EgoDex, EgoVerse, EgoScale, Being-H0, DYNA-2, DreamDojo, UniDex, EgoVLA), scaling-law evidence (EgoScale R2=0.9983 +54%, DYNA-2 transfer law), the transfer question (emergence/decoupling/synthesis + two-stage recipe), and open challenges. Cross-linked from Reviews and Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 25, 2026
  • Add in-depth review of NVIDIA RoboTTT (context scaling via TTT in GR00T N1.7) Review-RoboTTT (arXiv 2607.15275, NVIDIA GEAR + Stanford + UT Austin): 8K-timestep visuomotor context at constant latency by adding Test-Time- Training fast-weight layers to GR00T N1.7's DiT action head. Detailed GR00T implementation section: TTT layer after self/cross-attn in each of 16 DiT layers (~10M each -> ~690M), tanh-gated to preserve pretrained skills; register tokens (N=16) carry compressed VL history through TTT while VL tokens bypass; fast weights = 2-layer GeLU MLP updated per step (W_t <- W_{t-1} - eta*grad MSE(f(K),V)), read via Q; training recipe = flow matching + sequence action forcing (per-step tau) + TBPTT (fast weights carried, gradients detached at segment boundaries); 30 Hz on RTX 5090 (YAM bimanual). Results, new capabilities (one-shot in-context video imitation, DAgger- distillation self-improvement, perturbation robustness), limitations. Cross-linked from Reviews, Home lab programs, and Review-GR00T-Series. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 19, 2026
  • Add in-depth Review pages for HiMoE-VLA and DyGRO-VLA (multi-task VLA cluster) - Review-HiMoE-VLA: hierarchical depth-wise MoE (shallow AS-MoE per action space, deeper HB-MoE per embodiment, central dense consolidation; AS-Reg/ HB-Reg; 32 experts top-4). Negative-transfer ablation (dense pi0 -0.259 vs HiMoE +0.186), LIBERO 97.8 / real xArm7 75.0 / ALOHA 63.7. Framed as the MoE-routing cluster of multi-task fixes. - Review-DyGRO-VLA: cross-task RL scaling — protect a shared latent (offline info-theoretic pretrain) then optimize grouped RL residuals (a=Δa+a_base, K-critic ensemble, entropy load-balance). LIBERO 92.7->97.1, Long 85.2->95.0. - Repointed Multi-Task VLA links to the new Review pages; added in-depth pointers to the ICLR/ICML venue entries; catalog entries in Reviews. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 14, 2026
  • Add Multi-Task VLA review + MergeVLA page - Review-Multitask-VLA: why one VLA fails across many tasks (6 failure modes: negative transfer/gradient conflict, non-mergeability, multi-task conflict, catastrophic forgetting, routing confusion, instruction collapse) + a 7-cluster solution landscape (merging, MoE, gradient control, skill decomposition, instruction grounding, continual, adapters) with a decision guide and open questions. Anchored on MergeVLA (CVPR 2026). - CVPR-2026-MergeVLA: per-paper page (non-mergeability diagnosis + task-masked LoRA / cross-attention-only action expert / test-time task router). - Cross-linked from Reviews catalog and Home topic reviews. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 14, 2026
  • Add in-depth page for Genesis AI GENE-26.5 (glove-first dexterous FM, vendor) Review-Genesis-GENE: Genesis AI's May 2026 launch. 1:1:1 tactile-glove <-> human-hand <-> robot-hand co-design (collapses L4 retargeting), data engine (glove demos + ego + internet video + simulation), sub-hour robot fine-tune (secondary-sourced). Flagged throughout as vendor claims: no paper/weights/ benchmarks, undisclosed hand DoF and architecture, unquantified capabilities. Linked from the Dex-Hand Data Pyramid (matrix, 3b, links) and Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add in-depth pages for T-Rex and RLDX-1 (tactile-reactive + industry dexterity FM) - Review-T-Rex (arXiv 2606.17055): variable-rate Mixture-of-Transformers, slow Action Expert + fast Tactile Expert; temporal tactile VQ-VAE; Dexmate Vega-1 + 2x Sharpa Wave 22-DoF, 5 fingertip sensors, 300 Hz; 12 tasks 65% avg vs EgoScale 35% (+30pts); 100h/22-primitive dataset. - Review-RLDX-1 (RLWRLD, Seoul, vendor): dexterity-first foundation model, Multi-Stream Action Transformer (vision/motion/memory/torque), bare- human-hand + five-finger retargeting (>200 demos/hr) + synthetic aug; ALLEX/Franka/OpenArm; vendor benchmarks vs pi0.5/GR00T N1.6 (flagged as company claims, not peer-reviewed). - Linked both from the Dex-Hand Data Pyramid (matrix + tactile/industry deep-dives) and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add in-depth pages for YUBI and DexEXO (dexterous capture interfaces) - Review-YUBI (arXiv 2606.10244, AIST/UTokyo consortium): handheld finger- driven bidigital gripper, VR 6-DoF tracking, mount-on-robot deploy with NO retargeting; 8,434 h / 1.20M ep / 119 tasks bimanual dataset, single policy across UR/Franka/ELEY. Flagged 2-finger (not 5-finger) caveat. - Review-DexEXO (arXiv 2603.17323, UCLA RoMeLa): wearability-first finger exoskeleton, visual-match to 6-DoF OYMotion ROHand, direct wrist-cam RGB policy training; beats DexUMI/teleop (scissors 0.79 vs 0.00; piano 0.96 vs 0.62); operator-agnostic 140-217mm, n=14 user study; no tactile. - Linked both from the Dex-Hand Data Pyramid (matrix rows split by L4 behavior + capture-interface deep-dives) and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add in-depth pages for Do-As-I-Do and AnyDexRT (dexterous retargeting) - Review-DoAsIDo (arXiv 2606.19333, Berkeley/NYU): device-free monocular video -> dexterous data. HaWoR+SAM-3D+guided-diffusion 4D reconstruction, dynamics-aware MPPI retargeting in MuJoCo Warp (warmup/force-perturb/ transition-reward) on 22-DoF Sharpa Wave; in-the-wild 25->71%, OakInk2 72->81%; 500 verified trajectories; 4% clip survival honesty. - Review-AnyDexRT (arXiv 2607.08341, SJTU/Cewu Lu): calibration-free cross-hand retargeting. Self-supervised fingertip mapper (partial-Chamfer + distance + local-motion preservation) + few-shot anchors + pinch contact classifier; 7 hands, LMC 90.2% / pinch 62.0% vs GeoRT, 293 Hz, stable under +-90deg rotations. - Linked both from the Dex-Hand Data Pyramid (matrix + L4 bridge sections) and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add Review-Dexterous-Hand-Data-Pyramid: data types + approach×data matrix + current/future insight New page classifying data-acquisition/generation methods for humanoid 5-finger dexterous hands as a 6-tier data pyramid (web-video → egocentric → wearable glove/exo → retargeting → sim/synthetic → target-hand teleop), with tactile/force as a cross-cutting axis. Includes: - per-data-type description + pros/cons table - ONE summary matrix: approach (paper/company) × data-type used (EgoScale, DYNA-2, UniDex, Being-H0, DO-AS-I-DO, DexUMI, YUBI/DexEXO, AnyDexRT, Dex1B, DexGrasp-Zero, DexNDM, shared-autonomy, MANUS teleop, RLDX-1, pi0.5/0.7, T-Rex, One-Hand) - current most-common recipe (teleop+retarget+sim) and future directions (human-video scaling laws, device-free RGB, calib-free retargeting, tactile-first, cross-hand co-design, unified WAM) w/ latest 2026 papers. Cross-linked from Review-Dexterous-Manipulation, Reviews, Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add Review-Single-Checkpoint-Multi-Robot: deploy-side cross-embodiment (one frozen checkpoint, many robots) New page separating the DEPLOYMENT question (one unchanged checkpoint controls multiple physical robots at inference) from the training-data cut in Review-Cross-Embodiment. Organized by three tiers: - Tier 2 (zero-shot to unseen robot): LAP-3B (actions-as-language, first substantial zero-shot to unseen), Green-VLA, Gemini Robotics 1.5, DreamZero/DYNA-2 (WAM), Contact-Anchored Policies, One-Hand. - Tier 1 (routed seen-robot generalist): RT-X, CrossFormer, RDT-1B, UniAct, GR00T N1, pi0.5->pi0.7, Motus. - Tier 3 (contrast, needs per-robot fit): Octo, HPT, X-VLA; plus MergeVLA (merge specialists into one checkpoint). Includes a routing-mechanism taxonomy, honest limits (AnyBody, unreplicated 2026 zero-shot claims, 'single checkpoint != nothing per robot'), and a design guide. Cross-linked from Review-Cross-Embodiment, Reviews, Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 12, 2026
  • Add in-depth pages for HALO and BAGEL; add architecture diagram to Motus - Review-HALO (arXiv 2602.21157, ICML'26): three-expert MoT (AR understanding + diffusion visual-subgoal + flow-matching action), EM-CoT think→imagine→act, ~4.5B (Qwen2.5-1.5B×3); RoboTwin2.0 Easy 80.5% (+34.1 over pi0); full ablation analysis isolating the vision tower's value. Cross-linked with ICML-2026-HALO. - Review-BAGEL (arXiv 2505.14683, ByteDance): the base unified multimodal MoT recipe (understanding+generation, VAE+ViT dual encoders, separate QKV/FFN + shared attention, 7B/14B) that BagelVLA/HALO inherit; framed for VLA relevance. - Review-MOTUS: added mermaid architecture diagram (4 towers + scheduler). - Linked both from Review-VLA-Hybrid-Architectures core table/§8, ICML-2026-HALO, and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add Review-VLA-Hybrid-Architectures: three-expert MoT (vision as separate tower) comparison + design guide New themed in-depth page on the WAM+VLA hybrid convergence, centered on the papers that split vision into its own expert tower in a three-expert Mixture-of-Transformers: - Core: Motus (understanding+video-gen+action, 8B, open), BagelVLA (LLM+ generation+action, RFG 1-step foresight), HALO (reasoning+visual- foresight+action, +34.1% over pi0), BAGEL (base MoT recipe: VAE+ViT dual vision encoders, separate QKV/FFN + shared attention). - Variants contrasted: DYNA-2, Being-H0.7, Cosmos 3, omega-0, LingBot-VLA. - Comparative analysis across 5 axes (vision granularity, per-tower mechanism, inference contract, action-from-video supervision, scale/ openness) + pros/cons of splitting vision + a 6-step design guide + open questions. Cross-linked from Review-VLA-Architecture §4.2b and the Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add in-depth pages for Motus, Cortex 2.0, Being-H0.7 (WAM+VLA hybrids) Three new source-verified reviews of the WAM+VLA hybrid exemplars named in NVIDIA's WAM blog: - Review-MOTUS: THU-ML unified scheduled MoT (8B), optical-flow latent actions, RoboTwin2.0 88.66% vs X-VLA/pi0.5, open weights (arXiv 2512.13030) - Review-Cortex2: Sereact modular foresight-planning hybrid (WM+PRO scorer +flow head, 30Hz), deployment-reported vs pi0.5 (arXiv 2604.20246) - Review-Being-H07: BeingBeyond unified+reactive latent world-action model, posterior/prior latent queries, no rollout at inference (arXiv 2605.00078) Cross-linked from Review-VLA-Architecture §4.2b/§7, Review-NVIDIA-WAM- Cosmos3 §2.6, Latest-Papers, and Reviews catalog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Fix broken wikilinks in Latest-Papers table (escape pipe for GitHub-wiki table cells); surface latest-paper reviews in Reviews topic sections Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add DYNA-2 in-depth review (Dyna Robotics World-Action Model launch) Company announcement (Aug 10, 2026), not a paper: WAM on ~1M h human egocentric video with no robot data in pre-training, joint next-frame+ next-action, claimed first human-to-robot scaling law smooth over 1k->1M h (~50x EgoScale), 87% vs 46% zero-shot over DYNA-1. Reviewed with an explicit vendor-claim caveat (no technical paper/benchmark/ weights). Filed under Latest Papers; cross-linked from World-Models, Human-Video-Transfer, Reviews. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add Latest Papers tracker + omega-0 and Stellar VLA in-depth reviews New Latest-Papers.md preprint tracker (pre-publication reviews) and two figure-illustrated in-depth reviews: omega-0 (arXiv 2608.06375, whole- body humanoid latent-predictive World Action Model; 81.8% on 11 household tasks vs 44.5% psi-0; ships 40h omega-HOME dataset) and Stellar VLA (arXiv 2511.18085, continual imitation learning with a Dirichlet-Process knowledge space + knowledge-routed MoE, 1% replay). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA, World-Models. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add DreamZero in-depth review (World Action Models are Zero-shot Policies) NVIDIA's 14B video-diffusion World Action Model (arXiv 2602.15922): jointly predicts video+action, >2x over SOTA VLAs on unseen-env/ unseen-object real-robot evals, 38x inference stack (DreamZero-Flash) for 7 Hz closed-loop control, video-only cross-embodiment transfer. Fig. 4 architecture embedded. Cross-linked from World Models review, Home lab-programs, Reviews catalog, and sidebar. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 10, 2026
  • Survey-depth topic pages: 12 State-of-the-Field updates + 3 new surveys Each decision-map topic's detail page now carries a dated July-2026 survey section: trend arc through the latest venues, approach taxonomy with definitions and trade-offs, and current limitations. Three previously page-less topics get dedicated surveys: Human-Video Transfer (emergence/decoupling/synthesis fork + decision guide), VLA Evaluation (indictment + 2026 toolkit + emerging norms), Real-Time Execution (RTC->Legato arc + approach comparison). Home fold-outs link the full surveys; Reviews catalog updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 27, 2026
  • Restructure navigation: top-level sidebar + Reviews catalog page New Reviews.md holds the complete in-depth-review catalog (topic reviews, lab programs, per-paper long-forms, RSS 2026 figure pages). Sidebar slimmed to top-level only: reviews hub + six star topics, model lineages, ML hub, one link per venue year (venue pages already index their papers), foundational refs. Home merges its two review sections into one compact section pointing at the catalog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 25, 2026