Skip to content

History / Home

Revisions

  • Add CoRL 2026 survey (preliminary, community-sourced; official list pending) CoRL 2026 (Austin, Nov 9-12; 687 accepted / 32.8% / 2,094 submitted; official per-paper program not yet public). Preliminary manipulation-centric survey with explicit acceptance-confidence marking: - Confirmed CoRL 2026: SG-WAM (geometry-aware WAM), FiberTune (robustness- preserving VLA fine-tune), Dex-X (visual-tactile from human video via sim), Touch2Trace (tactile IL, cable tracing). - Reported/unverified: StellaVLA (in-context VLA), Choice Policies (Berkeley/ Malik, whole-body humanoid), Weave (whole-body dexterous loco-manip). - Excluded ManiFlow (it's CoRL 2025, not 2026). Prominent caveat + confidence legend; to be promoted to a full session-taxonomy survey when the program publishes. Linked from CoRL hub, Home venue table, sidebar. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 28, 2026
  • Split Reviews into topic-reviews (Reviews) + per-paper long-forms (Reviews-Per-Paper) Durable fix for the recurring 'too long to render' on the growing catalog: - Reviews.md keeps cross-paper topic reviews, lab/series programs, latest-paper reviews (78 links). - New Reviews-Per-Paper.md holds all single-paper long-forms — architecture/ runtime, data/training, world-models/tactile, hybrid-MoT, dexterous-hand data, multi-task/in-context, IROS 2026 full-paper analyses, RSS pointer (56 links). - Both now safely under GitHub-wiki's ~100-link render limit. - Linked Per-Paper from Home and the sidebar; added the IROS 2026 + hybrid + dexterous-data long-forms that weren't catalogued before. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 12, 2026
  • Add IROS 2026 VLA & Manipulation survey (official program, pre-conference) IROS-2026-VLA-Manipulation-Survey: built from the public official program (2026.ieee-iros.org, 1,900+ papers, Pittsburgh Sep 27-Oct 1). Session-level taxonomy of ~40 in-scope manipulation/VLA/dexterous/imitation sessions + sampled papers extracted per theme (VLA: AnyCamVLA/LangGap/OG-VLA/BFA++/ Safe-Night-VLA; scaling/RL: VLA-RL/LAR-MoE/flow-matching; in-context imitation: RoboSSM/ICLR-visual-reasoning/IMLE-VLA/ESPADA; policy: MaskVLA/SynthLA; language-in-loop: NL2SpaTiaL/AURORA). Clearly flagged title-level (abstracts pending Xplore post-conference), not abstract-verified. Added to Home venue table + top pointer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 5, 2026
  • Add in-depth survey: In-Context Imitation & Demo-Following Review-In-Context-Imitation: watch-a-demo-and-reproduce-it (no per-task FT) as a memory-conditioning problem. Taxonomy by how the demo is conditioned — (A) cross-attention/video-conditioned (Vid2Robot, VLBiMan, See-Once-Then-Act), (B) recurrent/query memory (HAMLET, MemoryVLA, RememVLA, ContextVLA), (C) fast-weight/TTT (RoboTTT), (D) retrieval (MemER, MAP-VLA, Memory-Retrieval, KEMO, Long-Context-IL), (E) token-sequence ICL (ICRT, Behavior Prompting), (F) play-video ICL (MimicDroid). Mapped to RoboMME's Imitation (procedural memory) suite; comparison table + design axes + open challenges. Cross-linked from Reviews and Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 31, 2026
  • Add in-depth survey: Egocentric Video for VLA Pre-Training Review-Egocentric-Video-Pretraining: how label-free first-person human video becomes a pretraining signal. Taxonomy of methods (A pseudo-action extraction: hand-pose/keypoints/optical-flow/part-motion; B latent-action models; C world-model/video-prediction; D reconstruct-then-retarget; E auxiliary-modality recovery), the dataset landscape (EgoDex, EgoVerse, EgoScale, Being-H0, DYNA-2, DreamDojo, UniDex, EgoVLA), scaling-law evidence (EgoScale R2=0.9983 +54%, DYNA-2 transfer law), the transfer question (emergence/decoupling/synthesis + two-stage recipe), and open challenges. Cross-linked from Reviews and Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 25, 2026
  • Add in-depth review of NVIDIA RoboTTT (context scaling via TTT in GR00T N1.7) Review-RoboTTT (arXiv 2607.15275, NVIDIA GEAR + Stanford + UT Austin): 8K-timestep visuomotor context at constant latency by adding Test-Time- Training fast-weight layers to GR00T N1.7's DiT action head. Detailed GR00T implementation section: TTT layer after self/cross-attn in each of 16 DiT layers (~10M each -> ~690M), tanh-gated to preserve pretrained skills; register tokens (N=16) carry compressed VL history through TTT while VL tokens bypass; fast weights = 2-layer GeLU MLP updated per step (W_t <- W_{t-1} - eta*grad MSE(f(K),V)), read via Q; training recipe = flow matching + sequence action forcing (per-step tau) + TBPTT (fast weights carried, gradients detached at segment boundaries); 30 Hz on RTX 5090 (YAM bimanual). Results, new capabilities (one-shot in-context video imitation, DAgger- distillation self-improvement, perturbation robustness), limitations. Cross-linked from Reviews, Home lab programs, and Review-GR00T-Series. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 19, 2026
  • Add Multi-Task VLA review + MergeVLA page - Review-Multitask-VLA: why one VLA fails across many tasks (6 failure modes: negative transfer/gradient conflict, non-mergeability, multi-task conflict, catastrophic forgetting, routing confusion, instruction collapse) + a 7-cluster solution landscape (merging, MoE, gradient control, skill decomposition, instruction grounding, continual, adapters) with a decision guide and open questions. Anchored on MergeVLA (CVPR 2026). - CVPR-2026-MergeVLA: per-paper page (non-mergeability diagnosis + task-masked LoRA / cross-attention-only action expert / test-time task router). - Cross-linked from Reviews catalog and Home topic reviews. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 14, 2026
  • Add Review-Dexterous-Hand-Data-Pyramid: data types + approach×data matrix + current/future insight New page classifying data-acquisition/generation methods for humanoid 5-finger dexterous hands as a 6-tier data pyramid (web-video → egocentric → wearable glove/exo → retargeting → sim/synthetic → target-hand teleop), with tactile/force as a cross-cutting axis. Includes: - per-data-type description + pros/cons table - ONE summary matrix: approach (paper/company) × data-type used (EgoScale, DYNA-2, UniDex, Being-H0, DO-AS-I-DO, DexUMI, YUBI/DexEXO, AnyDexRT, Dex1B, DexGrasp-Zero, DexNDM, shared-autonomy, MANUS teleop, RLDX-1, pi0.5/0.7, T-Rex, One-Hand) - current most-common recipe (teleop+retarget+sim) and future directions (human-video scaling laws, device-free RGB, calib-free retargeting, tactile-first, cross-hand co-design, unified WAM) w/ latest 2026 papers. Cross-linked from Review-Dexterous-Manipulation, Reviews, Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 13, 2026
  • Add Review-Single-Checkpoint-Multi-Robot: deploy-side cross-embodiment (one frozen checkpoint, many robots) New page separating the DEPLOYMENT question (one unchanged checkpoint controls multiple physical robots at inference) from the training-data cut in Review-Cross-Embodiment. Organized by three tiers: - Tier 2 (zero-shot to unseen robot): LAP-3B (actions-as-language, first substantial zero-shot to unseen), Green-VLA, Gemini Robotics 1.5, DreamZero/DYNA-2 (WAM), Contact-Anchored Policies, One-Hand. - Tier 1 (routed seen-robot generalist): RT-X, CrossFormer, RDT-1B, UniAct, GR00T N1, pi0.5->pi0.7, Motus. - Tier 3 (contrast, needs per-robot fit): Octo, HPT, X-VLA; plus MergeVLA (merge specialists into one checkpoint). Includes a routing-mechanism taxonomy, honest limits (AnyBody, unreplicated 2026 zero-shot claims, 'single checkpoint != nothing per robot'), and a design guide. Cross-linked from Review-Cross-Embodiment, Reviews, Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 12, 2026
  • Embed paper architecture figures in Motus/HALO/BAGEL; add hybrid link to Home topic reviews - Motus: embed the paper's tri-expert architecture figure (Video Gen / Action / Understanding + Tri-modal Joint Attention), arXiv 2512.13030. - BAGEL: embed the MoT figure (Und/Gen experts + Multi-modal Self-Attention, dual encoders), arXiv 2505.14683. - HALO: embed Fig.1 (three-expert MoT) and Fig.2 (EM-CoT data pipeline). All figures credited to the authors alongside the schematic mermaids. - Home: add [[VLA Hybrid Architectures]] to the Topic reviews row. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Escape pipe in all in-table wikilinks wiki-wide (202 links, 17 files) GitHub-wiki table cells read a wikilink's separator | as a column delimiter, splitting the cell and breaking the link. Escape to \| in every table-row wikilink (Home nav, RSS-2026-Papers, topic surveys). Prose wikilinks left as plain | (render correctly outside tables). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add Latest Papers tracker + omega-0 and Stellar VLA in-depth reviews New Latest-Papers.md preprint tracker (pre-publication reviews) and two figure-illustrated in-depth reviews: omega-0 (arXiv 2608.06375, whole- body humanoid latent-predictive World Action Model; 81.8% on 11 household tasks vs 44.5% psi-0; ships 40h omega-HOME dataset) and Stellar VLA (arXiv 2511.18085, continual imitation learning with a Dirichlet-Process knowledge space + knowledge-routed MoE, 1% replay). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA, World-Models. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 11, 2026
  • Add DreamZero in-depth review (World Action Models are Zero-shot Policies) NVIDIA's 14B video-diffusion World Action Model (arXiv 2602.15922): jointly predicts video+action, >2x over SOTA VLAs on unseen-env/ unseen-object real-robot evals, 38x inference stack (DreamZero-Flash) for 7 Hz closed-loop control, video-only cross-embodiment transfer. Fig. 4 architecture embedded. Cross-linked from World Models review, Home lab-programs, Reviews catalog, and sidebar. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 10, 2026
  • Weave ICML 2026 evidence into deep-dive surveys; revise three verdicts 14 State-of-the-Field sections gain ICML 2026 findings from the 99-paper index: recipes and latent-action supervision (VLANeXt, From-Pixels-to-Tokens, XR-1), MoT dual-systems and shortcut counters, the 9-paper efficiency cluster (Reflex 50Hz, GridS -76% FLOPs, XPU profile, latent reasoning -90%), reward/critic and model-based RL (VLAC, VLAW +39.2%), memory (HiMe/SOMA/CAPS), world models (DreamDojo 44kh, LAC-WM, dWorldEval), dexterous (DexMachina/DECO/Tabero/CTSRL), cross-embodiment (OXE-AugE, latent motion codes), evaluation (LIBERO-Gen, VLA-Arena, FixBench, TRAP). Verdicts revised: forgetting milder than assumed; discrete-token verdict scoped to robot-action auxiliaries; WM-evaluator action gap first crack. Home synced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Decision map: prominent deep-dive links + uniform detail-page template Home fold-outs now lead with a heading-level "Deep dive ->" link and compress trend/approaches/limitations into a labeled 3-row table. The 12 detail-page State-of-the-Field sections are rewritten to one template (Verdict quote + Trend + Approaches-and-trade-offs + optional Established-findings + Limitations, dated Aug 2026); the three standalone surveys get matching headers with structure legends. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 5, 2026
  • Survey-depth topic pages: 12 State-of-the-Field updates + 3 new surveys Each decision-map topic's detail page now carries a dated July-2026 survey section: trend arc through the latest venues, approach taxonomy with definitions and trade-offs, and current limitations. Three previously page-less topics get dedicated surveys: Human-Video Transfer (emergence/decoupling/synthesis fork + decision guide), VLA Evaluation (indictment + 2026 toolkit + emerging norms), Real-Time Execution (RTC->Legato arc + approach comparison). Home fold-outs link the full surveys; Reviews catalog updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 27, 2026
  • Rewrite Home decision map as an insight map Each of the 15 topics is now a collapsible entry carrying the trend arc through the latest venues, competing approaches with definitions and trade-offs, and current limitations - replacing the plain link table. All claims use wiki-verified numbers (OAT/Legato, RECAP, LDA-1B/mimic-video, VLM4VLA +18.1, DexGrasp-Zero/OHRA, Psi-0, LIBERO-X pyramid, LBM verdicts, camera-frame EEF scaling law). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 27, 2026
  • Redesign Home: themed decision map, stat strip; move rules to Maintenance Decision map split into four themed tables (Building / Running & improving / Data & evaluation / Embodiment) with a current-answer column; stat strip and emoji headers added; reviews section as a compact table; Page Format / Maintenance Rule section moved to the new Maintenance page, linked from the footer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 25, 2026
  • Restructure navigation: top-level sidebar + Reviews catalog page New Reviews.md holds the complete in-depth-review catalog (topic reviews, lab programs, per-paper long-forms, RSS 2026 figure pages). Sidebar slimmed to top-level only: reviews hub + six star topics, model lineages, ML hub, one link per venue year (venue pages already index their papers), foundational refs. Home merges its two review sections into one compact section pointing at the catalog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 25, 2026
  • Refresh Home research decision map with RSS 2026 findings Four new question rows (improvement-from-experience, human-video-vs- robot-data fork, co-training data selection, credible evaluation) and five rows updated with RSS evidence (Legato/OAT, LDA-1B/mimic-video, ViTacFormer/CGP, cross-hand transfer, Psi-0). Header stats and start-here pointer updated to the RSS 2026 survey. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 25, 2026
  • Add RSS 2026 conference survey, Psi-0 in-depth, and 14 per-paper pages RSS 2026 (Sydney, Jul 13-17): 210 accepted papers parsed from the official program, ~116 manipulation/hand/humanoid papers in scope, all abstracts verified. Survey covers six threads: RL-from-experience (pi*0.6/RECAP flagship), human-video transfer (emergence vs decoupling), video/world models vs VLA backbones, contact-as-representation, cross-embodiment dexterous hands, and evaluation infrastructure. New: RSS venue hub, Review-Psi0 (full-paper in-depth: 800h human video + 30h robot data beats 10x corpora by >40pp on Unitree G1), per-paper pages for LDA-1B, H2R-Emergence, mimic-video, LBM co-training study, ViTacFormer, DexGrasp-Zero, One-Hand, Contact-Grounded Policy, PolaRiS, LIBERO-X, OAT, Legato, HoMMI. PI-RECAP updated with RSS camera-ready results. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 25, 2026
  • Add in-depth Qwen-RobotNav and Qwen-RobotWorld reviews; update program page RobotNav (2606.18112): parameterized observation interface, 15.6M corpus, agentic EQA SOTA, supersedes Qwen-VLA nav by ~15pp. RobotWorld (2606.17030): 20B double-stream MMDiT, frozen Qwen2.5-VL action encoder, EWK 8.6M corpus, Scene2Robot; 1st on EWMBench/DreamGen. Program page: suite rows upgraded to deep-read, backbone/lambda doctrine qualified, System-2 slot marked demonstrated, watch-list items 4-5 updated; sibling cross-links added. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 24, 2026
  • Add cross-paper review: Qwen Team's VLA Program VLM4VLA -> Qwen-VLA -> Qwen-Robot Suite (RobotManip/RobotNav/RobotWorld): shared doctrine (Qwen3.5-4B, flow matching, lambda=0.1 VL co-training, synthetic-data scaling, language-as-interface), diagnostic-to-flagship trace, internal contradictions between the two flagship VLAs, and lab-program positioning vs PI/GR00T/TRI/Gemini. Cross-links from the three constituent reviews; indexed in Home/Changelog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 24, 2026
  • Add in-depth Qwen-RobotManip review (arXiv 2606.17846) Alignment-first scaling thesis, camera-frame delta EEF + CaPE, 38,100h open-data corpus with 24,808h human-to-robot synthesis, RoboTwin-IF/XE benchmarks, RoboChallenge Table30-v1 generalist #1. Index in Home/Changelog; cross-link from Qwen-VLA review. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jul 24, 2026
  • Add dedicated NVIDIA WAM + Cosmos 3 review; slim §10b to a pointer - New Review-NVIDIA-WAM-Cosmos3: self-contained bundle of the "imagine->act" blog + a detailed Cosmos 3 analysis, with original recreated schematics (mermaid) and data tables (RoboArena, paradigms, tiers, roles, benchmarks). NVIDIA copyrighted figures are linked, not embedded. - WAM review §10b trimmed to a brief summary + pointer to the new page (keeps the review-specific verdict section C). - Wired into Home (Evaluation & Robustness) and the sidebar Per-paper list. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 25, 2026
  • Add cross-topic review: Independent Visual Representation in VLAs New Review-Independent-Visual-Representation: collects research that builds the visual representation OUTSIDE the VLM, motivated by the insufficiency of VLM (language-aligned) vision for action. Taxonomy: G (spatial/geometric), P (predictive/world-model latents), M (temporal memory), D (dense/SSL encoders) x a 5-rung injection ladder (alignment loss -> fused tokens -> AdaLN modulation -> latent interface -> replace-VLM/geometry-first). Wired into Home (reviews list + nav table) and the sidebar. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 20, 2026
  • Add Humanoid VLA + FTP-1 reviews; update Tactile & Dexterous cross-reviews - New: Review-Humanoid-VLA (whole-body/bipedal loco-manip: balance, high-DoF action space, data collection) and Review-FTP-1 (cross-sensor tactile foundation policy). - Tactile review: integrate FTP-1 as new taxonomy axis G (cross-sensor), third fault line, comparison table, limitations. - Dexterous review: add ICML 2026 cluster (§4c) + ICLR 2026 completeness pass (DemoGrasp, DexMove, D-REX, OmniReset, House of Dextra). - Home, _Sidebar, System-0-1-2: nav links + backlinks to the new reviews. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 16, 2026
  • Add in-depth review: Demystifying Action Space (EEF vs Joint, absolute vs delta) Long-form review of arXiv 2602.23408 (ICML 2026): the action-abstraction taxonomy (joint vs task/EEF space; absolute vs delta; chunk-wise vs step-wise), the EEF-vs-joint verdict (joint=stability & scales; EEF=generalization/transfer), the O(k)-vs-O(1) noise-propagation result behind chunk-wise delta, and the paper's practical guidelines. Embeds paper Figures 1/2/3/5 (attributed). Linked from the one-pager, sidebar, Home deep-dives, and Changelog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 11, 2026
  • Home §4: make the Survey/Index column consistent The column mixed three value types (Survey / bare years for CVPR / "Index (99 manip)" for ICML). Normalize to one year per row and a single uniform link per cell: "Survey" for survey venues, "Index" for ICML. CVPR 2025 link and ICML's 99-paper count moved into the Focus column so nothing is lost. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 11, 2026
  • Restructure Home around a Research Decision Map; split out Changelog New front-page structure: scale/freshness header + start-here, (1) intent-based Research Decision Map, (2) Latest Updates (moved up, dated, links to Changelog), (3) Core Topic Reviews (cross-paper only, one-liners restored), (4) merged venue survey table, (5) per-paper/series deep-dives, (6) foundational refs, (7) page-format/maintenance rule. Adds the new VLA Training Frameworks and RoboMME pages; de-dups OpenVLA; splits a dated Changelog.md (sidebar-linked). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 11, 2026