Add CoRL 2026 survey (preliminary, community-sourced; official list pending)
CoRL 2026 (Austin, Nov 9-12; 687 accepted / 32.8% / 2,094 submitted; official
per-paper program not yet public). Preliminary manipulation-centric survey with
explicit acceptance-confidence marking:
- Confirmed CoRL 2026: SG-WAM (geometry-aware WAM), FiberTune (robustness-
preserving VLA fine-tune), Dex-X (visual-tactile from human video via sim),
Touch2Trace (tactile IL, cable tracing).
- Reported/unverified: StellaVLA (in-context VLA), Choice Policies (Berkeley/
Malik, whole-body humanoid), Weave (whole-body dexterous loco-manip).
- Excluded ManiFlow (it's CoRL 2025, not 2026). Prominent caveat + confidence
legend; to be promoted to a full session-taxonomy survey when the program
publishes. Linked from CoRL hub, Home venue table, sidebar.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Split Reviews into topic-reviews (Reviews) + per-paper long-forms (Reviews-Per-Paper)
Durable fix for the recurring 'too long to render' on the growing catalog:
- Reviews.md keeps cross-paper topic reviews, lab/series programs, latest-paper
reviews (78 links).
- New Reviews-Per-Paper.md holds all single-paper long-forms — architecture/
runtime, data/training, world-models/tactile, hybrid-MoT, dexterous-hand data,
multi-task/in-context, IROS 2026 full-paper analyses, RSS pointer (56 links).
- Both now safely under GitHub-wiki's ~100-link render limit.
- Linked Per-Paper from Home and the sidebar; added the IROS 2026 + hybrid +
dexterous-data long-forms that weren't catalogued before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add IROS 2026 VLA & Manipulation survey (official program, pre-conference)
IROS-2026-VLA-Manipulation-Survey: built from the public official program
(2026.ieee-iros.org, 1,900+ papers, Pittsburgh Sep 27-Oct 1). Session-level
taxonomy of ~40 in-scope manipulation/VLA/dexterous/imitation sessions +
sampled papers extracted per theme (VLA: AnyCamVLA/LangGap/OG-VLA/BFA++/
Safe-Night-VLA; scaling/RL: VLA-RL/LAR-MoE/flow-matching; in-context imitation:
RoboSSM/ICLR-visual-reasoning/IMLE-VLA/ESPADA; policy: MaskVLA/SynthLA;
language-in-loop: NL2SpaTiaL/AURORA). Clearly flagged title-level (abstracts
pending Xplore post-conference), not abstract-verified. Added to Home venue
table + top pointer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth survey: In-Context Imitation & Demo-Following
Review-In-Context-Imitation: watch-a-demo-and-reproduce-it (no per-task FT) as
a memory-conditioning problem. Taxonomy by how the demo is conditioned —
(A) cross-attention/video-conditioned (Vid2Robot, VLBiMan, See-Once-Then-Act),
(B) recurrent/query memory (HAMLET, MemoryVLA, RememVLA, ContextVLA),
(C) fast-weight/TTT (RoboTTT), (D) retrieval (MemER, MAP-VLA, Memory-Retrieval,
KEMO, Long-Context-IL), (E) token-sequence ICL (ICRT, Behavior Prompting),
(F) play-video ICL (MimicDroid). Mapped to RoboMME's Imitation (procedural
memory) suite; comparison table + design axes + open challenges. Cross-linked
from Reviews and Home.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth survey: Egocentric Video for VLA Pre-Training
Review-Egocentric-Video-Pretraining: how label-free first-person human video
becomes a pretraining signal. Taxonomy of methods (A pseudo-action extraction:
hand-pose/keypoints/optical-flow/part-motion; B latent-action models; C
world-model/video-prediction; D reconstruct-then-retarget; E auxiliary-modality
recovery), the dataset landscape (EgoDex, EgoVerse, EgoScale, Being-H0, DYNA-2,
DreamDojo, UniDex, EgoVLA), scaling-law evidence (EgoScale R2=0.9983 +54%,
DYNA-2 transfer law), the transfer question (emergence/decoupling/synthesis +
two-stage recipe), and open challenges. Cross-linked from Reviews and Home.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth review of NVIDIA RoboTTT (context scaling via TTT in GR00T N1.7)
Review-RoboTTT (arXiv 2607.15275, NVIDIA GEAR + Stanford + UT Austin):
8K-timestep visuomotor context at constant latency by adding Test-Time-
Training fast-weight layers to GR00T N1.7's DiT action head. Detailed GR00T
implementation section: TTT layer after self/cross-attn in each of 16 DiT
layers (~10M each -> ~690M), tanh-gated to preserve pretrained skills;
register tokens (N=16) carry compressed VL history through TTT while VL
tokens bypass; fast weights = 2-layer GeLU MLP updated per step (W_t <-
W_{t-1} - eta*grad MSE(f(K),V)), read via Q; training recipe = flow matching
+ sequence action forcing (per-step tau) + TBPTT (fast weights carried,
gradients detached at segment boundaries); 30 Hz on RTX 5090 (YAM bimanual).
Results, new capabilities (one-shot in-context video imitation, DAgger-
distillation self-improvement, perturbation robustness), limitations.
Cross-linked from Reviews, Home lab programs, and Review-GR00T-Series.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Multi-Task VLA review + MergeVLA page
- Review-Multitask-VLA: why one VLA fails across many tasks (6 failure modes:
negative transfer/gradient conflict, non-mergeability, multi-task conflict,
catastrophic forgetting, routing confusion, instruction collapse) + a
7-cluster solution landscape (merging, MoE, gradient control, skill
decomposition, instruction grounding, continual, adapters) with a decision
guide and open questions. Anchored on MergeVLA (CVPR 2026).
- CVPR-2026-MergeVLA: per-paper page (non-mergeability diagnosis + task-masked
LoRA / cross-attention-only action expert / test-time task router).
- Cross-linked from Reviews catalog and Home topic reviews.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Review-Dexterous-Hand-Data-Pyramid: data types + approach×data matrix + current/future insight
New page classifying data-acquisition/generation methods for humanoid
5-finger dexterous hands as a 6-tier data pyramid (web-video → egocentric
→ wearable glove/exo → retargeting → sim/synthetic → target-hand teleop),
with tactile/force as a cross-cutting axis. Includes:
- per-data-type description + pros/cons table
- ONE summary matrix: approach (paper/company) × data-type used
(EgoScale, DYNA-2, UniDex, Being-H0, DO-AS-I-DO, DexUMI, YUBI/DexEXO,
AnyDexRT, Dex1B, DexGrasp-Zero, DexNDM, shared-autonomy, MANUS teleop,
RLDX-1, pi0.5/0.7, T-Rex, One-Hand)
- current most-common recipe (teleop+retarget+sim) and future directions
(human-video scaling laws, device-free RGB, calib-free retargeting,
tactile-first, cross-hand co-design, unified WAM) w/ latest 2026 papers.
Cross-linked from Review-Dexterous-Manipulation, Reviews, Home.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Review-Single-Checkpoint-Multi-Robot: deploy-side cross-embodiment (one frozen checkpoint, many robots)
New page separating the DEPLOYMENT question (one unchanged checkpoint
controls multiple physical robots at inference) from the training-data cut
in Review-Cross-Embodiment. Organized by three tiers:
- Tier 2 (zero-shot to unseen robot): LAP-3B (actions-as-language, first
substantial zero-shot to unseen), Green-VLA, Gemini Robotics 1.5,
DreamZero/DYNA-2 (WAM), Contact-Anchored Policies, One-Hand.
- Tier 1 (routed seen-robot generalist): RT-X, CrossFormer, RDT-1B,
UniAct, GR00T N1, pi0.5->pi0.7, Motus.
- Tier 3 (contrast, needs per-robot fit): Octo, HPT, X-VLA; plus MergeVLA
(merge specialists into one checkpoint).
Includes a routing-mechanism taxonomy, honest limits (AnyBody, unreplicated
2026 zero-shot claims, 'single checkpoint != nothing per robot'), and a
design guide. Cross-linked from Review-Cross-Embodiment, Reviews, Home.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Embed paper architecture figures in Motus/HALO/BAGEL; add hybrid link to Home topic reviews
- Motus: embed the paper's tri-expert architecture figure (Video Gen /
Action / Understanding + Tri-modal Joint Attention), arXiv 2512.13030.
- BAGEL: embed the MoT figure (Und/Gen experts + Multi-modal Self-Attention,
dual encoders), arXiv 2505.14683.
- HALO: embed Fig.1 (three-expert MoT) and Fig.2 (EM-CoT data pipeline).
All figures credited to the authors alongside the schematic mermaids.
- Home: add [[VLA Hybrid Architectures]] to the Topic reviews row.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Escape pipe in all in-table wikilinks wiki-wide (202 links, 17 files)
GitHub-wiki table cells read a wikilink's separator | as a column
delimiter, splitting the cell and breaking the link. Escape to \| in
every table-row wikilink (Home nav, RSS-2026-Papers, topic surveys).
Prose wikilinks left as plain | (render correctly outside tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Latest Papers tracker + omega-0 and Stellar VLA in-depth reviews
New Latest-Papers.md preprint tracker (pre-publication reviews) and two
figure-illustrated in-depth reviews: omega-0 (arXiv 2608.06375, whole-
body humanoid latent-predictive World Action Model; 81.8% on 11
household tasks vs 44.5% psi-0; ships 40h omega-HOME dataset) and
Stellar VLA (arXiv 2511.18085, continual imitation learning with a
Dirichlet-Process knowledge space + knowledge-routed MoE, 1% replay).
Cross-linked from Home, Reviews, sidebar, Humanoid-VLA, World-Models.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add DreamZero in-depth review (World Action Models are Zero-shot Policies)
NVIDIA's 14B video-diffusion World Action Model (arXiv 2602.15922):
jointly predicts video+action, >2x over SOTA VLAs on unseen-env/
unseen-object real-robot evals, 38x inference stack (DreamZero-Flash)
for 7 Hz closed-loop control, video-only cross-embodiment transfer.
Fig. 4 architecture embedded. Cross-linked from World Models review,
Home lab-programs, Reviews catalog, and sidebar.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Weave ICML 2026 evidence into deep-dive surveys; revise three verdicts
14 State-of-the-Field sections gain ICML 2026 findings from the
99-paper index: recipes and latent-action supervision (VLANeXt,
From-Pixels-to-Tokens, XR-1), MoT dual-systems and shortcut counters,
the 9-paper efficiency cluster (Reflex 50Hz, GridS -76% FLOPs, XPU
profile, latent reasoning -90%), reward/critic and model-based RL
(VLAC, VLAW +39.2%), memory (HiMe/SOMA/CAPS), world models (DreamDojo
44kh, LAC-WM, dWorldEval), dexterous (DexMachina/DECO/Tabero/CTSRL),
cross-embodiment (OXE-AugE, latent motion codes), evaluation
(LIBERO-Gen, VLA-Arena, FixBench, TRAP). Verdicts revised: forgetting
milder than assumed; discrete-token verdict scoped to robot-action
auxiliaries; WM-evaluator action gap first crack. Home synced.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Decision map: prominent deep-dive links + uniform detail-page template
Home fold-outs now lead with a heading-level "Deep dive ->" link and
compress trend/approaches/limitations into a labeled 3-row table.
The 12 detail-page State-of-the-Field sections are rewritten to one
template (Verdict quote + Trend + Approaches-and-trade-offs +
optional Established-findings + Limitations, dated Aug 2026); the
three standalone surveys get matching headers with structure legends.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Survey-depth topic pages: 12 State-of-the-Field updates + 3 new surveys
Each decision-map topic's detail page now carries a dated July-2026
survey section: trend arc through the latest venues, approach
taxonomy with definitions and trade-offs, and current limitations.
Three previously page-less topics get dedicated surveys:
Human-Video Transfer (emergence/decoupling/synthesis fork + decision
guide), VLA Evaluation (indictment + 2026 toolkit + emerging norms),
Real-Time Execution (RTC->Legato arc + approach comparison). Home
fold-outs link the full surveys; Reviews catalog updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rewrite Home decision map as an insight map
Each of the 15 topics is now a collapsible entry carrying the trend
arc through the latest venues, competing approaches with definitions
and trade-offs, and current limitations - replacing the plain link
table. All claims use wiki-verified numbers (OAT/Legato, RECAP,
LDA-1B/mimic-video, VLM4VLA +18.1, DexGrasp-Zero/OHRA, Psi-0,
LIBERO-X pyramid, LBM verdicts, camera-frame EEF scaling law).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Redesign Home: themed decision map, stat strip; move rules to Maintenance
Decision map split into four themed tables (Building / Running &
improving / Data & evaluation / Embodiment) with a current-answer
column; stat strip and emoji headers added; reviews section as a
compact table; Page Format / Maintenance Rule section moved to the
new Maintenance page, linked from the footer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Restructure navigation: top-level sidebar + Reviews catalog page
New Reviews.md holds the complete in-depth-review catalog (topic
reviews, lab programs, per-paper long-forms, RSS 2026 figure pages).
Sidebar slimmed to top-level only: reviews hub + six star topics,
model lineages, ML hub, one link per venue year (venue pages already
index their papers), foundational refs. Home merges its two review
sections into one compact section pointing at the catalog.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Refresh Home research decision map with RSS 2026 findings
Four new question rows (improvement-from-experience, human-video-vs-
robot-data fork, co-training data selection, credible evaluation) and
five rows updated with RSS evidence (Legato/OAT, LDA-1B/mimic-video,
ViTacFormer/CGP, cross-hand transfer, Psi-0). Header stats and
start-here pointer updated to the RSS 2026 survey.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add RSS 2026 conference survey, Psi-0 in-depth, and 14 per-paper pages
RSS 2026 (Sydney, Jul 13-17): 210 accepted papers parsed from the official
program, ~116 manipulation/hand/humanoid papers in scope, all abstracts
verified. Survey covers six threads: RL-from-experience (pi*0.6/RECAP
flagship), human-video transfer (emergence vs decoupling), video/world
models vs VLA backbones, contact-as-representation, cross-embodiment
dexterous hands, and evaluation infrastructure. New: RSS venue hub,
Review-Psi0 (full-paper in-depth: 800h human video + 30h robot data
beats 10x corpora by >40pp on Unitree G1), per-paper pages for LDA-1B,
H2R-Emergence, mimic-video, LBM co-training study, ViTacFormer,
DexGrasp-Zero, One-Hand, Contact-Grounded Policy, PolaRiS, LIBERO-X,
OAT, Legato, HoMMI. PI-RECAP updated with RSS camera-ready results.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add in-depth Qwen-RobotNav and Qwen-RobotWorld reviews; update program page
RobotNav (2606.18112): parameterized observation interface, 15.6M corpus,
agentic EQA SOTA, supersedes Qwen-VLA nav by ~15pp. RobotWorld (2606.17030):
20B double-stream MMDiT, frozen Qwen2.5-VL action encoder, EWK 8.6M corpus,
Scene2Robot; 1st on EWMBench/DreamGen. Program page: suite rows upgraded to
deep-read, backbone/lambda doctrine qualified, System-2 slot marked
demonstrated, watch-list items 4-5 updated; sibling cross-links added.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add cross-paper review: Qwen Team's VLA Program
VLM4VLA -> Qwen-VLA -> Qwen-Robot Suite (RobotManip/RobotNav/RobotWorld):
shared doctrine (Qwen3.5-4B, flow matching, lambda=0.1 VL co-training,
synthetic-data scaling, language-as-interface), diagnostic-to-flagship
trace, internal contradictions between the two flagship VLAs, and
lab-program positioning vs PI/GR00T/TRI/Gemini. Cross-links from the
three constituent reviews; indexed in Home/Changelog.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add in-depth Qwen-RobotManip review (arXiv 2606.17846)
Alignment-first scaling thesis, camera-frame delta EEF + CaPE,
38,100h open-data corpus with 24,808h human-to-robot synthesis,
RoboTwin-IF/XE benchmarks, RoboChallenge Table30-v1 generalist #1.
Index in Home/Changelog; cross-link from Qwen-VLA review.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add dedicated NVIDIA WAM + Cosmos 3 review; slim §10b to a pointer
- New Review-NVIDIA-WAM-Cosmos3: self-contained bundle of the "imagine->act"
blog + a detailed Cosmos 3 analysis, with original recreated schematics
(mermaid) and data tables (RoboArena, paradigms, tiers, roles, benchmarks).
NVIDIA copyrighted figures are linked, not embedded.
- WAM review §10b trimmed to a brief summary + pointer to the new page
(keeps the review-specific verdict section C).
- Wired into Home (Evaluation & Robustness) and the sidebar Per-paper list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add cross-topic review: Independent Visual Representation in VLAs
New Review-Independent-Visual-Representation: collects research that builds
the visual representation OUTSIDE the VLM, motivated by the insufficiency of
VLM (language-aligned) vision for action. Taxonomy: G (spatial/geometric),
P (predictive/world-model latents), M (temporal memory), D (dense/SSL
encoders) x a 5-rung injection ladder (alignment loss -> fused tokens ->
AdaLN modulation -> latent interface -> replace-VLM/geometry-first).
Wired into Home (reviews list + nav table) and the sidebar.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Humanoid VLA + FTP-1 reviews; update Tactile & Dexterous cross-reviews
- New: Review-Humanoid-VLA (whole-body/bipedal loco-manip: balance, high-DoF
action space, data collection) and Review-FTP-1 (cross-sensor tactile
foundation policy).
- Tactile review: integrate FTP-1 as new taxonomy axis G (cross-sensor),
third fault line, comparison table, limitations.
- Dexterous review: add ICML 2026 cluster (§4c) + ICLR 2026 completeness
pass (DemoGrasp, DexMove, D-REX, OmniReset, House of Dextra).
- Home, _Sidebar, System-0-1-2: nav links + backlinks to the new reviews.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth review: Demystifying Action Space (EEF vs Joint, absolute vs delta)
Long-form review of arXiv 2602.23408 (ICML 2026): the action-abstraction
taxonomy (joint vs task/EEF space; absolute vs delta; chunk-wise vs step-wise),
the EEF-vs-joint verdict (joint=stability & scales; EEF=generalization/transfer),
the O(k)-vs-O(1) noise-propagation result behind chunk-wise delta, and the
paper's practical guidelines. Embeds paper Figures 1/2/3/5 (attributed).
Linked from the one-pager, sidebar, Home deep-dives, and Changelog.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Home §4: make the Survey/Index column consistent
The column mixed three value types (Survey / bare years for CVPR / "Index
(99 manip)" for ICML). Normalize to one year per row and a single uniform link
per cell: "Survey" for survey venues, "Index" for ICML. CVPR 2025 link and
ICML's 99-paper count moved into the Focus column so nothing is lost.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Restructure Home around a Research Decision Map; split out Changelog
New front-page structure: scale/freshness header + start-here, (1) intent-based
Research Decision Map, (2) Latest Updates (moved up, dated, links to Changelog),
(3) Core Topic Reviews (cross-paper only, one-liners restored), (4) merged
venue survey table, (5) per-paper/series deep-dives, (6) foundational refs,
(7) page-format/maintenance rule. Adds the new VLA Training Frameworks and
RoboMME pages; de-dups OpenVLA; splits a dated Changelog.md (sidebar-linked).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>