In-context imitation: add in-depth pages for ICRT, Behavior Prompting, MimicDroid (w/ paper figures + arXiv)
- Review-ICRT (arXiv 2408.15980, ICRA'25, Berkeley): causal-transformer next-
token ICL over sensorimotor tokens — the cluster-E anchor; method figure.
Framed as the reference the newer papers improve on (SSM/visual-reasoning/data).
- Review-Behavior-Prompting (arXiv 2606.30457, Stanford/Shuran Song): single-demo
'behavior prompt' via cross-attn prompt encoder + diffusion decoder; key finding
= task diversity drives prompting; iPhUMI handheld data + DrawAnything/LIBERO-Gen.
Architecture figure.
- Review-MimicDroid (arXiv 2509.09769, ICRA'26, UT Austin RPL): learns ICL from
unlabeled human play video (no teleop) via similar-behavior pairs + wrist-pose
retargeting + patch masking; ~2x real success. Overview figure.
- Linked all three from Review-In-Context-Imitation (clusters E/F + comparison table).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
IROS 2026: add full-paper pages for RoboSSM and TempoFit; link from survey + reviews
- IROS-2026-RoboSSM (KAIST + UT Austin/Peter Stone): scalable in-context
imitation via state-space models (Longhorn SSM, linear-time, length
extrapolation, LIBERO, open-source). The in-context/memory efficiency answer.
- IROS-2026-TempoFit (XJTLU): training-free temporal retrofit that reuses a
frozen VLA's prefix-attention K/V as content-addressable memory (layer-wise
FIFO K/V + Frame-Gap Temporal Bias); LIBERO-Long +4.0%, near-real-time.
- Linked both from the IROS survey (§3.3, §5.2), Review-In-Context-Imitation,
and Review-VLA-Memory (structured-vs-parametric row).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
IROS 2026: add 5 full-paper analyses + reflect into existing in-depth reviews
Full-paper pages (abstract-verified from official program):
- IROS-2026-AtomVLA: subtask-aware VLA + latent-WM scoring for offline GRPO (WAM lens)
- IROS-2026-3D-FlowMatch-Actor: CMU/NVIDIA unified single/dual-arm 3D policy,
+41.4% PerAct2, ~30x faster (bimanual SOTA)
- IROS-2026-EquiBim: symmetry-equivariant bimanual policy
- IROS-2026-IMLE-VLA: single-step cIMLE action head, 55Hz, LIBERO 98.0% (efficiency)
- IROS-2026-ICLR-Visual-Reasoning: in-context imitation with image-space reasoning traces
Reflected IROS 2026 into existing reviews:
- Review-World-Models: WAM-as-critic row (AtomVLA offline GRPO)
- Review-In-Context-Imitation: ICLR-visual-reasoning + RoboSSM
- Review-VLA-Memory: structured-vs-parametric memory row (GaussMemory/PROMPT vs TempoFit/RoboSSM)
- Review-Multitask-VLA: VLA-RL / LAR-MoE / AtomVLA / MoE-humanoid
- Review-Realtime-Execution: single-step head row (IMLE-VLA)
- Review-Humanoid-VLA: IROS bimanual/whole-body trend (3DFA/EquiBim/ULTRA/CEER/MoE-VLA)
Linked all from the IROS 2026 survey.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth survey: In-Context Imitation & Demo-Following
Review-In-Context-Imitation: watch-a-demo-and-reproduce-it (no per-task FT) as
a memory-conditioning problem. Taxonomy by how the demo is conditioned —
(A) cross-attention/video-conditioned (Vid2Robot, VLBiMan, See-Once-Then-Act),
(B) recurrent/query memory (HAMLET, MemoryVLA, RememVLA, ContextVLA),
(C) fast-weight/TTT (RoboTTT), (D) retrieval (MemER, MAP-VLA, Memory-Retrieval,
KEMO, Long-Context-IL), (E) token-sequence ICL (ICRT, Behavior Prompting),
(F) play-video ICL (MimicDroid). Mapped to RoboMME's Imitation (procedural
memory) suite; comparison table + design axes + open challenges. Cross-linked
from Reviews and Home.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>