Skip to content

History / Review In Context Imitation

Revisions

  • In-context imitation: add in-depth pages for ICRT, Behavior Prompting, MimicDroid (w/ paper figures + arXiv) - Review-ICRT (arXiv 2408.15980, ICRA'25, Berkeley): causal-transformer next- token ICL over sensorimotor tokens — the cluster-E anchor; method figure. Framed as the reference the newer papers improve on (SSM/visual-reasoning/data). - Review-Behavior-Prompting (arXiv 2606.30457, Stanford/Shuran Song): single-demo 'behavior prompt' via cross-attn prompt encoder + diffusion decoder; key finding = task diversity drives prompting; iPhUMI handheld data + DrawAnything/LIBERO-Gen. Architecture figure. - Review-MimicDroid (arXiv 2509.09769, ICRA'26, UT Austin RPL): learns ICL from unlabeled human play video (no teleop) via similar-behavior pairs + wrist-pose retargeting + patch masking; ~2x real success. Overview figure. - Linked all three from Review-In-Context-Imitation (clusters E/F + comparison table). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 7, 2026
  • IROS 2026: add full-paper pages for RoboSSM and TempoFit; link from survey + reviews - IROS-2026-RoboSSM (KAIST + UT Austin/Peter Stone): scalable in-context imitation via state-space models (Longhorn SSM, linear-time, length extrapolation, LIBERO, open-source). The in-context/memory efficiency answer. - IROS-2026-TempoFit (XJTLU): training-free temporal retrofit that reuses a frozen VLA's prefix-attention K/V as content-addressable memory (layer-wise FIFO K/V + Frame-Gap Temporal Bias); LIBERO-Long +4.0%, near-real-time. - Linked both from the IROS survey (§3.3, §5.2), Review-In-Context-Imitation, and Review-VLA-Memory (structured-vs-parametric row). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 5, 2026
  • IROS 2026: add 5 full-paper analyses + reflect into existing in-depth reviews Full-paper pages (abstract-verified from official program): - IROS-2026-AtomVLA: subtask-aware VLA + latent-WM scoring for offline GRPO (WAM lens) - IROS-2026-3D-FlowMatch-Actor: CMU/NVIDIA unified single/dual-arm 3D policy, +41.4% PerAct2, ~30x faster (bimanual SOTA) - IROS-2026-EquiBim: symmetry-equivariant bimanual policy - IROS-2026-IMLE-VLA: single-step cIMLE action head, 55Hz, LIBERO 98.0% (efficiency) - IROS-2026-ICLR-Visual-Reasoning: in-context imitation with image-space reasoning traces Reflected IROS 2026 into existing reviews: - Review-World-Models: WAM-as-critic row (AtomVLA offline GRPO) - Review-In-Context-Imitation: ICLR-visual-reasoning + RoboSSM - Review-VLA-Memory: structured-vs-parametric memory row (GaussMemory/PROMPT vs TempoFit/RoboSSM) - Review-Multitask-VLA: VLA-RL / LAR-MoE / AtomVLA / MoE-humanoid - Review-Realtime-Execution: single-step head row (IMLE-VLA) - Review-Humanoid-VLA: IROS bimanual/whole-body trend (3DFA/EquiBim/ULTRA/CEER/MoE-VLA) Linked all from the IROS 2026 survey. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Sep 5, 2026
  • Add in-depth survey: In-Context Imitation & Demo-Following Review-In-Context-Imitation: watch-a-demo-and-reproduce-it (no per-task FT) as a memory-conditioning problem. Taxonomy by how the demo is conditioned — (A) cross-attention/video-conditioned (Vid2Robot, VLBiMan, See-Once-Then-Act), (B) recurrent/query memory (HAMLET, MemoryVLA, RememVLA, ContextVLA), (C) fast-weight/TTT (RoboTTT), (D) retrieval (MemER, MAP-VLA, Memory-Retrieval, KEMO, Long-Context-IL), (E) token-sequence ICL (ICRT, Behavior Prompting), (F) play-video ICL (MimicDroid). Mapped to RoboMME's Imitation (procedural memory) suite; comparison table + design axes + open challenges. Cross-linked from Reviews and Home. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 31, 2026