-
Notifications
You must be signed in to change notification settings - Fork 0
Review In Context Imitation
Question: how does a policy watch a demonstration and reproduce it at test time β ideally without per-task fine-tuning β by holding that demo (and its own history) in memory/context? Framing: in-context imitation is fundamentally a memory-conditioning problem. The demo is external context the policy attends to; long-horizon history is self-context. The controlled benchmark for both is RoboMME (its Imitation suite = replicate a demonstrated motion strategy, i.e. procedural memory). Companion: VLA Memory Β· RoboTTT.
Classic imitation learning trains on demonstrations. In-context imitation instead conditions on a demo held in the model's context/memory and produces actions that copy its strategy β the way in-context learning works for LLMs. Two properties define the family:
- Test-time skill acquisition β a new task is specified by showing one (or a few) demos, no gradient step per task.
- Memory is the mechanism β the demo must be encoded and retrieved during rollout, so every method is really a choice of how to store and attend to context.
This is exactly the faculty RoboMME's Imitation suite isolates (MoveCube / InsertPeg / PatternLock / RouteStick β pick-place vs push vs hook, linear vs circular), scored as procedural memory.
Encode the demo video and cross-attend from the current robot state to it.
- Vid2Robot (2403.12943) β prompt-video encoder + robot-state encoder joined by cross-attention transformers; imitates a human video without per-task fine-tuning, +20% over BC-Z, with cross-object motion transfer. The canonical demo-conditioned policy.
- VLBiMan β vision-language-anchored one-shot demonstration β generalizable bimanual manipulation.
- See Once, Then Act (2512.07582) β VLA that learns a task from a single video demonstration at test time.
Compress rollout (and demo) history into a recurrent latent or a set of learnable queries.
- HAMLET β "switch your VLA into a history-aware policy."
- MemoryVLA β perceptual-cognitive memory in a VLA.
- RememVLA (2026) β memory via dual-level recurrent queries.
- ContextVLA (2510.04246) β amortized multi-frame context.
Turn context into a parametric recurrent state updated by gradient descent at inference.
- RoboTTT (NVIDIA GEAR) β 8K-timestep context via TTT fast weights in GR00T N1.7's DiT; enables one-shot in-context imitation from a human video (Circuit 6/10 vs a baseline's 0/10) at constant latency. The strongest recent unification of long memory + demo-following.
Store many experiences/keyframes and retrieve the ones relevant to the current step.
- Memory Retrieval in Visuomotor Policies (long-horizon control) Β· MemER (scale memory via experience retrieval) Β· MAP-VLA (memory-augmented prompting) Β· KEMO (2606.23589, event-driven keyframe memory) Β· Long-Context IL (focus on key history frames).
Treat [demo tokens; rollout tokens] as one sequence and predict the next action token β LLM-style ICL.
- In-Context Robot Transformer / ICRT (ICRA 2025) β in-context imitation via next-token prediction over sensorimotor tokens (the cluster-E anchor). Keypoint Action Tokens (KAT) and Instant Policy are related graph/token-ICL variants.
- Behavior Prompting Policy (2606.30457) β demonstrations as prompts for manipulation; finds task diversity (not quantity) drives the prompting ability.
- MimicDroid (ICRA 2026) β in-context learning for humanoid manipulation from human play videos β learns to ICL from unlabeled video, no teleop; ~2Γ real success.
RoboMME turns "which memory helps which task?" into a controlled sweep: 14 memory-augmented policies on one Ο0.5 backbone (3 memory representations Γ 3 integration mechanisms). Its Imitation (procedural) suite is the in-context-imitation testbed; its temporal/spatial/object suites test self-history memory. Headline: no single memory design wins everywhere (best non-oracle = FrameSamp+Modulator, 44.5% avg; humans 90.5%) β so the Β§2 mechanisms are complementary, not competing. Design-space detail: VLA Memory.
| Approach | Demo / context form | Conditioning mechanism | No per-task FT? | Memory representation |
|---|---|---|---|---|
| Vid2Robot | human prompt video | cross-attention (A) | β | encoded demo tokens |
| VLBiMan | one-shot demo | vision-language anchor (A) | β | anchored demo |
| RoboTTT | human video (one-shot) | fast-weight TTT (C) | β | parametric (fast weights) |
| HAMLET / MemoryVLA | own history | recurrent / cognitive memory (B) | n/a (history) | latent / token buffer |
| MemER / MAP-VLA | stored experiences | retrieval (D) | β (retrieve) | external memory store |
| ICRT / Behavior Prompting | demo token sequence | next-token / cross-attn+diffusion ICL (E) | β | in-context tokens |
| MimicDroid | human play video | in-context (F) | β | in-context |
(β = specifies a new task by showing a demo, no gradient step per task.)
- Clean expert video demo, want motion transfer? β cross-attention (Vid2Robot, VLBiMan).
- Very long demo/rollout, must stay real-time? β fast-weight/TTT (RoboTTT) β memory is parametric, latency constant.
- Many past experiences, need the relevant one? β retrieval (MemER, MAP-VLA, KEMO).
- Long-horizon own-history dependence (counting, re-finding)? β recurrent/keyframe memory (HAMLET, Long-Context-IL) β and consult RoboMME for which representation fits which memory type.
- Unstructured/noisy demos (human play)? β play-video ICL (MimicDroid).
Unifying view: in-context imitation and history-memory are the same operation β attend to non-current context β differing only in whose context (a demonstrator's vs the robot's own) and where it's stored (attention tokens Β· recurrent latent Β· fast weights Β· external store).
- One demo β robust policy β sensitivity to demo viewpoint, object identity, and speed; generalization beyond the shown instance is uneven.
- Context length vs latency β attention over long demos is quadratic; fast-weight/retrieval are the escape hatches but under-benchmarked at scale (RoboTTT Β§6).
- Which memory for which task is unsolved β RoboMME shows no universal winner; humans still lead 90.5% vs ~44%.
- Humanβrobot demo gap β video demos carry the same embodiment/retargeting gap as egocentric-video pretraining.
- Evaluation β few benchmarks isolate in-context skill acquisition from ordinary multi-task competence (RoboMME's Imitation suite is a start).
- Demo-conditioned / one-shot: RoboTTT Β· VLBiMan Β· Vid2Robot (2403.12943) Β· See-Once-Then-Act (2512.07582) Β· Behavior Prompting (2606.30457) Β· MimicDroid (2026)
- IROS 2026 π: ICLR: In-Context Imitation with Visual Reasoning (EΓvisual-reasoning β image-space intent traces, USC) Β· RoboSSM (scalable in-context imitation via state-space models β linear-cost long context) β context: IROS 2026 survey Β§5.2
- Memory-augmented VLAs: HAMLET Β· MemoryVLA Β· MemER Β· MAP-VLA Β· Memory Retrieval Β· Long-Context IL Β· KEMO (2606.23589)
- Benchmark & synthesis: RoboMME Β· VLA Memory Β· VLA Architectures
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)