-
Notifications
You must be signed in to change notification settings - Fork 0
Review ICRT
Paper: "In-Context Imitation Learning via Next-Token Prediction" β arXiv 2408.15980 Β· ICRA 2025 Β· UC Berkeley (Letian Fu et al.) Β· code. The foundational token-sequence datapoint of in-context imitation β a causal transformer over sensorimotor trajectories that learns a new task from a prompt of demonstrations at test time, no fine-tuning, no language, no reward. Companions: In-Context Imitation (this is its cluster-E anchor) Β· RoboSSM Β· Behavior Prompting Β· MimicDroid.

Robots should learn a new task by being shown a few demonstrations at test time β the way LLMs do in-context learning β without updating policy parameters and without language or reward supervision. The question ICRT answers: can next-token prediction over raw sensorimotor trajectories give a real robot this in-context ability?
ICRT is a causal transformer trained by autoregressive prediction on sensorimotor trajectories (image observations, proprioceptive states, actions) β no linguistic data, no reward.
-
Tokenization: multi-view images (third-person + wrist) β ViT β attention pooling β a state token
f_s; proprioception β MLP; each action β MLP βf_a. A trajectory becomes an interleaved[f_sΒΉ, f_aΒΉ, f_sΒ², f_aΒ², β¦]sequence. -
Prompt-then-rollout as one sequence: at test time the model is prompted with the new task's trajectory tokens (
π―_prompts), then continues the sequence β autoregressively emitting actions for the current rollout (π―β, π―β, β¦). This is training-free task specification β the task is the prompt, not a gradient step. - Training data: teleoperated multi-task trajectories; the model learns how to imitate from context, not any single task.
- On a Franka Emika robot, ICRT adapts to new tasks specified purely by the prompt, even in environment configurations different from both the prompt and the training data.
- In a multitask setup, ICRT significantly outperforms prior next-token-prediction robot models on generalizing to unseen tasks.
ICRT is the anchor of cluster E (token-sequence ICL) in In-Context Imitation β it established that LLM-style next-token prediction over [demo; rollout] tokens gives a real robot genuine in-context imitation, with no language, no reward, no test-time training. Everything after refines this recipe:
- RoboSSM swaps ICRT's Transformer for an SSM to fix its long-prompt collapse (in RoboSSM's own experiments ICRT degrades to ~0 as demos grow past training length β the direct motivation).
- ICLR: Visual Reasoning adds image-space intent traces to ICRT-style prompts (action-only β action+intent).
- Behavior Prompting and MimicDroid attack ICRT's data bottleneck (teleop β handheld interface / human play video).
So ICRT is best read as the reference point the newer in-context papers improve on β architecture (SSM), prompt content (visual reasoning), and data source (video/handheld).
Limitations. Transformer prompt length caps how many/long demos help β degrades past the training length (RoboSSM's finding); needs teleoperated training trajectories (the scalability wall Behavior Prompting/MimicDroid target); no language conditioning, so tasks must be shown, not told.
- Paper: arXiv 2408.15980 Β· HF Β· code Β· IEEE Xplore
- Family: In-Context Imitation (cluster E) Β· RoboSSM Β· Behavior Prompting Β· MimicDroid Β· ICLR: Visual Reasoning
β Back to In-Context Imitation Β· Reviews Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)