-
Notifications
You must be signed in to change notification settings - Fork 0
Review Behavior Prompting
Paper: "Behavior Prompting Policy: Demonstrations as Prompts for Manipulation" β arXiv 2606.30457 (Jun 2026) Β· Stanford Β· UC Berkeley (Patel, Pekarek, Castro Hernandez, Shuran Song) Β· project Β· code. The "demo-as-prompt" datapoint β perform a new task from a single human demonstration ("behavior prompt") at inference, and it identifies task diversity as the primary driver of the prompting ability. Companions: In-Context Imitation Β· ICRT Β· MimicDroid.

In-context imitation should let a robot do a new task from a single demonstration at test time β but two things block it: (1) the architecture must translate an arbitrary demo + the current observation into actions, and (2) what training data actually confers the prompting ability was unclear.
(Algorithm) Behavior Prompting Policy (BPP) β an in-context visuomotor architecture:
- The behavior prompt (one demo) is chunked into
(observation, action-segment, proprioception)chunks; each chunk is attention-pooled into a prompt token. - A prompt encoder (transformer) has the current observation self-attend, then cross-attend to the prompt chunks, yielding a prompt embedding.
- An action decoder (diffusion) denoises the current action chunk conditioned on
current obs + prompt embedding.
(Data) Task diversity is the driver. The paper's key empirical finding: task diversity β not sheer quantity β is the primary driver of prompting capability. To collect diverse data cheaply, they introduce iPhUMI, a handheld manipulation interface (a UMI-style device β cf. DexUMI/YUBI).
(Evaluation) New test-time-adaptation benchmarks: DrawAnything (unseen drawing tasks) and LIBERO-Gen (unseen tabletop manipulation) β probes for genuine test-time adaptation to unseen tasks.
BPP advances the In-Context Imitation family on the axis ICRT left open β where the prompting ability comes from. Its answer, task diversity > quantity, is a training-data recipe that reframes the whole family: to get in-context generalization you need many different tasks, and a cheap diverse-data interface (iPhUMI) is the enabler β connecting in-context imitation to the Dexterous-Hand Data Pyramid's L3 capture-interface thread. Architecturally it's a cross-attention (prompt-encoder) + diffusion (decoder) hybrid β sitting between cluster A (cross-attention on the demo) and the diffusion-head mainstream, distinct from ICRT's pure next-token sequence. The DrawAnything / LIBERO-Gen benchmarks give the field a cleaner test of unseen-task adaptation.
Limitations (reviewer). Single-demo prompting is powerful but sensitive to prompt-demo quality/coverage; the diversity finding is shown on their tasks/benchmarks; handheld-interface data still needs collection (cheaper than teleop, not free); diffusion decoder adds inference cost vs a single-step head.
- Paper: arXiv 2606.30457 Β· project Β· code Β· HF
- Family: In-Context Imitation Β· ICRT Β· MimicDroid Β· RoboSSM
- Data interface kin: DexUMI Β· YUBI Β· Dexterous-Hand Data Pyramid
β Back to In-Context Imitation Β· Reviews Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)