-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 ICLR Visual Reasoning
Venue: IROS 2026 (Pittsburgh) Β· paper #268 Β· University of Southern California Β· Autodesk Research (Nguyen, Yuan, Wei, Li, Seita, Wang). The in-context-imitation datapoint of IROS 2026 β augment demonstration prompts with visual reasoning traces (anticipated future trajectories in image space) so the policy mimics intent, not just actions. (Note: "ICLR" here is the paper's name, not the conference.) Companions: In-Context Imitation Β· RoboMME Β· IROS 2026 survey.
In-context imitation learning lets robots adapt to new tasks from a few demos without additional training β but existing approaches condition only on stateβaction trajectories and lack an explicit representation of task intent. In ambiguous settings the same actions can serve different objectives, so action-only prompts underperform.
ICLR augments the demonstration prompt with structured visual reasoning traces β anticipated future robot trajectories rendered in image space β and jointly learns to generate the reasoning traces and the low-level actions in one autoregressive transformer. The policy thus mimics not only action prediction but the reasoning process that leads to those actions β embodied visual chain-of-thought for in-context demo-following.
- Sim + real-world manipulation: consistent improvements in success rate and generalization to unseen tasks and novel object configurations vs other in-context imitation methods.
- Take-away: embodied visual reasoning is a promising axis for robust robotic in-context learning.
ICLR maps directly onto the In-Context Imitation taxonomy as an token-sequence ICL (E) Γ visual-reasoning hybrid: the demo is context, and the added image-space reasoning trace is the "intent" channel that pure state-action prompts lack β the same gap RoboMME's Imitation (procedural-memory) suite probes. It pairs with IROS 2026's other in-context entries β RoboSSM (state-space long-context ICL) and RoboTTT-style fast-weight memory β to show the survey Β§5.2 point: in-context imitation is maturing from "condition on trajectories" to "condition on reasoned intent." The visual-reasoning-trace idea also echoes the subgoal/keyframe-imagination line (Ο0.7, HALO's EM-CoT).
Limitations (reviewer): generating image-space traces adds an inference step (latency vs pure action ICL); trace quality bounds action quality; evaluated on manipulation tasks, not long-horizon compositional suites.
- Official program: IROS 2026 (paper #268) Β· survey: IROS 2026
- Related: In-Context Imitation Β· RoboMME Β· RoboTTT Β· VLA Memory
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)