Skip to content

IROS 2026 ICLR Visual Reasoning

hwoo.han edited this page Sep 5, 2026 · 1 revision

IROS 2026 β€” ICLR: In-Context Imitation Learning with Visual Reasoning

Venue: IROS 2026 (Pittsburgh) Β· paper #268 Β· University of Southern California Β· Autodesk Research (Nguyen, Yuan, Wei, Li, Seita, Wang). The in-context-imitation datapoint of IROS 2026 β€” augment demonstration prompts with visual reasoning traces (anticipated future trajectories in image space) so the policy mimics intent, not just actions. (Note: "ICLR" here is the paper's name, not the conference.) Companions: In-Context Imitation Β· RoboMME Β· IROS 2026 survey.

1. Problem

In-context imitation learning lets robots adapt to new tasks from a few demos without additional training β€” but existing approaches condition only on state–action trajectories and lack an explicit representation of task intent. In ambiguous settings the same actions can serve different objectives, so action-only prompts underperform.

2. Method

ICLR augments the demonstration prompt with structured visual reasoning traces β€” anticipated future robot trajectories rendered in image space β€” and jointly learns to generate the reasoning traces and the low-level actions in one autoregressive transformer. The policy thus mimics not only action prediction but the reasoning process that leads to those actions β€” embodied visual chain-of-thought for in-context demo-following.

3. Results

  • Sim + real-world manipulation: consistent improvements in success rate and generalization to unseen tasks and novel object configurations vs other in-context imitation methods.
  • Take-away: embodied visual reasoning is a promising axis for robust robotic in-context learning.

4. Why it matters (in-context / memory lens)

ICLR maps directly onto the In-Context Imitation taxonomy as an token-sequence ICL (E) Γ— visual-reasoning hybrid: the demo is context, and the added image-space reasoning trace is the "intent" channel that pure state-action prompts lack β€” the same gap RoboMME's Imitation (procedural-memory) suite probes. It pairs with IROS 2026's other in-context entries β€” RoboSSM (state-space long-context ICL) and RoboTTT-style fast-weight memory β€” to show the survey Β§5.2 point: in-context imitation is maturing from "condition on trajectories" to "condition on reasoned intent." The visual-reasoning-trace idea also echoes the subgoal/keyframe-imagination line (Ο€0.7, HALO's EM-CoT).

Limitations (reviewer): generating image-space traces adds an inference step (latency vs pure action ICL); trace quality bounds action quality; evaluated on manipulation tasks, not long-horizon compositional suites.

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally