-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 RoboSSM
Venue: IROS 2026 (Pittsburgh) Β· paper #3133 Β· KAIST Β· UT Austin (Yoo, Hu, Zhu, Liu, Liu, MartΓn-MartΓn, Peter Stone). Paper: arXiv 2509.19658 Β· OpenReview Β· code. The state-space datapoint for in-context imitation β replace the Transformer backbone of in-context imitation learning with an SSM (Longhorn) for linear-time inference and strong long-prompt extrapolation. Companions: In-Context Imitation Β· RoboTTT Β· VLA Memory Β· IROS 2026 survey.

In-context imitation learning (ICIL) lets a robot learn a task from a prompt of a few demonstrations β no deployment-time parameter updates, so it supports few-shot adaptation to novel tasks. But recent ICIL methods are Transformer-based, which has quadratic computational limits and underperforms when the prompt is longer than those seen at training (poor length extrapolation).
RoboSSM is a scalable ICIL recipe built on state-space models:
- Replaces the Transformer with Longhorn β a state-of-the-art SSM giving linear-time inference and strong extrapolation β making it well-suited to long-context prompts (many/long demos).
- The recurrent SSM state is the memory: context is compressed into a fixed-size state rather than an ever-growing attention cache β the same "memory as parametric state, not cache" idea as RoboTTT's fast weights, but via SSM recurrence.
- On LIBERO, RoboSSM processes prompts up to 16Γ longer than seen in training and outperforms the Transformer-based ICIL baseline (ICRT) on unseen and long-horizon tasks β the gap widens exactly as the demonstration count grows (see figure: ICRT collapses past its training length while RoboSSM holds).
- "For the first time, SSMs are an efficient and scalable backbone for ICIL." Code released.
RoboSSM is IROS 2026's clearest evidence that the long-context latency wall of in-context imitation (In-Context Imitation Β§6) has an architectural escape: SSMs give linear cost + length extrapolation, so more/longer demos help instead of blowing up compute. It sits in the In-Context Imitation Β§2 taxonomy as a token-sequence-ICL (E) method with an SSM (not Transformer) backbone, and pairs with the other two IROS "efficient long-context memory" answers β TempoFit (KV-cache memory) and RoboTTT (fast-weight TTT). Together they show the field converging on parametric/recurrent context over ever-larger attention windows.
Limitations (reviewer): LIBERO-only evaluation; SSM extrapolation on real multi-hour prompts and contact-rich tasks untested; Longhorn's fixed-size state may bottleneck very information-dense demos.
- Official program: IROS 2026 (paper #3133) Β· code: GitHub Β· survey: IROS 2026
- Related: In-Context Imitation Β· TempoFit Β· RoboTTT Β· ICLR: Visual Reasoning Β· VLA Memory
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)