Skip to content

IROS 2026 RoboSSM

hwoo.han edited this page Sep 7, 2026 · 2 revisions

IROS 2026 β€” RoboSSM: Scalable In-Context Imitation Learning via State-Space Models

Venue: IROS 2026 (Pittsburgh) Β· paper #3133 Β· KAIST Β· UT Austin (Yoo, Hu, Zhu, Liu, Liu, MartΓ­n-MartΓ­n, Peter Stone). Paper: arXiv 2509.19658 Β· OpenReview Β· code. The state-space datapoint for in-context imitation β€” replace the Transformer backbone of in-context imitation learning with an SSM (Longhorn) for linear-time inference and strong long-prompt extrapolation. Companions: In-Context Imitation Β· RoboTTT Β· VLA Memory Β· IROS 2026 survey.

RoboSSM vs the Transformer-based ICL baseline (ICRT) as the number of test-time demonstrations grows to 32: RoboSSM (orange) stays flat/rising across LIBERO-Object and LIBERO-90 scenes, while ICRT (purple) collapses toward 0 once prompts exceed the training length β€” SSMs extrapolate to long prompts where Transformers do not (results figure from Yoo et al., arXiv 2509.19658, Β© the authors)

1. Problem

In-context imitation learning (ICIL) lets a robot learn a task from a prompt of a few demonstrations β€” no deployment-time parameter updates, so it supports few-shot adaptation to novel tasks. But recent ICIL methods are Transformer-based, which has quadratic computational limits and underperforms when the prompt is longer than those seen at training (poor length extrapolation).

2. Method

RoboSSM is a scalable ICIL recipe built on state-space models:

  • Replaces the Transformer with Longhorn β€” a state-of-the-art SSM giving linear-time inference and strong extrapolation β€” making it well-suited to long-context prompts (many/long demos).
  • The recurrent SSM state is the memory: context is compressed into a fixed-size state rather than an ever-growing attention cache β€” the same "memory as parametric state, not cache" idea as RoboTTT's fast weights, but via SSM recurrence.

3. Results

  • On LIBERO, RoboSSM processes prompts up to 16Γ— longer than seen in training and outperforms the Transformer-based ICIL baseline (ICRT) on unseen and long-horizon tasks β€” the gap widens exactly as the demonstration count grows (see figure: ICRT collapses past its training length while RoboSSM holds).
  • "For the first time, SSMs are an efficient and scalable backbone for ICIL." Code released.

4. Why it matters (in-context / memory lens)

RoboSSM is IROS 2026's clearest evidence that the long-context latency wall of in-context imitation (In-Context Imitation Β§6) has an architectural escape: SSMs give linear cost + length extrapolation, so more/longer demos help instead of blowing up compute. It sits in the In-Context Imitation Β§2 taxonomy as a token-sequence-ICL (E) method with an SSM (not Transformer) backbone, and pairs with the other two IROS "efficient long-context memory" answers β€” TempoFit (KV-cache memory) and RoboTTT (fast-weight TTT). Together they show the field converging on parametric/recurrent context over ever-larger attention windows.

Limitations (reviewer): LIBERO-only evaluation; SSM extrapolation on real multi-hour prompts and contact-rich tasks untested; Longhorn's fixed-size state may bottleneck very information-dense demos.

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally