-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 VLA Forgetting
Pretrained VLAs are Surprisingly Resistant to Forgetting β continual learning in large-scale Vision-Language-Action models
Venue: ICML 2026 (Oral) Category: Analysis-Insight Affiliations: Huihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu, Yuke Zhu (2026) Traction (2026-06): 4 citations (arXiv)

Continual learning β acquiring new skills over time without catastrophically forgetting old ones β has been studied mostly in small behavior-cloning (BC) policies trained from scratch, where forgetting is severe. Whether the same dynamics hold for modern large-scale pretrained Vision-Language-Action (VLA) models was underexplored. This empirical study asks whether the conventional stabilityβplasticity trade-off still governs VLAs, and what role large-scale pretraining plays.
The authors run continual-learning experiments on LIBERO (four task suites: Spatial, 10, Object, Goal), training a separate model per suite with a fixed task ordering and carrying weights across tasks. They evaluate two pretrained VLA backbones β Pi0 and GR00T N1.5 (differing in architecture, parameter count, and pretraining data) β against a non-pretrained BC-Transformer (plus BC-ViT, BC-Diffusion-Policy). The main continual-learning strategy is simple Experience Replay (ER): at task k, training mixes the current task with a buffer of M randomly sampled past-task transitions (default M=1000). Metrics are average success rate (SR) and Negative Backward Transfer (NBT) β positive NBT = forgetting, β€0 = retention or positive transfer.
To isolate pretraining, they compare three Pi0 variants β VL+Action (full robot-data pretraining on a PaliGemma backbone), VL-only (PaliGemma backbone, no robot pretraining), and from scratch β sweeping replay-buffer size. A knowledge-retention probe (Figure 6) decomposes the model into a vision-language (VL) backbone and an action head, then swaps components across training stages and measures recovery speed when re-finetuning a "forgotten" task.

Pretrained VLAs with ER achieve near-zero or even positive backward transfer across LIBERO suites β learning new tasks can improve prior-task performance, challenging the classic stabilityβplasticity trade-off. The effect is consistent across both Pi0 and GR00T N1.5, suggesting it is a general property of pretrained VLAs rather than an architecture artifact. The advantage is starkest in low-replay regimes: at a 2% buffer (100 samples/task), pretrained VLAs hold NBT around 0.1β0.2 while non-pretrained baselines deteriorate to 0.4β0.5 (2β4Γ more forgetting), and the small models need ~20% replay to match. A Pareto-frontier analysis (forgetting vs. buffer size) shows robot-data pretraining shifts the curve toward zero forgetting. Ablations indicate the training objective barely matters (Pi0 flow-matching vs. β2 on LIBERO-Spatial: NBT β0.0003 vs. 0.016), whereas model size does β from-scratch Pi0 NBT improves from 0.110 (17M LLM + ResNet vision) to β0.052 (250M + SigLIP-So400M/14). Finally, the component-swap study shows seemingly forgotten skills are retained, not erased: re-finetuning recovers prior-task performance rapidly (Pi0 recovers within ~20% of the training steps that BC-Transformer needs).
The paper reframes continual learning for VLAs: large-scale pretraining fundamentally changes the dynamics, making simple Experience Replay a surprisingly strong baseline that can reach zero forgetting at small buffer sizes. The insight that degraded performance reflects suppressed but retained knowledge β recoverable with minimal finetuning β suggests practitioners need far less replay data than the from-scratch literature implies, and that scale, not specialized anti-forgetting regularizers (e.g., EWC), drives robustness.
- arXiv: 2603.03818
- ICML 2026: https://icml.cc/virtual/2026/oral/71117
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)