-
Notifications
You must be signed in to change notification settings - Fork 0
Review Stellar VLA
In-Depth Review β Stellar VLA: Continually Evolving Skill Knowledge in Vision-Language-Action Models
Paper: Continually Evolving Skill Knowledge in Vision Language Action Model Authors: Yuxuan Wu, Guangming Wang, Zhiheng Yang, Tianchen Deng, Maoqing Yao, Brian Sheil, Hesheng Wang Affiliations: Shanghai Jiao Tong University Β· Shanghai Innovation Institute Β· University of Cambridge Β· Beihang Β· NTU Β· MIT SMART Β· AgiBot arXiv: 2511.18085 (v4 May 8, 2026; v1 Nov 2025, cs.RO) Β· Project: stellarvla.github.io Status: preprint β indexed here under Latest Papers
Companion reviews: RL for VLA Β· LBM Co-training Β· VLA Evaluation Β· VLA Architectures Β· VLA Memory. Related insight: ICML 2026's Pretrained VLAs Resist Forgetting (Oral) β the finding this paper builds against.
- Continual imitation learning (CIL) for VLAs without growing the network. Stellar VLA lets a fixed-size ~1B VLA learn a stream of tasks while retaining prior skills, using only 1% data replay β versus conventional CIL methods that bolt on adapters/task-modules (parameter growth, storage cost) and VLA-replay work that needs ~20% replay.
- A self-evolving knowledge space is the core idea. Task-relevant knowledge is organized as a Dirichlet-Process-based cluster space with an unbounded number of components, so new task/skill clusters emerge automatically as tasks arrive β no predefined component count, no manual annotation of task identity.
- Two variants, flat vs hierarchical. T-Stellar uses a Dirichlet Process Mixture Model (DPMM) for flat task-centric clusters; TS-Stellar extends to a Hierarchical Dirichlet Process (HDP) that models a taskβskill structure where reusable subskills are shared across tasks (task distribution = aggregation over shared skill atoms).
- Knowledge-guided expert routing = specialization at no extra params. A diffusion-based MoE action head (MoDE-style) is routed by the knowledge space β via a knowledge-relation embedding (distance of the task latent to cluster centers, weighted by posterior membership) and Top-K semantic embeddings β instead of MoDE's noise-level routing, giving task-specific parameter sharing/differentiation without increasing model size.
- State-of-the-art CIL on LIBERO + real dual-arm. Against both VLA baselines (MoDE, UniVLA, Ο0, Ο0.5) and CIL baselines (ER, SeqLoRA, LoTUS, IsCiL), Stellar variants take the best AUC and Final SR across LIBERO-goal / -long / -30*, with >50% average AUC/Final-SR improvement in the from-scratch setting and ~20% Final-SR improvement over CIL baselines; on a real dual-arm 7-task suite TS-Stellar reaches 90.0% Final SR (7.4 NBT) vs Ο0.5's 72.9% (25.8 NBT).
- It reframes VLA continual learning as knowledge modeling, not parameter management. The lifelong-robot problem is usually attacked with replay buffers or parameter isolation. Stellar argues the leverage is capturing task/skill relationships so related tasks share parameters and unrelated ones don't interfere β pushing the RL/continual discussion from "how much to replay / how many adapters" toward "what structure to learn."
- It's the constructive answer to the ICML 2026 "VLAs resist forgetting" Oral. That insight paper showed large pretrained VLAs barely forget with simple replay β implicitly challenging complex CIL designs. Stellar takes the challenge seriously (fixed size, minimal replay) but shows a structured knowledge space still buys real gains, especially from scratch and on long-horizon/hierarchical tasks where plain replay is weakest.
- Non-parametric Bayesian structure meets VLA MoE. Bringing DPMM/HDP (unbounded, auto-discovering clusters; memoized variational Bayes for incremental updates) into VLA routing is a genuinely different mechanism from the fixed-expert MoEs elsewhere in Review-VLA-Architecture β the number of task/skill clusters grows with experience rather than being a hyperparameter.
- Hierarchical skill sharing is validated where it should help. TS-Stellar's edge concentrates on long-horizon and compositional manipulation (LIBERO-long, real bimanual handovers), the regime where reusing subskills across tasks is the natural inductive bias.

Figure 2 of the paper. A CLIP/FiLM-conditioned encoder turns language + image into embeddings; a VAE latent encoder/decoder learns a task-centric latent z (reconstruction + DP-aware KL). The Knowledge Space (center) co-evolves with z: new-task latents update Dirichlet-Process clusters while a L_KL term pulls z toward current clusters ("Co-Evolution"). The learned KS-prior routing then guides an MoE Diffusion Transformer action head (Attention + Expert blocks ΓK) that outputs the action chunk. Bottom strip: past-task latents aggregate into clusters; a new task spawns/updates knowledge components.
Standard CIL: tasks {T_j} arrive sequentially, each with expert demos (language, obs, actions); the agent must learn new tasks with limited access to past data while retaining old skills. Stellar uses Experience Replay with a tiny buffer (1% of past demos; 5% for the real-world experiments) and adds structure so that small replay doesn't drift.
- DP prior G βΌ DP(Ξ±, Gβ) supports clustering with an unbounded number of components β clusters are shared across tasks but new ones emerge as tasks evolve.
- T-Stellar (DPMM): task latent z_j βΌ F_task(ΞΈ_j), ΞΈ_j βΌ G β dynamic clustering of task representations.
- TS-Stellar (HDP): each task is a distribution over skills (z_ji βΌ F_skill(ΞΈ_ji), ΞΈ_ji βΌ G_j, G_j βΌ DP(Ξ³,G), G βΌ DP(Ξ±,Gβ)); the task-level Gaussian parameters are aggregated from shared skill atoms with task-specific mixture weights β so subskills transfer across tasks.
A hierarchical variational-inference loop (Algorithm 1): a VAE infers task-centric latents from vision-language input (reconstruction loss L_recon + DP-aware L_KL toward current clusters); periodically the knowledge distribution Ξ is updated from sampled latents via memoized variational Bayes (memoVB) for efficient incremental global-statistic sharing. TS-Stellar decodes task latents β language goals and skill latents β visual observations separately, with an HDP-structured KL. The result is a self-reinforcing cycle that retains old and discovers new task/skill knowledge.
A diffusion MoE action head (Γ la MoDE) where routing is conditioned on the knowledge space instead of denoising level:
- Knowledge-relation embedding f_R = Ξ£_k p_k Β· |z β ΞΌ_k| (posterior membership p_k Γ distance to cluster centers) β a fixed-dim summary despite variable cluster counts.
- Top-K semantic embeddings for the most relevant clusters. These route experts to give task-specific specialization + related-task sharing without adding parameters β the whole model stays ~1B.
Metrics: FWT (forward transfer), NBT (negative backward transfer = forgetting, lower better), AUC (success-rate-curve stability), Final SR (after all tasks). 100 trials Γ 50 init states, 3 seeds on -goal/-long.
| Benchmark | Best VLA baseline (Final SR) | Best CIL baseline | T-Stellar | TS-Stellar |
|---|---|---|---|---|
| LIBERO-goal (scratch) | Ο0 35.7 | β | 67.9 | 64.2 |
| LIBERO-long (scratch) | Ο0 12.0 | β | 34.2 | 35.0 |
| LIBERO-30* (scratch) | MoDE 28.5 | β | 42.9 | 42.6 |
| LIBERO-goal (CIL, pretrained on -90) | ER 55.3 | ER 55.3 | 62.1 | 57.3 |
| LIBERO-long (CIL, pretrained on -90) | ER 16.1 | LoTUS 31.8 | 36.3 | 40.9 |
Reported summary: Stellar variants take best AUC and Final SR across all scenarios; >50% average AUC/Final-SR improvement over all baselines in the scratch setting; ~20% Final-SR improvement over CIL baselines with only 1% replay and no parameter growth (vs LoTUS/IsCiL which add parameters). TS-Stellar leads specifically on long-horizon tasks. Some baselines post low NBT only because their FWT is also very low (the forward/backward transfer trade-off).
| Metric | ER | UniVLA | Ο0 | Ο0.5 | T-Stellar | TS-Stellar |
|---|---|---|---|---|---|---|
| FWT β | 97.1 | 70.0 | 98.6 | 95.7 | 98.6 | 98.6 |
| NBT β | 21.9 | 37.1 | 34.6 | 25.8 | 12.4 | 7.4 |
| AUC β | 79.9 | 43.6 | 72.6 | 75.6 | 89.6 | 93.4 |
| Final SR β | 70.0 | 21.4 | 57.1 | 72.9 | 84.3 | 90.0 |
TS-Stellar shows the lowest forgetting (NBT 7.4) and highest retention (Final SR 90.0) on a new embodiment with compositional/bimanual tasks (e.g., "Handover Toy" after training on "Pull Stick from Bag"), confirming the hierarchical-skill hypothesis transfers to real hardware.
- A structured middle path in the CIL debate. Between "just replay, VLAs barely forget" (ICML Oral) and "add adapters/modules per task" (LoTUS, IsCiL), Stellar keeps the model fixed-size with 1% replay yet recovers large gains β strongest exactly where plain replay is weakest (from-scratch, long-horizon, hierarchical). It refines, rather than overturns, the forgetting-resistance finding.
- Non-parametric knowledge as the routing signal. Auto-discovering task/skill clusters and using them to route a diffusion MoE is a distinct mechanism from fixed-expert or noise-routed MoEs in Review-VLA-Architecture; it grows structure with experience without growing parameters β relevant to the "lifelong VLA" thread the field is opening.
- Ties to the data-efficiency agenda. 1% replay / ~1B params is a storage-and-compute argument aligned with the co-training and evaluation-cost concerns in Review-LBM-Cotraining and Review-VLA-Evaluation.
- LIBERO-30* is single-run (cost), so those numbers carry less statistical weight than the 3-seed -goal/-long results.
- Cross-embodiment pretraining is harder for the method β the paper notes pretraining on LIBERO-90 (aligned embodiment) transfers more cleanly than pretraining on 1k+ cross-embodiment tasks, which needs more parameter updates; the clean-transfer story is strongest within-embodiment.
- Task/skill count and horizon are modest (LIBERO suites + 7 real tasks); very long task streams (hundreds of tasks) aren't tested.
- The DP/HDP + memoVB machinery adds conceptual and implementation complexity; the paper's own framing (against "increasingly complex CIL designs") invites the question of whether the knowledge-space gains justify it versus tuned replay at larger buffers β an ablation of replay-rate Γ structure would sharpen this.
- Backbone is a ~1B CLIP/FiLM+ResNet+diffusion-MoE stack, not a frontier VLM-initialized VLA; whether the knowledge-space benefit persists on top of a strong pretrained VLM (Ο/GR00T-class) is untested.
- Multiple arXiv versions (v1 Nov 2025 β v4 May 2026); numbers here are from the latest revision β treat as an evolving preprint.
- arXiv: https://arxiv.org/abs/2511.18085 Β· Project: https://stellarvla.github.io/
- Pretrained VLAs Resist Forgetting β the forgetting-resistance Oral this builds against
- RL for VLA β continual-learning-under-repeated-adaptation is the open item there
- LBM Co-training β the data-efficiency / mixture evidence base
- VLA Architectures β where knowledge-routed MoE sits among action heads
- Latest Papers β the preprint tracker this is filed under
β Back to Latest Papers Β· Home Β· Reviews
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)