Skip to content

Review Stellar VLA

hwoo.han edited this page Aug 11, 2026 · 1 revision

In-Depth Review β€” Stellar VLA: Continually Evolving Skill Knowledge in Vision-Language-Action Models

Paper: Continually Evolving Skill Knowledge in Vision Language Action Model Authors: Yuxuan Wu, Guangming Wang, Zhiheng Yang, Tianchen Deng, Maoqing Yao, Brian Sheil, Hesheng Wang Affiliations: Shanghai Jiao Tong University Β· Shanghai Innovation Institute Β· University of Cambridge Β· Beihang Β· NTU Β· MIT SMART Β· AgiBot arXiv: 2511.18085 (v4 May 8, 2026; v1 Nov 2025, cs.RO) Β· Project: stellarvla.github.io Status: preprint β€” indexed here under Latest Papers

Companion reviews: RL for VLA Β· LBM Co-training Β· VLA Evaluation Β· VLA Architectures Β· VLA Memory. Related insight: ICML 2026's Pretrained VLAs Resist Forgetting (Oral) β€” the finding this paper builds against.


1. TL;DR

  1. Continual imitation learning (CIL) for VLAs without growing the network. Stellar VLA lets a fixed-size ~1B VLA learn a stream of tasks while retaining prior skills, using only 1% data replay β€” versus conventional CIL methods that bolt on adapters/task-modules (parameter growth, storage cost) and VLA-replay work that needs ~20% replay.
  2. A self-evolving knowledge space is the core idea. Task-relevant knowledge is organized as a Dirichlet-Process-based cluster space with an unbounded number of components, so new task/skill clusters emerge automatically as tasks arrive β€” no predefined component count, no manual annotation of task identity.
  3. Two variants, flat vs hierarchical. T-Stellar uses a Dirichlet Process Mixture Model (DPMM) for flat task-centric clusters; TS-Stellar extends to a Hierarchical Dirichlet Process (HDP) that models a task→skill structure where reusable subskills are shared across tasks (task distribution = aggregation over shared skill atoms).
  4. Knowledge-guided expert routing = specialization at no extra params. A diffusion-based MoE action head (MoDE-style) is routed by the knowledge space β€” via a knowledge-relation embedding (distance of the task latent to cluster centers, weighted by posterior membership) and Top-K semantic embeddings β€” instead of MoDE's noise-level routing, giving task-specific parameter sharing/differentiation without increasing model size.
  5. State-of-the-art CIL on LIBERO + real dual-arm. Against both VLA baselines (MoDE, UniVLA, Ο€0, Ο€0.5) and CIL baselines (ER, SeqLoRA, LoTUS, IsCiL), Stellar variants take the best AUC and Final SR across LIBERO-goal / -long / -30*, with >50% average AUC/Final-SR improvement in the from-scratch setting and ~20% Final-SR improvement over CIL baselines; on a real dual-arm 7-task suite TS-Stellar reaches 90.0% Final SR (7.4 NBT) vs Ο€0.5's 72.9% (25.8 NBT).

2. Why this paper matters

  • It reframes VLA continual learning as knowledge modeling, not parameter management. The lifelong-robot problem is usually attacked with replay buffers or parameter isolation. Stellar argues the leverage is capturing task/skill relationships so related tasks share parameters and unrelated ones don't interfere β€” pushing the RL/continual discussion from "how much to replay / how many adapters" toward "what structure to learn."
  • It's the constructive answer to the ICML 2026 "VLAs resist forgetting" Oral. That insight paper showed large pretrained VLAs barely forget with simple replay β€” implicitly challenging complex CIL designs. Stellar takes the challenge seriously (fixed size, minimal replay) but shows a structured knowledge space still buys real gains, especially from scratch and on long-horizon/hierarchical tasks where plain replay is weakest.
  • Non-parametric Bayesian structure meets VLA MoE. Bringing DPMM/HDP (unbounded, auto-discovering clusters; memoized variational Bayes for incremental updates) into VLA routing is a genuinely different mechanism from the fixed-expert MoEs elsewhere in Review-VLA-Architecture β€” the number of task/skill clusters grows with experience rather than being a hyperparameter.
  • Hierarchical skill sharing is validated where it should help. TS-Stellar's edge concentrates on long-horizon and compositional manipulation (LIBERO-long, real bimanual handovers), the regime where reusing subskills across tasks is the natural inductive bias.

3. Method

Stellar VLA architecture (Figure 2 of arXiv 2511.18085, Β© the authors)

Figure 2 of the paper. A CLIP/FiLM-conditioned encoder turns language + image into embeddings; a VAE latent encoder/decoder learns a task-centric latent z (reconstruction + DP-aware KL). The Knowledge Space (center) co-evolves with z: new-task latents update Dirichlet-Process clusters while a L_KL term pulls z toward current clusters ("Co-Evolution"). The learned KS-prior routing then guides an MoE Diffusion Transformer action head (Attention + Expert blocks Γ—K) that outputs the action chunk. Bottom strip: past-task latents aggregate into clusters; a new task spawns/updates knowledge components.

3.1 Setting

Standard CIL: tasks {T_j} arrive sequentially, each with expert demos (language, obs, actions); the agent must learn new tasks with limited access to past data while retaining old skills. Stellar uses Experience Replay with a tiny buffer (1% of past demos; 5% for the real-world experiments) and adds structure so that small replay doesn't drift.

3.2 Dirichlet-Process knowledge space

  • DP prior G ∼ DP(Ξ±, Gβ‚€) supports clustering with an unbounded number of components β€” clusters are shared across tasks but new ones emerge as tasks evolve.
  • T-Stellar (DPMM): task latent z_j ∼ F_task(ΞΈ_j), ΞΈ_j ∼ G β€” dynamic clustering of task representations.
  • TS-Stellar (HDP): each task is a distribution over skills (z_ji ∼ F_skill(ΞΈ_ji), ΞΈ_ji ∼ G_j, G_j ∼ DP(Ξ³,G), G ∼ DP(Ξ±,Gβ‚€)); the task-level Gaussian parameters are aggregated from shared skill atoms with task-specific mixture weights β€” so subskills transfer across tasks.

3.3 Self-evolution (co-training z and the knowledge space)

A hierarchical variational-inference loop (Algorithm 1): a VAE infers task-centric latents from vision-language input (reconstruction loss L_recon + DP-aware L_KL toward current clusters); periodically the knowledge distribution Θ is updated from sampled latents via memoized variational Bayes (memoVB) for efficient incremental global-statistic sharing. TS-Stellar decodes task latents β†’ language goals and skill latents β†’ visual observations separately, with an HDP-structured KL. The result is a self-reinforcing cycle that retains old and discovers new task/skill knowledge.

3.4 Knowledge-guided expert routing

A diffusion MoE action head (Γ  la MoDE) where routing is conditioned on the knowledge space instead of denoising level:

  • Knowledge-relation embedding f_R = Ξ£_k p_k Β· |z βˆ’ ΞΌ_k| (posterior membership p_k Γ— distance to cluster centers) β€” a fixed-dim summary despite variable cluster counts.
  • Top-K semantic embeddings for the most relevant clusters. These route experts to give task-specific specialization + related-task sharing without adding parameters β€” the whole model stays ~1B.

4. Results

4.1 LIBERO CIL (Final SR, higher better)

Metrics: FWT (forward transfer), NBT (negative backward transfer = forgetting, lower better), AUC (success-rate-curve stability), Final SR (after all tasks). 100 trials Γ— 50 init states, 3 seeds on -goal/-long.

Benchmark Best VLA baseline (Final SR) Best CIL baseline T-Stellar TS-Stellar
LIBERO-goal (scratch) Ο€0 35.7 β€” 67.9 64.2
LIBERO-long (scratch) Ο€0 12.0 β€” 34.2 35.0
LIBERO-30* (scratch) MoDE 28.5 β€” 42.9 42.6
LIBERO-goal (CIL, pretrained on -90) ER 55.3 ER 55.3 62.1 57.3
LIBERO-long (CIL, pretrained on -90) ER 16.1 LoTUS 31.8 36.3 40.9

Reported summary: Stellar variants take best AUC and Final SR across all scenarios; >50% average AUC/Final-SR improvement over all baselines in the scratch setting; ~20% Final-SR improvement over CIL baselines with only 1% replay and no parameter growth (vs LoTUS/IsCiL which add parameters). TS-Stellar leads specifically on long-horizon tasks. Some baselines post low NBT only because their FWT is also very low (the forward/backward transfer trade-off).

4.2 Real-world dual-arm (7 tasks, 5% replay, 10 trials each)

Metric ER UniVLA Ο€0 Ο€0.5 T-Stellar TS-Stellar
FWT ↑ 97.1 70.0 98.6 95.7 98.6 98.6
NBT ↓ 21.9 37.1 34.6 25.8 12.4 7.4
AUC ↑ 79.9 43.6 72.6 75.6 89.6 93.4
Final SR ↑ 70.0 21.4 57.1 72.9 84.3 90.0

TS-Stellar shows the lowest forgetting (NBT 7.4) and highest retention (Final SR 90.0) on a new embodiment with compositional/bimanual tasks (e.g., "Handover Toy" after training on "Pull Stick from Bag"), confirming the hierarchical-skill hypothesis transfers to real hardware.


5. Significance & positioning

  • A structured middle path in the CIL debate. Between "just replay, VLAs barely forget" (ICML Oral) and "add adapters/modules per task" (LoTUS, IsCiL), Stellar keeps the model fixed-size with 1% replay yet recovers large gains β€” strongest exactly where plain replay is weakest (from-scratch, long-horizon, hierarchical). It refines, rather than overturns, the forgetting-resistance finding.
  • Non-parametric knowledge as the routing signal. Auto-discovering task/skill clusters and using them to route a diffusion MoE is a distinct mechanism from fixed-expert or noise-routed MoEs in Review-VLA-Architecture; it grows structure with experience without growing parameters β€” relevant to the "lifelong VLA" thread the field is opening.
  • Ties to the data-efficiency agenda. 1% replay / ~1B params is a storage-and-compute argument aligned with the co-training and evaluation-cost concerns in Review-LBM-Cotraining and Review-VLA-Evaluation.

6. Limitations

6.1 Visible in the paper

  • LIBERO-30* is single-run (cost), so those numbers carry less statistical weight than the 3-seed -goal/-long results.
  • Cross-embodiment pretraining is harder for the method β€” the paper notes pretraining on LIBERO-90 (aligned embodiment) transfers more cleanly than pretraining on 1k+ cross-embodiment tasks, which needs more parameter updates; the clean-transfer story is strongest within-embodiment.
  • Task/skill count and horizon are modest (LIBERO suites + 7 real tasks); very long task streams (hundreds of tasks) aren't tested.

6.2 Reviewer's notes

  • The DP/HDP + memoVB machinery adds conceptual and implementation complexity; the paper's own framing (against "increasingly complex CIL designs") invites the question of whether the knowledge-space gains justify it versus tuned replay at larger buffers β€” an ablation of replay-rate Γ— structure would sharpen this.
  • Backbone is a ~1B CLIP/FiLM+ResNet+diffusion-MoE stack, not a frontier VLM-initialized VLA; whether the knowledge-space benefit persists on top of a strong pretrained VLM (Ο€/GR00T-class) is untested.
  • Multiple arXiv versions (v1 Nov 2025 β†’ v4 May 2026); numbers here are from the latest revision β€” treat as an evolving preprint.

7. Links & related pages

← Back to Latest Papers Β· Home Β· Reviews

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally