-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 LDA 1B
Venue: RSS 2026 (Imitation Learning session) Β· Authors: Jiangran Lyu, Kai Liu, Xuheng Zhang, Haoran Liao, β¦ Ming-Yu Liu, Zhizheng Zhang, et al. (PKU / Galbot / NVIDIA lineage) Β· arXiv: 2602.12215 Category: World-model-based robot foundation model Trend tag: RSS 2026 thread 3 β video/world models vs the VLA backbone
Compiled from the verified RSS 2026 abstract; arXiv ID cross-confirmed via Qwen-RobotManip's reference list.

Figure 1 of the LDA-1B paper. Left: the three-tier data taxonomy of EI-30k β high-quality action data (real/sim robot + human), suboptimal/noisy action data, and actionless human videos, 30k+ hours in a unified LeRobot-style format with aligned coordinate systems. Center: each tier feeds a different objective β all tasks for high-quality data, dynamics for noisy data, forecasting for actionless video β into one model (the figure states 1.6B parameters, ~10Γ prior UWM instantiations), with visual forecasting in DINO latent space rather than pixels. Right: headline real-robot comparisons vs Ο0.5 β contact-rich 42β63, dexterous 37β85, long-horizon 27β50 β and the data-efficient fine-tuning result (adding non-expert data: Ο0.5 drops 55β40 while LDA-1B rises 60β70).
Behavior-cloning robot foundation models imitate expert actions but discard the transferable dynamics knowledge embedded in heterogeneous embodied data β failures, suboptimal rollouts, action-free video. The Unified World Model (UWM) formulation could exploit all of it, but prior instantiations don't scale to foundation level: coarse data usage, fragmented datasets, pixel-space redundancy.
- Joint objectives: dynamics learning + policy learning + visual forecasting in one model, with distinct roles assigned to data of different quality tiers (expert data supervises the policy; low-quality data still teaches dynamics).
- EI-30k: a standardized embodied-interaction corpus of >30,000 hours of human and robot trajectories in a unified format.
- Structured DINO latent space for dynamics β avoids modeling pixel-level appearance, which is what previously blocked scaling.
- Mixed-frequency multi-modal diffusion transformer handling asynchronous vision and action streams; stable at 1B parameters.
- Outperforms prior methods (incl. Ο0.5) by up to +21% (contact-rich), +48% (dexterous), +23% (long-horizon) in simulation and real world.
- Data-efficient fine-tuning: +10% by ingesting the 30% of low-quality trajectories that BC pipelines normally discard as harmful.
- Code and data release stated.
The strongest scaling result yet for the "learn dynamics, not just actions" camp (Review-World-Models) β and the direct empirical counter to pure behavior cloning at foundation scale. The low-quality-data result operationalizes what LBM-style curation studies only hint at: bad trajectories are dynamics data. Cited as a representation-alignment reference by Qwen-RobotManip ("LDA-1B... universal embodied data ingestion"); sits alongside mimic-video as RSS 2026's two-pronged challenge to the VLM-backbone monopoly.
β RSS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)