-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 FOCA
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Duc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho, Doanh Le Thien, Quang Nguyen, Thien-Loc Ha, Tran Van Nhiem, Bao Thach, An Thai Le, Daniel Sonntag, Mathias Niepert, Vien Ngo, et al.
VisionβLanguageβAction (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. When only a handful of demonstrations are available per task, adaptation quality degrades sharply β the model struggles with long-horizon reasoning and tends to overfit the short, narrow demonstration set rather than learning where the task is headed.
FOCA (Future-Oriented Conditioning for data-efficient Adaptation) augments VLA fine-tuning with a forward-looking objective that combines two complementary signals:
- Explicit prediction of task-grounded future interaction embeddings β the model is trained to anticipate compact embeddings of upcoming interactions rather than raw pixels.
- Implicit alignment to future goal observations, giving the policy long-horizon reasoning without any expensive pixel-level video prediction.
Conceptually, this amounts to learning a future-conditioned, value-like representation: the policy is shaped by where the trajectory is going, not just the immediate next action. Crucially, the framework also supports action-free co-training with synthetic videos generated by video world models, letting FOCA absorb additional future-dynamics signal without needing extra teleoperated action labels.
flowchart LR
O[Current obs + language] --> VLA[VLA backbone]
VLA --> A[Action head]
VLA --> F[Future interaction embedding predictor]
F -. explicit prediction .-> FE[Task-grounded future embedding]
VLA -. implicit alignment .-> G[Future goal observation]
WM[Video world model<br/>synthetic videos] -. action-free co-training .-> F
- 95.7% success with only 20 demonstrations on LIBERO, demonstrating strong data efficiency in the few-shot regime.
- 7β12% improvements on RoboCasa over baselines.
- Up to 26% absolute gains on real robots, indicating the future-oriented conditioning transfers beyond simulation.
These results consistently show that conditioning on future interaction/goal representations recovers much of the performance lost when demonstrations are scarce.
FOCA targets one of the most practical bottlenecks for deploying VLAs: adapting a pretrained model to a new task from only a few demonstrations. By replacing costly pixel-prediction world-modeling with lightweight future-embedding prediction and goal alignment β and by allowing action-free co-training on synthetic world-model videos β FOCA offers a data-efficient adaptation recipe that improves both simulated and real-robot few-shot imitation.
- ICML 2026: https://icml.cc/virtual/2026/poster/66754
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)