-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 LAGEA
LAGEA: Language Guided Embodied Agents for Robotic Manipulation β Turning VLM self-reflections into time-grounded shaping rewards
Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: University of Dhaka, Bangladesh Traction (2026-06): 1 citation (arXiv)

Robotic manipulation increasingly benefits from foundation models that describe goals, but agents still lack a principled way to learn from their own mistakes. Sparse-reward, long-horizon tasks make exploration brittle: dense reward shaping from vision-language models (VLMs) can destabilize training or invite reward hacking, while contrastive reward-alignment approaches such as FuRL can suffer when early misalignment compounds and misdirects exploration. LaGEA asks whether natural language can instead serve as an error-reasoning signal β feedback that helps an embodied agent diagnose what went wrong and correct course.
LaGEA (Language Guided Embodied Agents) turns episodic, schema-constrained reflections from a VLM into temporally grounded guidance for reinforcement learning. The pipeline has four stages:
-
Keyframe selection β after each rollout, causal moments in the trajectory are identified and per-step weights
$\hat{w}_t$ are computed, localizing the decisive frames. - Schema-constrained self-reflection β a VLM is queried on those frames and returns a concise, structured language summary of what happened (an error taxonomy constrains the format).
- Visual-language alignment β feedback is aligned with visual state in a shared representation, co-trained with BCE / InfoNCE objectives so the embedding space becomes control-relevant.
-
Delta-based shaping rewards β a Goal Potential
$\phi_t$ aligns the current state$z_t$ with the goal image$z_g$ and instruction$z_y$ , and a Feedback Potential$\psi_t$ measures agreement with the reflection. These are converted into bounded, step-wise shaping rewards.

The combined shaping term is modulated by an adaptive, failure-aware coefficient
On the Meta-World MT10 embodied manipulation benchmark (average success across five random seeds), LaGEA improves average success over state-of-the-art methods by 9.0% on random goals and 5.3% on fixed goals, while converging faster. Baselines include SAC, LIV, LIV-Proj, Relay, and FuRL (with and without goal image). Across eight Meta-World tasks (Figure 3), LaGEA reaches high success in far fewer environment steps than FuRL and SAC, which plateau late or stall. Ablations confirm that (a) structured feedback beats free-form feedback, (b) keyframe selection matters (drawer-open study), and (c) removing any of
LaGEA supports the hypothesis that language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes. Rather than using a VLM as a one-shot reward labeler, it converts episodic reflections into bounded, decaying, failure-aware shaping signals β a recipe that improves both sample efficiency and final success on sparse-reward manipulation.
- arXiv: 2509.23155
- ICML 2026: https://icml.cc/virtual/2026/poster/60801
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)