-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2026 VLS
Venue: CoRL 2026 (Austin, TX, Nov 9β12). Paper: arXiv 2602.03973. Representative of: training-free inference-time policy steering β keep the policy frozen; a VLM-generated reward steers action denoising. Companions: RL for VLA Β· Real-Time Execution Β· CoRL 2026 survey.
Authors: Shuo Liu, Ishneet Sukhvinder Singh, Yiqing Xu, Jiafei Duan, Ranjay Krishna.

Pretrained diffusion / flow-matching policies fail when the same task moves near an obstacle, onto a shifted support surface, or into mild clutter. These failures are not missing motor skills β they expose imitation learning's coupling of action generation to training-specific spatial configurations and task specifications. Retraining or fine-tuning is costly and conceptually misaligned: the needed behavior already exists in the policy but cannot be selectively summoned at test time.
Vision-Language Steering (VLS) treats adaptation as an inference-time control problem over a frozen generative policy β no parameter updates. Given out-of-distribution observation-language inputs, VLS:
- Grounds the OOD scene into task-relevant 3D keypoints (SAM + DINOv2 features).
- Has a VLM synthesize a stage-aware, trajectory-differentiable reward as PyTorch operations.
- Steers denoising with the reward: gradient-based refinement (MCMC), RBF repulsive forces for trajectory diversity, and gradient-free FeynmanβKac resampling.
- Runs closed-loop, with adaptive guidance strength and Schmitt-trigger stage switching driven by reward feedback.
- CALVIN: ~+31% success over the frozen base policy; 94% on movable objects and 87% on articulated parts, beating DynaGuide and ITPS by ~15β25 points.
- LIBERO-PRO: Οβ.β + VLS reaches 36.81% overall, up to +13% under spatial / semantic perturbations vs. the frozen baseline.
- Real Franka: +19% in-distribution (69% avg); holds up under appearance shift; 40% on a novel-mug substitution where the baseline fails.
VLS shows that much of a pretrained policy's OOD "failure" is a retrieval problem, not a capability gap: a VLM-authored reward can steer denoising to recover latent skills without any gradient step on the policy. That reframes robustness as an inference-time, plug-in layer usable on top of any frozen diffusion/flow VLA.
Limitations (reviewer): batch sampling, MCMC steps, and FeynmanβKac resampling add heavy per-step inference overhead in the denoising loop; quality hinges on the VLM's reward synthesis and the accuracy of keypoint grounding.
- arXiv 2602.03973
- Survey: CoRL 2026 Β· Related: RL for VLA Β· Real-Time Execution
β Back to CoRL 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)