-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Visual Verification Enables Inference time Steering
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 1 Β· paper #79 Authors: Mingtong Zhang, Dhruv Shah (Princeton University) arXiv: 2606.18247 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: Overview of VERITAS. A pre-trained generalist policy (left) acts as a stochastic generator that samples multiple short-horizon action chunks per decision step; a gradient-free visual verifier scores the candidates and the best-of-N action is executed ("Inference-Time Steering"). Successful verifier-guided rollouts are logged as "Verified Rollouts" and reused to fine-tune the policy ("Policy Improvement"), forming a self-improvement flywheel with minimal human supervision.
Robot foundation models are trained almost entirely on human-expert demonstrations, so improving them scales linearly with human labor. The paper asks how to improve generalist policies without collecting more human data, by letting the robot practice and learn from its own experience.
The paper proposes VERITAS, a generatorβverifier framework. A pre-trained generalist policy is treated as a "generator" that samples diverse candidate action chunks; a gradient-free "visual verifier" (VLM-based or heuristic) scores candidates on task alignment and physical plausibility, and the highest-scoring action is executed via best-of-N selection β steering the policy at inference time with no parameter updates. Verifier-approved rollouts are then logged and used as on-policy supervision to fine-tune the base policy for offline improvement. Experiments use Ο0-Bridge in SimplerEnv simulation and Ο0-DROID / Ο0.5-DROID on a real FR3-DROID platform, benchmarked against a V-GPS-DROID learned-value baseline.
Across simulation and real-world settings (3 policies, 1160 total evaluation episodes), verifier-guided execution improved success rates by an average of 12.6% in simulation and 35% in real-world deployment with no fine-tuning. For offline improvement, VERITAS fine-tuning raised average success by 9.7% over the base Ο0-Bridge policy across 4 SimplerEnv tasks (960 episodes), with the largest gain on Stack Blocks (31.3% β 59.2%, +27.9%). Post-training on verified autonomous rollouts matched the efficiency of learning from human-expert demonstrations, without any human intervention.
Positions inference-time verification as a scalable, policy- and verifier-agnostic mechanism for autonomous policy improvement, complementing test-time-scaling and co-training trends in Review-LBM-Cotraining and Review-Human-Video-Transfer.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)