Skip to content

ICLR 2026 Policy Contrastive Decoding

Heungwoo edited this page Jun 1, 2026 · 1 revision

Policy Contrastive Decoding β€” training-free decoding to kill spurious visual correlations in robot policies

Venue: ICLR 2026 Β· Authors: Shihan Wu, Xu Luo, Ji Zhang, Junlin Xie, Jingkuan Song, Heng Tao Shen, Lianli Gao Β· arXiv:2505.13255 Β· Category: reasoning / inference-time decoding for VLA Β· Trend tag: generalization & robustness.

Approach diagram

flowchart LR
  Img[Original visual input] --> Pol1[Robot policy]
  ImgM[Object-masked visual input] --> Pol2[Same policy]
  Pol1 --> D1[Action distribution p_orig]
  Pol2 --> D2[Action distribution p_masked]
  D1 --> C{Contrast: p_orig - p_masked}
  D2 --> C
  C --> Act[Object-grounded action]
Loading

Problem

Generalist robot policies (robotic foundation models) tend to learn spurious correlations from pre-training trajectories β€” e.g., latching onto backgrounds or table layouts rather than the task-relevant object. This hurts generalization beyond the training distribution.

Method

Policy Contrastive Decoding (PCD) redirects the policy's focus toward object-relevant visual clues by contrasting two action probability distributions: one from the original image and one from an object-masked image. The difference amplifies action components that genuinely depend on the manipulated object and suppresses background-driven ones.

PCD is training-free and works as a plug-in: it requires no fine-tuning and no access to model weights, so it can wrap heterogeneous policy types β€” both autoregressive and diffusion-based.

Results

  • Evaluated on three open-source policies: OpenVLA (autoregressive), Octo (diffusion), and Ο€β‚€ (diffusion/flow).
  • On Ο€β‚€: +8.9% in simulation and +108% in real-world manipulation (relative improvement reported by the authors).
  • Consistent gains across policy families, in both simulation and real-world settings.

Significance

Demonstrates that contrastive-decoding ideas from VLM hallucination mitigation transfer to action models, offering a cheap, model-agnostic robustness boost without retraining β€” a practical lever for deploying existing robot foundation models more reliably.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally