Skip to content

ICLR 2026 VITA

Heungwoo edited this page Jun 1, 2026 · 3 revisions

VITA β€” Zero-Shot Value Functions via Test-Time Adaptation of VLMs

Venue: ICLR 2026 Authors: Christos Ziakas, Alessandra Russo (Imperial College London) Source: arXiv 2506.10085 Β· project page Β· OpenReview Category: RL for VLA β€” Reward Modeling Trend tag: Trend 3

Approach diagram

flowchart LR
  Traj[Trajectory frames<br/>+ goal description] --> TTA[Test-time adaptation<br/>gradient step on meta-learned<br/>self-supervised loss, no labels]
  VLM[Frozen CLIP encoder +<br/>small adaptation MLP] --> TTA
  TTA --> AdaptedVLM[Adapted module<br/>= zero-shot goal-conditioned<br/>value function]
  AdaptedVLM --> R[Estimated task progress<br/>= reward signal]
  R --> RL[Reward shaping for<br/>offline RL policy]
Loading

Problem

RL needs reward functions, but hand-designing rewards for each new task is slow and error-prone. Learned rewards typically need lots of labeled data, defeating the purpose of zero-shot deployment.

Method

Turn a frozen contrastive VLM (OpenCLIP ViT-B/32) into a zero-shot goal-conditioned value function that estimates task progress from a goal description and current observation. The key idea is test-time adaptation (TTA): a small adaptation module (a two-layer residual MLP with GELU, projection dim dβ€²=64) is updated at inference via a gradient step on a meta-learned self-supervised loss. The self-supervised objective is itself meta-learned (gradient-based meta-learning) so that each test-time update provably improves downstream value estimation, rather than relying on a hand-picked auxiliary task.

By applying these updates sequentially over a trajectory, VITA encodes history into the module's parameters, giving the otherwise frame-independent CLIP encoder temporal reasoning. To prevent the adaptation from latching onto spurious shortcuts, training uses a dissimilarity-based sampling strategy that selects semantically diverse trajectory segments.

Results

  • Real-world manipulation (Value-Order Correlation, VOC): generalizing from a single training environment, VITA scores 0.782 in-distribution, 0.725 under environment shift, and 0.820 under embodiment shift β€” far above the autoregressive-VLM baseline GVL (Gemini 1.5 Pro), which sits around 0.21–0.31 across the same settings.
  • Meta-World MT10 offline RL: using VITA's zero-shot value estimates for reward shaping yields a multi-task policy with IQM 0.815 [0.785, 0.838], exceeding both CLIP-based baselines (VLM-CL, VLM-RM, CLIP-FT, CLIP-GRU) and the simulator's hand-engineered fuzzy-logic dense rewards (0.779).

The decisive result is that a zero-shot VLM-derived reward beats a task-specific hand-designed dense reward, validating that VLMs encode enough task-progress structure to serve as reward models β€” but only once TTA supplies the missing temporal/generalization signal.

Significance

Removes the reward-design bottleneck that limited RL-based VLA fine-tuning. Combined with VLA-RFT (environment) and PLD (training loop), VITA provides the missing third component β€” the reward.

Links

Related pages

← Back to ICLR-2026 Β· Topic: RL

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally