Skip to content

ICML 2026 TapSampling

hwoo.han edited this page Jun 11, 2026 · 1 revision

TapSampling β€” Inference-time sampling with a task-progress verifier for robot policies

Venue: ICML 2026 (Poster) Category: Efficiency / Inference-Time Scaling Affiliations: (from arXiv) Sizhe Zhao, Shengping Zhang, Shuo Yang, Weiyu Zhao, Shuigen Wang, Xiangyang Ji Traction (2026-06): 0 citations (arXiv)

TapSampling extends single-shot inference with multi-sample generation and task-progress-guided verification (Figure 1 from Zhao et al., 2026)

Problem

Generalist robot policies built on diffusion or autoregressive (next-token) backbones are non-deterministic: under identical conditions they succeed on some trials and fail on others. Because deployment typically commits to a single action per decision step (single-shot inference), there is no mechanism to correct this stochastic variance, fundamentally capping performance. A motivating study on CALVIN with VPP shows the policy reaches 4.39 average success length without retries but climbs to ~4.72 when allowed up to 4 retries β€” evidence of a large untapped upper bound. The paper asks whether inference-time computation, rather than more training data or parameters, can close this gap. Doing so requires two hard components: (1) a candidate sampling strategy that produces plausible alternatives cheaply, and (2) a verifier that can judge low-level actions β€” for which no off-the-shelf evaluator exists.

Method

TapSampling is a plug-and-play, policy-agnostic framework with two parts.

TapSampling overview: Action-VAE samples candidates from a learned posterior; a task-progress verifier scores them (Figure 2 from Zhao et al., 2026)

Action sampling with a learned posterior. An Action-VAE (transformer encoder/decoder) compresses each policy-generated action chunk into a low-dimensional Gaussian posterior, trained with reconstruction + KL loss. At inference, a small set of N=4 policy actions is encoded into a mixed posterior q_mix = (1/N) Ξ£ q_E(z|a_i), from which an arbitrary number of latents are drawn and decoded into candidates. Because the posterior captures intra-chunk dimensional correlations, candidates stay close to the true policy distribution β€” and sampling is ~5Γ— faster than re-querying the policy.

Action selection via task-progress verification. Under a linear-progress assumption, each timestep i in a trajectory of length t is labeled p_i = i/t. Forward action sub-sequences yield positive labels (+k/t); reversed sub-sequences yield negatives (βˆ’k/t), giving automatic positive/negative pairs with no manual annotation. A verifier built on VLA-Adapter (Qwen2.5-0.5B backbone) regresses the task-progress change Ξ”p of a query action via L1 loss. At inference, candidates below a threshold are discarded and a score-weighted average of survivors is executed. The VLM backbone runs once; candidates are scored in a batch through the lightweight head.

Results

  • CALVIN ABCβ†’D (avg. success length): Diffusion Policy 2.41β†’2.58, OpenVLA 3.30β†’3.51, VPP 4.39β†’4.46 β€” consistent plug-and-play gains across diffusion, autoregressive, and video-prediction policies.
  • LIBERO-Long: Ο€0.5 improves 96.8%β†’98.0%.
  • Real-world (Franka FR3, 3 task families): Ο€0 average 78.3%β†’83.3%, with the largest gains on Stack-Unseen (+10.0) and Knock-Down-Unseen (+6.6).
  • Efficiency: Policy Sampling gives the best quality (avg. len 4.50) but ~20Γ— latency; Gaussian and Learned-Posterior Sampling add <0.01 s. Learned-Posterior best balances the two (4.46) and yields lower MMD to the policy distribution than Gaussian (e.g., 0.064 vs 0.098 at Ξ³=2.0). Verification is ~12Γ— faster than RoboMonkey at k=16.

Significance

TapSampling reframes test-time scaling β€” well established for LLMs and image diffusion β€” for robotic control, supplying the two missing ingredients (a distribution-faithful sampler and an interpretable, semantically grounded action verifier) without fine-tuning the base policy. The progress-change verifier is notable for producing scores with explicit meaning (negative = impedes task, larger positive = faster completion), enabling interpretable selection and reusable across heterogeneous policies.

Links

← Back to ICML-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally