Skip to content

ICML 2026 Speedup Patch

hwoo.han edited this page Jun 11, 2026 · 1 revision

Speedup Patch (SuP) β€” Plug-and-play offline-RL acceleration for frozen manipulation policies

Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Zhichao Wu, Junyin Ye, Zhilong Zhang, Yihao Sun, Haoxin Lin, Jiaheng Luo, Haoxiang Ren, Lei Yuan, Yang Yu (Nanjing University) Traction (2026-06): 2 citations (arXiv)

Plug-and-Play Speedup via Scheduler Policy: a scheduler predicts a downsampling rate k that compresses the frozen policy's action chunk into a shorter chunk (Figure 1 from Wu et al., 2026)

Problem

Modern embodied policies (ACT, diffusion policies, VLAs such as Ο€0.5) inherit the "tardy pacing" of the human teleoperation data they imitate, so even when they succeed they execute slowly. Existing acceleration methods either retrain the policy on entropy-resampled demonstrations (e.g., DemoSpeedup) or require costly online interaction β€” both of which do not scale to large frozen foundation models whose weights are downloaded and never touched. The paper asks: can we accelerate an arbitrary, frozen embodied policy using only offline data and no policy retraining?

Method

SuP (SpeedUp Patch) wraps a frozen base policy Ο€_base with a lightweight external scheduler that adaptively decides, per action chunk, a downsampling rate k β€” keeping every k-th action to produce a shorter chunk. The key is choosing k aggressively where motion is redundant but conservatively near contact-rich or precise phases.

The authors formalize scheduler learning as a Constrained Markov Decision Process (S, K, P, r, c, h, Ξ³): maximize an efficiency reward while keeping a safety cost below threshold. Because true task success cannot be evaluated offline, SuP introduces a world-model-based state deviation surrogate: a learned (recurrent) world model predicts the counterfactual trajectory the un-downsampled policy would have produced, and the deviation between downsampled and counterfactual states acts as the constraint cost. Training is purely offline in three phases β€” (1) recurrent world-model learning, (2) data synthesis, (3) scheduler optimization via IQL (offline RL).

The three-phase offline training process of SuP: recurrent world model learning, data synthesis, and IQL scheduler optimization (Figure 3 from Wu et al., 2026)

Results

Evaluated across diverse frozen architectures (ACT and Diffusion Policy on BiGym humanoid tasks; pre-trained Ο€0.5 and VLA-Adapter on LIBERO), plus three real-world dual-arm (Aloha-like) tasks. Cells report success rate and average steps-to-completion.

  • Overall ~1.8Γ— execution speedup while preserving original success rates.
  • LIBERO (Ο€0.5): base 0.969 avg SR at 1.00Γ—; SuP reaches 1.94Γ— on Long suite and maintains comparable success, vs. fixed downsample (-ds2) which drops to 0.928 SR at 1.72Γ—.
  • BiGym (ACT): SuP improves both success and speed (e.g., Sandwich 0.45β†’0.64 SR), outperforming Vanilla Downsample and DemoSpeedup, which suffer >5% success drops.
  • Real-world (Ο€0.5, dual-arm): average success 0.589β†’0.611 at 2.17Γ— speedup, beating fixed -ds2/-ds3 (which collapse to 0.356 SR at 2.19Γ—) and matching DemoSpeedup on speed without retraining.

A case study links violation count (world-model deviation events) to conditional success-rate drops, validating the surrogate cost.

Significance

SuP decouples acceleration from the policy itself: a small scheduler trained offline can be patched onto any frozen manipulation policy β€” including downloaded VLA foundation models β€” to roughly halve execution time without sacrificing success or touching the base weights. The CMDP-plus-world-model formulation gives a principled, retraining-free path to deploy-time efficiency.

Links

← Back to ICML-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally