-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 Sparse ActionGen
Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Tsinghua University Traction (2026-06): 2 citations (arXiv)

Diffusion Policy is the dominant action-generation head for visuomotor and VLA control because it models multi-modal action distributions, but its multi-step denoising is too slow for real-time control. The paper notes that on an RTX 4090, 50 denoising steps at ~1 ms/step take 50 ms, capping execution at 20 Hz β far below the 50β1000 Hz a Franka arm needs. Existing caching-based accelerators (e.g. EfficientVLA's uniform schedule, BAC's task-specific block-wise schedule) rely on static caching schedules fixed offline that do not adapt to the dynamics of robotβenvironment interaction. A leave-one-out study on the Square task (Figure 1) shows that a fixed schedule performs inconsistently across rollout iterations and that the optimal schedule differs per iteration, so any static schedule constrains the performance/efficiency tradeoff.
Sparse ActionGen (SAG) is a rollout-adaptive prune-then-reuse mechanism operating over three nested levels β rollout (robotβenvironment interaction), denoising (diffusion inference), and block (DiT forward). In each rollout iteration it globally identifies prunable computations and substitutes them on the fly with cached activations.
-
Real-time diffusion pruner. SAG parameterizes a pruner
G_Οthat, instead of profiling post-forward activation similarities (which would negate any speedup), learns to predict the sparsity pattern a priori. Because the computational pattern of action generation is strongly correlated with the visual input, the pruner is conditioned on the current observationo_t, giving environment-aware adaptation. It is built with a parameter- and inference-efficient design that serves all blocks with a single network in a single forward pass for the whole denoising trajectory. - Global sparsity loss. An end-to-end objective guides the pruner to non-uniformly allocate compute across both timesteps and blocks under a strict budget, challenging the standard block-wise caching paradigm and its overlooked inter-block redundancy.
- One-for-all reusing strategy. Motivated by strong cross-block activation similarity (Figure 4a), SAG reuses cached activations across both blocks and timesteps in a zig-zag manner, minimizing global redundancy. The model parameters are never updated.
On Diffusion Policy (transformer variant, DP-T) over robomimic and Kitchen tasks (Lift, Can, Square, Transport, Tool hang, Kitchen), SAG prunes over 90% of computations and achieves a 3.6β4Γ speedup without sacrificing performance:
- Proficient Human (PH) data (Table 1): SAG reaches Lift 100, Can 98, Square 89, Transport 85, Tool 50 β all at ~3.4β3.7Γ speedup β with an average performance gain of 13% over the full-precision baseline, matching or beating BAC and far exceeding EfficientVLA/L2C.
- Mixed Human (MH) data (Table 2): average performance gain of 29%, with >3.7Γ speedup across tasks (e.g. Square 79, Transport 50).
- Multi-stage Kitchen (Table 3): a lossless 4.03Γ speedup, retaining p1βp4 success of 100/100/100/99, where competitors like CP collapse at high speedups and EfficientVLA fails entirely.
Ablations confirm each component contributes β the real-time pruner, the one-for-all reusing strategy, and the global sparsity loss β and a real-world pick-and-release task (Figure 5) validates improved inference frequency at maintained success.
SAG reframes diffusion-policy acceleration from offline static caching to an online, observation-conditioned pruning problem aligned with closed-loop control. By predicting sparsity before inference and reusing activations across both timesteps and blocks, it delivers near-lossless ~4Γ speedups without retraining, making diffusion policies far more practical as high-frequency action heads for robotics and VLA systems.
- arXiv: 2601.12894
- ICML 2026: https://icml.cc/virtual/2026/poster/65503
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)