Skip to content

ICRA 2026 Token Pruning

hwoo.han edited this page Jun 11, 2026 · 2 revisions

Differentiable Token Pruning for Efficient VLA β€” "The Better You Learn, the Smarter You Prune"

Venue: ICRA 2026 Β· Authors: Titong Jiang, Xuefeng Jiang, Yuan Ma, Xin Wen, Bailin Li, Kun Zhan, Peng Jia, Yahui Liu, Sheng Sun, Xianpeng Lang Β· arXiv: 2509.12594 Category: Efficient VLA (token pruning) Trend tag: Differentiable token pruning

The method is called LightVLA. Built on OpenVLA-OFT, it learns which visual tokens to keep β€” pruning is driven by task performance rather than heuristics, so accuracy improves while compute drops.

Approach diagram

flowchart LR
  IMG[Input images] --> ViT[ViT encoder<br/>visual tokens H_v]
  LANG[Language tokens H_l] --> Q
  ViT --> Q["Query generation<br/>Q = softmax(H_v H_l^T/√D) H_l"]
  Q --> S[Token scoring<br/>S = Q H_v^T / √D]
  S --> GS[Gumbel-softmax selection<br/>hard argmax + soft path<br/>differentiable in backward]
  GS --> KEEP[Keep informative tokens<br/>~78 avg vs full set]
  KEEP --> LLM[VLA backbone<br/>OpenVLA-OFT]
  LLM --> ACT[Robot action]
Loading

Problem

VLA inference is bottlenecked by attention over large sets of visual tokens. Prior visual-token pruning was designed for VLMs and, when ported to VLA, tends to underperform: it relies on heuristic keep-ratios and "magic numbers" that ignore which tokens actually matter for action execution.

Method

  • Dynamic queries are generated by cross-attention between visual and language tokens, then used to score every visual token's importance.
  • Gumbel-softmax selection makes the discrete keep/drop decision differentiable end-to-end, combining a hard argmax (forward) with a soft path (backward), so the model learns an adaptive, input-dependent token budget during fine-tuning.
  • No magic numbers, no extra trainable parameters in the base variant β€” keeping it compatible with standard inference frameworks. A LightVLA* variant adds optional learnable query / LayerNorm weights.

Results

On the LIBERO benchmark, built on OpenVLA-OFT:

  • Success rate: 97.4% avg (Spatial 98.4 / Object 98.4 / Goal 98.2 / Long 94.6) vs OpenVLA-OFT 94.5% β†’ +2.9%.
  • Efficiency: βˆ’59.1% FLOPs, βˆ’38.2% latency, retaining ~78 visual tokens on average.
  • vs other pruning methods (success rate): FlashVLA 73.7%, SP-VLA 74.9%, VLA-Cache 74.7% β€” all far below LightVLA's 97.4%, showing VLM-oriented pruners degrade on VLA tasks.

Significance

LightVLA inverts the usual efficiency/accuracy trade-off: by making pruning performance-driven and differentiable, it simultaneously cuts compute and raises success rate. It is the parameter-free, hyperparameter-free counterpart to scheduler- and quantization-based VLA efficiency lines, slotting cleanly into the small/efficient VLA category.

Links

Related pages

← Back to ICRA-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally