Skip to content

ICLR 2026 AutoQVLA

Heungwoo edited this page Jun 1, 2026 · 2 revisions

QVLA β€” Not All Channels Are Equal (Channel-Aware Quantization)

Venue: ICLR 2026 (poster) Category: VLA Architecture β€” Efficiency Trend tag: Efficiency Paper: "QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization" (arXiv 2602.03782). AutoQVLA is the name the paper gives to the automated bit-allocation algorithm; the framework itself is QVLA. Authors: Yuhao Xu, Yantai Yang, Zhenyang Fan, Yufan Liu, Yuming Li, Bing Li, Zhipeng Zhang (Shanghai Jiao Tong University, Chinese Academy of Sciences, Alipay)

Approach diagram

flowchart LR
  W[VLA weights] --> S[Per-channel<br/>action-space sensitivity<br/>Jacobian / Taylor]
  S --> G[Greedy demotion<br/>bits in 0,2,4,8,16]
  G --> H[High-sensitivity channels]
  G --> L[Low-sensitivity channels]
  H --> HK[Keep in high precision]
  L --> LQ[Aggressively quantize<br/>0-bit = prune]
  HK --> M[Mixed-precision VLA<br/>29.2% VRAM, 98.9% perf, 1.49x]
  LQ --> M
Loading

Problem

LLM-derived quantization (e.g., SmoothQuant, AWQ, OmniQuant) optimizes for data fidelity β€” minimizing weight/activation reconstruction error β€” but ignores that in a VLA, minor action deviations compound into catastrophic task failure. Uniform-bit quantization is therefore mismatched to embodied control. The paper's analysis shows significant intra-layer channel heterogeneity: individual channels contribute very unequally to the final action output, so treating them identically is wasteful.

Method

Channel-wise mixed-precision bit allocation guided by action-space sensitivity:

  • For each weight channel, the impact on the final action output is estimated via a first-order Taylor / Jacobian approximation, yielding a per-channel importance score in action space (not weight-reconstruction space).
  • A greedy demotion algorithm starts every channel at full precision and iteratively lowers the least-sensitive channels through the bit-width set {0, 2, 4, 8, 16} until the memory budget is met.
  • 0-bit = pruning, so quantization and structured pruning are unified in one framework. Activations are kept at a uniform bit-width (e.g., 8-bit) for hardware efficiency.

Results

On LIBERO (4 task suites), base model OpenVLA-OFT, W4A4 setting:

  • VRAM drops to 29.2% of the original (4.5 GB vs 15.4 GB β†’ 70.8% reduction) while retaining 98.9% of original performance, with a 1.49Γ— speedup. W8A8 gives a βˆ’0.7% drop at 7.2 GB / 1.36Γ—.
  • Strongly beats LLM-derived baselines at equal budget: at W4A4, QVLA βˆ’0.5% vs SmoothQuant βˆ’13.3%; at W4A16, QVLA βˆ’0.4% vs AWQ βˆ’4.5%; vs OmniQuant βˆ’3.2%.
  • Ablations: channel-wise > layer-wise (76.5% vs 74.8% at INT4); enabling 0-bit pruning matches/exceeds no-pruning while lowering VRAM (76.8% @ 7.0 GB vs uniform-8bit 74.6%).
  • Validated beyond LIBERO: UniVLA-7B (W4A16) 95.1% vs AWQ 92.6%; real-robot Ο€β‚€ retains its 63.3% success rate (W8A16, 1.28Γ— speedup).

Significance

Reframes VLA compression around action-centric rather than data-fidelity error, the first systematic quantization study tailored to VLAs. Makes larger VLAs deployable on commodity GPUs / on-device. The per-channel action-sensitivity map is itself a reusable artifact for future efficient-VLA work.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally