-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 HyperVLA
Venue: ICLR 2026 Category: VLA Architecture β Efficiency Trend tag: Trend 1 / efficiency
flowchart LR
T[Language instruction + initial image oβ<br/>frozen T5 + frozen DINOv2 class token] --> H[Transformer context encoder<br/>HYPERNETWORK 216M<br/>runs ONCE per task]
H --> W[Generated weights for<br/>compact ViT policy ~0.1M activated]
W --> P[Compact policy<br/>runs every control step]
Obs[Observation] --> P
P --> A[Action]
Billion-parameter VLAs deliver strong generalization but demand significant inference compute, limiting deployment to high-end hardware. Distillation and quantization help but trade accuracy.
Use a hypernetwork (HN): a Transformer context encoder reads the task once β conditioned on the frozen T5 language-instruction embedding plus the frozen DINOv2 class-token embedding of the episode's initial image oβ β and generates the weights of a small task-specific policy. The HN is not a large VLM; it is a high-capacity Transformer with linear output heads (216M HN params + 86M shared params at training time), and it generates only ~0.1M parameters that are actually activated per control step. The base policy is a Vision Transformer (ViT) with a DINOv2 image encoder, a small Transformer policy head, and a linear action head. At deployment only the compact policy runs per control step; the HN sits idle. This is structurally different from distillation β the compact policy is generated per task rather than learned once. Two key design features: an HN normalization technique and a linear (non-diffusion) action head.
Versus OpenVLA, HyperVLA reduces test-time activated parameters by ~90Γ and achieves a ~120Γ inference speedup, while matching or exceeding success rates:
- SIMPLER (zero-shot): Google Robot pick 58Β±3% (OpenVLA 10%), Google Robot move 73Β±1% (OpenVLA 72%), WidowX avg 40Β±5% (OpenVLA 36%).
- LIBERO (few-shot adaptation): avg 89% vs Octo 75% / OpenVLA 77% (Spatial 95, Object 94, Goal 92, Long 74).
Ablations (Table 4): removing the vision backbone drops avg from 63% to 31%; removing HN normalization degrades OOD tasks to 31%; replacing the linear action head with diffusion falls to 53% avg.
A third path (alongside distillation and quantization) for deploying large VLAs. Suggests an interesting research direction: if the hypernetwork can condition on more than the task description (specific robot, environment, user preferences), it could produce true per-deployment specialized policies without fine-tuning.
Authors flag as future work: real-robot evaluation (results are simulation-only on SIMPLER/LIBERO), scaling up HN model size, and training on larger/more recent robotic datasets.
- arXiv:2510.04898 (Xiong, Li, Wang, Jackson, Foerster, Whiteson β University of Oxford)
- OpenReview (ICLR 2026)
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)