-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 SP VLA
Venue: ICLR 2026 Category: VLA Architecture β Efficiency Trend tag: Efficiency / real-time deployment Affiliation: Tsinghua University Β· CUHK Β· UIUC Β· Beihang University
flowchart LR
Hist[Action buffer S_A<br/>last n=6 actions] --> SCH[Action-aware Scheduler<br/>velocity check + intuitive ratio Ο=0.5]
SCH -- "Deliberative path<br/>low speed or low intuitive ratio" --> TP[Spatio-Semantic Token Pruning<br/>velocity-adaptive prune ratio]
TP --> ENC[ViT Encoder]
ENC --> LLM[7B LLM<br/>e.g. Llama 2 / OpenVLA / CogACT]
LLM --> AH[Action Head<br/>D-tokenizer or diffusion]
SCH -- "Intuitive path<br/>v in v_min..v_max and N_G/N_A above Ο" --> RIDGE["Lightweight Generator<br/>Ridge Regression on action buffer<br/>under 1K params"]
RIDGE --> CHK[Validity check]
AH --> ACT[Robot action]
CHK --> ACT
ACT -. feedback .-> Hist
VLA models are too compute-heavy for real-time deployment. OpenVLA is 7B+, RT-X is up to 55B; even on a 4090 inference runs at ~4 Hz. Existing acceleration work attacks single-step compute via quantization (QAIL), early-exit (DeeR-VLA), token caching (VLA-Cache), parallel decoding (PD-VLA, OpenVLA-OFT), or speculative decoding β but all treat one frame in isolation.
SP-VLA argues two redundancies are systematically ignored:
- Temporal redundancy in the sequential action stream (most steps don't need full reasoning β humans alternate "deliberate" vs "intuitive" motor control).
- Spatial redundancy in visual tokens (most patches are background β but VLA-targeted token pruning has unique constraints).
Behavioral analysis. Across 50 pick-and-place trials, the authors observe the manipulator follows a 4-phase velocity profile: target β grasp β move β place. Slow alignment, then high-speed translation, then careful action, etc. They argue VLA models have implicitly learned this kinematic pattern, so the action stream is naturally split into:
- Deliberative actions β slow, precise (grasping, turning, placement). Need the full 7B VLA.
- Intuitive actions β fast, ballistic (point-to-point translation). Can be approximated by a tiny model.
Scheduler logic. Let a^t_d = (a_x, a_y, a_z) be the per-step end-effector translational velocity. Define:
- An action a is intuitive if all components |a_i| > v_min (speed threshold).
- The lightweight model is allowed when (a) a_{t-1} β [v_min, v_max] and (b) the ratio of recent VLA-generated actions in the buffer N_G / N_A > Ο (default Ο = 0.5).
LWM = 1 if both hold, else 0. (Equation 1.) This drives small-step, high-frequency model switching β even within an "intuitive" segment, the VLA is invoked periodically to correct drift.
Lightweight generator: Ridge Regression on the action buffer S_A = {a_{t-n}, β¦, a_{t-1}} (n=6).
- X = [T, 1] β R^{nΓ2}, T = [0,β¦,nβ1]α΅; Y = action buffer; Ξ² β R^{2Γβ}.
- Minimize J(Ξ²) = βXΞ² β YβΒ² + Ξ»βΞ²βΒ² (Tikhonov / ridge).
- Closed-form: Ξ² = (Xα΅ X + Ξ»I)β»ΒΉ Xα΅ Y.
- Predict a_t = x_t Ξ²*, x_t = [t 1]α΅.
- Gripper state is NOT regressed (it's binary) β instead reuse the tβ1 value, leaving binary state transitions to the VLA. Predicted intuitive actions go through a validity check before execution.
Insight from controlled experiments (Fig. 2b): Random token pruning degrades but doesn't destroy task performance β there is spatial redundancy. But two surprising failure modes:
- Reordering tokens by semantic importance (no actual pruning) causes complete task failure β relative position of tokens carries spatial meaning the auto-regressive VLA depends on.
- Pruning purely by semantic attention scores removes object-contour background tokens and also fails β object contours are critical for spatial grounding.
So pruning must (a) preserve relative ordering and (b) explicitly retain edge / contour tokens.
Semantic-aware token importance. From last-encoder-layer attention: Q,K,V = X W_q,k,v Attn = Softmax(QKα΅/βd_k) V AccuAttn = Β½ (eα΅ β I_M) vec(Attn). Select T_se = {x_i | AccuAttn_i > t_ks}.
Spatial-aware token importance. Apply Canny edge detector to the input image: X_s = Canny(X). Then T_sp = f_E(X_s) is the ordered set of edge-region tokens.
Order-preserving union: T_select = U(T_se, T_sp) β the union, with original token positions retained.
Velocity-adaptive prune rate. Pruning is disabled for low-speed (deliberative) actions to protect precision. For higher speeds: T_r(v) = 1 if v < v_pmin, else 1 β (v β v_pmin) / (v_pmax β v_pmin).
This couples the spatial dimension (how aggressively to prune) to the temporal mode β fast intuitive actions can tolerate aggressive pruning; precise grasps cannot.
- Buffer size n = 6.
- Deliberation/intuition ratio threshold Ο = 0.5.
- Velocity thresholds: v_min = 0.2, v_max = 0.5 (paper settings); token-pruning velocity threshold v_pmin = 0.5. Sensitivity study (Table 7) varies these Β±25% and finds accuracy robust to speed but sensitive to n and Ο.
- Hardware: NVIDIA A100 GPUs for experiments/training; NVIDIA RTX 4090 (40GB) for frequency/latency measurements (paper Sec. on Frequency and Latency).
| Method | Goal | Object | Spatial | Long | Avg | Speedup | FLOPs % |
|---|---|---|---|---|---|---|---|
| OpenVLA | 75.40 | 86.20 | 83.80 | 53.00 | 74.60 | 1.00Γ | 100 |
| SparseVLM | 74.20 | 84.00 | 83.40 | 52.80 | 73.60 | 1.33Γ | 75.55 |
| FoPru + R | 59.80 | 81.20 | 71.60 | 26.20 | 59.70 | 1.31Γ | 77.20 |
| PruMerge + R | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.36Γ | 73.63 |
| FastVLM + R + S | 73.20 | 77.00 | 79.80 | 36.60 | 66.65 | 1.16Γ | 86.22 |
| VisionZip + R + S | 46.00 | 47.40 | 34.20 | 4.60 | 33.05 | 1.21Γ | 81.95 |
| Ours (Speed-priority) | 73.60 | 82.40 | 80.00 | 51.60 | 71.90 | 1.50Γ | 66.51 |
| Ours (Acc-priority) | 75.40 | 85.60 | 84.40 | 54.20 | 74.90 | 1.35Γ | 73.64 |
"+R" = preserve relative token positions. "+S" = add Canny edges. SP-VLA is the only method to deliver acceleration with β₯ baseline accuracy. PruMerge is interesting β it gets a 1.36Γ speedup but drops to 0% on every suite (the model collapses).
Visual Matching split:
| Method | PickCan | MoveNear | Drawer | DrawerApple | Avg | Speedup | FLOPs % |
|---|---|---|---|---|---|---|---|
| CogACT | 91.30 | 85.00 | 71.80 | 50.90 | 74.80 | 1.00Γ | 100 |
| Random Drop | 9.70 | 20.40 | 53.50 | 0.00 | 20.90 | 1.20Γ | 58.50 |
| FastV | 92.60 | 81.40 | 69.80 | 52.40 | 74.10 | 1.21Γ | 42.00 |
| VLA-Cache | 92.00 | 83.30 | 70.50 | 51.60 | 74.40 | 1.38Γ | 80.10 |
| EfficientVLA | 93.30 | 81.30 | 68.20 | 53.80 | 74.20 | 1.93Γ | 28.90 |
| Ours | 90.00 | 82.08 | 75.35 | 52.78 | 75.05 | 2.15Γ | 38.15 |
Visual Aggregation split:
| Method | PickCan | MoveNear | Drawer | DrawerApple | Avg | Speedup |
|---|---|---|---|---|---|---|
| CogACT | 89.60 | 80.80 | 28.30 | 46.60 | 61.30 | 1.00Γ |
| EfficientVLA | 93.20 | 75.80 | 26.90 | 49.20 | 61.20 | 1.91Γ |
| Ours | 86.18 | 77.33 | 55.29 | 41.80 | 65.16 | 2.09Γ |
Note +27 pp on Drawer in Visual-Aggregation while still 1.81Γ faster β interpreted as error-correction (the scheduler/lightweight generator smooths trajectories where CogACT alone falters).
WidowX split (bottom of Table 2; paper labels it "WindowX"):
| Method | PutSpoon | PutCarrot | StackBlock | PutEggplant | Avg | Speedup |
|---|---|---|---|---|---|---|
| CogACT | 71.70 | 50.80 | 15.00 | 67.50 | 51.30 | 1.00Γ |
| Ours | 70.83 | 54.17 | 29.17 | 75.00 | 57.29 | 2.41Γ |
Stack Block jumps from 15 β 29 with 2.54Γ acceleration.
| Setting | CogACT freq | SP-VLA freq | CogACT latency | SP-VLA latency |
|---|---|---|---|---|
| SimplerEnv Visual-Matching avg | 3.77 Hz | 8.06 Hz | 0.27 s | 0.13 s |
| SimplerEnv WindowX avg | 3.96 Hz | 8.69 Hz | 0.25 s | 0.12 s |
β 2.2Γ frequency improvement, ~50% latency reduction.
The ICLR camera-ready / arXiv v3 paper (2506.12723v3, Oct 2025) reports no real-robot experiment table β all benchmarks are simulation (LIBERO + SimplerEnv). The project's GitHub README separately states a real Franka Panda result of ~2.5Γ end-to-end inference acceleration with only a 1% success-rate drop, but provides no per-task breakdown, FLOPs, or latency numbers. Treat any detailed real-robot figures with caution until the source table is located.
Audit note: a previously listed "Real Franka Research 3 (Table 4)" table with per-task success rates (80/74/77 vs 78/74/76), FLOPs 35.55, 0.27/0.13 s latency, "150 trajectories per task," and "20 morning / 10 noon / 20 evening" lighting splits was not found in any reachable source and has been removed as a likely fabrication.
| Variant | Goal | Object | Spatial | Long | Avg | Speedup |
|---|---|---|---|---|---|---|
| Full SP-VLA | 75.40 | 85.60 | 84.40 | 54.20 | 74.90 | 1.35Γ |
| w/o Pruning | 74.40 | 84.20 | 84.00 | 53.30 | 73.98 | 1.27Γ |
| w/o Scheduling | 77.31 | 81.80 | 79.00 | 48.00 | 71.52 | 1.21Γ |
| w/o Canny (no edge tokens) | 33.60 | 39.00 | 22.00 | 1.10 | 23.93 | 1.35Γ |
Without Canny edge tokens the model effectively collapses (74.9 β 23.9), confirming the central insight that VLA spatial perception relies on object-contour tokens.
- Token pruning identifies ~22.7% redundancy on LIBERO-Spatial.
- Model scheduling identifies ~21% redundancy on LIBERO-Object.
- Intuitive-action proportion grows with task length: 18% on LIBERO-Spatial β 28% on LIBERO-Long β corresponding speedups 1.18Γ β 1.39Γ.
The temporal and spatial dimensions are not orthogonal β different tasks expose redundancy along different axes β so combining both is multiplicative, not additive.
The authors flag a single explicit limitation:
- Intuitive-action generation is preliminary. Currently they only "lightweight-ify" the VLA via Ridge Regression; they have not achieved a complete behavioral separation between deliberative and intuitive generation. Making the VLA more human-like by structurally separating these two modes (rather than scheduling between a big and tiny model) is flagged as the key future direction.
Other implicit caveats:
- Hyperparameter v_min, v_max are device-dependent. The 1/4-3/4 max-task-speed heuristic (v_min = 0.2, v_max = 0.5 in the paper's settings) is validated only in simulation; the sensitivity study (Table 7) shows accuracy is robust to Β±25% speed perturbation but sensitive to buffer size n and the intuitive-action proportion Ο.
- Lossless acceleration claim is benchmark-specific. "1.5Γ lossless on LIBERO" is the Speed-priority preset with 71.9 vs 74.6 baseline (2.7 pp drop) β not strictly lossless. Authors' "Acc" preset gets 74.9 (no drop) at 1.35Γ.
- No real-robot evaluation in the paper. v3 reports only LIBERO + SimplerEnv (simulation); the lone real Franka claim lives in the repo README without a results table.
- Lightweight generator works because intuitive segments are approximately linear; on highly nonlinear / contact-rich intuitive segments (rare in pick-and-place but common in dexterous manipulation), Ridge Regression may break down.
- vs token-pruning-only methods (Action-aware Dynamic Pruning, FastV, VLA-Cache, EfficientVLA, SparseVLM, FoPru, FastVLM, VisionZip, PruMerge): SP-VLA's spatio-semantic-with-edge pruning prevents the catastrophic failure the paper documents in pure semantic pruners (PruMerge β 0%, VisionZip β 33%). Combined with scheduling it gets 2.4Γ SimplerEnv speedup at +0.25 pp accuracy, which no single-axis pruner achieves.
- vs quantization (AutoQVLA, QAIL): orthogonal β could be stacked.
- vs parallel decoding (OpenVLA-OFT, PD-VLA, FASTER): these accelerate inside one VLA forward pass; SP-VLA accelerates between forward passes by skipping them. Composable.
- vs hierarchical dual-system VLAs (Ο0.5, Hi-Robot, Helix, OneTwoVLA): dual-system designs route reasoning to a separate model; SP-VLA borrows the System-1 / System-2 metaphor but applies it to action-typing rather than reasoning. The "lightweight generator" is essentially a System-1 motor primitive (ridge regression) for intuitive arm trajectories, while the full VLA acts as System-2 for grasp-class moments.
- vs early-exit (DeeR-VLA): DeeR exits earlier in the LLM stack on easy frames; SP-VLA skips the LLM entirely on intuitive frames, which is a stronger speedup ceiling.
- First systematic treatment of temporal Γ spatial redundancy in VLA inference β the action-type indicator + edge-aware pruning ablation reveals previously unappreciated structural facts about VLA computation.
- Action-aware Dynamic Pruning β token-level efficiency
- AutoQVLA β channel-aware quantization
- FASTER
- OneTwoVLA β adaptive reasoning, kindred conceptual structure
- Survey: VLA & Manipulation
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)