-
Notifications
You must be signed in to change notification settings - Fork 0
Review Demystifying Action Space
Venue: ICML 2026 Β· arXiv: 2602.23408 Affiliations: Tsinghua University (IAIR / Wuxi Research Institute of Applied Technologies) Scale: 13,000+ real-world rollouts Β· 500+ trained models Β· 14 tasks (4 real + 10 sim) Β· 3 platforms One-pager: ICML-2026-Demystifying-Action-Space Β· See also VLA Architectures
This is the large-scale empirical study that asks: for an imitation-learning manipulation policy, should the action be in joint space or end-effector (task) space, and should it be absolute or delta? The answer, from 13k+ real rollouts, is not "one wins" β it is a decision rule:
- Temporal axis (absolute vs delta): delta wins, almost always β but only if implemented chunk-wise, not step-wise.
- Spatial axis (joint vs EEF): complementary. Joint space β control stability and gets better with scale; task/EEF space β generalization (cross-embodiment, transfer).
- Practical rule: single platform, maximize performance β joint + chunk-wise delta. Cross-embodiment / transfer β task-space (EEF).

Modern VLA / IL work pours effort into scaling data and model size, but the action space β the very thing the policy regresses to β is still chosen by "ad-hoc heuristics or legacy designs" (RT-1 used discretized EEF; Ο uses joint deltas; ACT uses joint absolute; etc.). The paper argues the action space "fundamentally shapes the optimization landscape," and sets out to measure it under controlled conditions instead of folklore.
The design space is factored into two orthogonal axes plus an implementation detail:

- Actuator space (torque / current) β the low-level controller layer.
- Joint space (configuration, ββΏ) β joint angles; reached from EEF via inverse kinematics (risk: singularities).
- Task space (end-effector SE(3) pose) β the gripper pose; reached from joints via forward kinematics.
- Absolute (0th order) β predict the global target position.
- Delta (1st order) β predict the relative displacement from a reference (needs integration to execute).
- Higher-order (accel / force) β requires accurate inertial modeling and high-frequency feedback (out of scope for standard IL).
When a policy predicts a chunk of future actions, a delta can be referenced two ways:
- Step-wise delta β each step relative to the previous predicted step (errors integrate along the chunk).
- Chunk-wise delta β every step relative to the single anchor frame at the chunk's start (errors stay independent).
- Policies: ACT (regression backbone) and DP (flow-matching / diffusion policy) β so conclusions aren't tied to one model class.
- Platforms: single-arm AgileX, bimanual AgileX, and RoboTwin 2.0 simulation.
- Tasks/data: 14 tasks (4 real + 10 sim); 250 demos/task (real), 50/task (sim); trained to 600 epochs (and 900/1200 for scaling).
- Scale knobs: data {100, 250, 500} trajectories Γ compute {600, 900, 1200} epochs; single- and multi-task; transfer from Ο0; cross-embodiment.
Chunk-wise delta consistently and significantly beats step-wise delta β the gap reaches upwards of 10% on average. The paper gives the mechanism as a stability theorem:
Proposition 4.1 (noise amplification). For a predicted chunk of length k with bounded per-step noise, the decoded execution error scales as O(k) for step-wise delta (the transform is the lower-triangular all-ones matrix Lβ, so errors accumulate along the horizon), but stays O(1) for chunk-wise delta and absolute actions (the transform is the identity Iβ, so errors propagate independently).
So step-wise delta structurally amplifies prediction noise as the horizon grows; chunk-wise does not. Takeaway: if you use delta actions, always anchor them chunk-wise.
Horizon coupling: absolute control prefers a longer execution horizon; delta control peaks at a shorter horizon (relative reps are sensitive to execution drift). They train at k=60 (2 s @ 30 Hz) and grid-search the execution window 15β60 β i.e., the chunk horizon is not a constant, it must be tuned to the temporal abstraction.

Temporal (delta vs absolute): with the best implementation of each, delta still consistently and significantly outperforms absolute across all platforms, tasks, and both ACT and DP. Why: (1) mapping high-dim vision β global coordinates has low local coherence, whereas predicting immediate displacement is a more tractable inductive bias; (2) absolute reps need long horizons that are hard to train.
Spatial (joint vs EEF) β the headline comparison: the two are complementary, not ranked:
- Joint space β the more robust / stable spatial representation in most single-platform cases.
- Task space (EEF) β competitive in low-data/limited-compute regimes and better for generalization.
- Delta stays superior as data and compute scale (sometimes by only a marginal margin).
- Joint-space superiority becomes more pronounced as epochs and data increase (especially for regression policies) β it "benefits disproportionately from stronger modeling and extensive training," apparently because it better captures the underlying kinematic manifold.
- Task-space (EEF) is competitive when data/compute is limited, and β critically β EEF becomes the superior choice under transfer learning (from Ο0) and cross-embodiment, where a body-agnostic SE(3) pose transfers across different kinematics while joint vectors do not.

| Joint space (ββΏ) | Task / EEF space (SE(3)) | |
|---|---|---|
| Strength | Control stability; captures the kinematic manifold | Generalization / transfer |
| Scaling | Improves with more data/compute (esp. regression) | Strong in low-data / limited-compute |
| Transfer & cross-embodiment | Weaker (joint vectors are body-specific) | Superior (SE(3) pose is body-agnostic) |
| Risks | needs IK; singularities | FK is well-defined |
| Best when | one hardware platform, maximize performance | many bodies, transfer, generalization |
And on the temporal axis, delta > absolute in essentially all standard settings β provided it is chunk-wise.
- The chunking execution horizon k is not a constant β adapt it to the temporal abstraction (longer for absolute, shorter for delta).
-
Single-platform, resource-rich, maximize performance β
joint space + chunk-wise deltais the most robust combination. -
Generalized setting (cross-embodiment / transfer learning) β
task space (EEF)is the superior spatial abstraction.
Most VLA papers inherit an action space without justifying it; this study turns that choice into a measured, reproducible decision and supplies a mechanistic reason (the O(k) vs O(1) noise-propagation result) for the single most common silent bug β step-wise delta chunking. It pairs naturally with VLA Architectures (the action-decoder axis) and the Ο-series choice of chunk-wise joint deltas, and it explains why cross-embodiment frameworks (Cross-Embodiment) gravitate to EEF/SE(3) action spaces. As an ICML-style "demystifying" study it is evidence, not a new model β but it is the reference to cite when defending an action-space choice.
No public code repository was found at review time; analysis is grounded in the arXiv paper (v1).
- arXiv: 2602.23408
- ICML 2026: https://icml.cc/virtual/2026/poster/62805
- One-page summary: ICML-2026-Demystifying-Action-Space
- Figures 1β5 reproduced from the paper (Tsinghua IAIR, 2026) for scholarly review; Β© the authors.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)