-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 Chunking Augmentation
Venue: ICLR 2026 Β· Authors: Thomas T. Zhang, Daniel Pfrommer, Chaoyi Pan, Nikolai Matni, Max Simchowitz Β· Paper: arXiv 2507.09061 β Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control (Jul 2025) Β· Category: Theory / empirical study Β· Trend tag: Control-theoretic foundations of imitation learning
Note: the ICLR index lists this under the working title "Action Chunking and Data Augmentation Yield Exponential Improvements..."; the confirmed arXiv title uses "Exploratory Data Collection" and "Continuous Control."
flowchart LR
BC[Behavior cloning<br/>continuous control] --> Comp[Error compounds<br/>EXPONENTIALLY with horizon H]
Comp --> I1[Intervention 1:<br/>Action chunking<br/>open-loop action sequences]
Comp --> I2[Intervention 2:<br/>Exploratory data collection<br/>augment expert demos]
I1 --> Stab{Control-theoretic<br/>stability}
I2 --> Stab
Stab --> Poly[Compounding error<br/>circumvented in different regimes]
Poly --> Bounds[Tighter statistical<br/>guarantees on IL error]
In continuous control, behavior cloning can suffer error that compounds exponentially with the task horizon β a worst-case that information-theoretic analyses capture but do not explain mechanistically. Two interventions are widely used in practice (action chunking; augmenting/exploring around expert demonstrations) and empirically help, but lacked a rigorous account of why and when.
A theoretical analysis with a control-theoretic lens:
- Action chunking = predicting sequences of actions executed open-loop. The paper shows this circumvents exponential compounding error in certain regimes.
- Exploratory data collection = augmenting expert demonstrations with exploratory data. This similarly avoids exponential blow-up, but in a different operating regime.
- The unifying mechanism identified is control-theoretic stability of the underlying dynamics; stability is what prevents the error cascade. This yields fine-grained insight into how compounding error arises and tighter statistical guarantees on imitation-learning error than information-theoretic bounds alone.
Theoretical predictions are validated on popular robot-learning benchmarks, matching the regimes in which each intervention is predicted to help. (The work is primarily analytical; precise bound constants and benchmark scores are omitted here pending the full tables.)
Provides a principled, stability-based explanation for two of the most impactful empirical tricks in modern robot imitation learning (chunking as used in ACT/Ο0-style policies; data augmentation/DAgger-like exploration). Reframes the compounding-error story from worst-case pessimism to a regime-dependent picture governed by control-theoretic stability β guidance for when chunking vs. exploratory data is the right lever.
- arXiv: https://arxiv.org/abs/2507.09061
- OpenReview: https://openreview.net/forum?id=jiWXDvw1Lf
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)