Skip to content

ICLR 2026 Chunking Augmentation

Heungwoo edited this page Jun 1, 2026 · 1 revision

Action Chunking & Exploratory Data β€” exponential improvements in BC

Venue: ICLR 2026 Β· Authors: Thomas T. Zhang, Daniel Pfrommer, Chaoyi Pan, Nikolai Matni, Max Simchowitz Β· Paper: arXiv 2507.09061 β€” Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control (Jul 2025) Β· Category: Theory / empirical study Β· Trend tag: Control-theoretic foundations of imitation learning

Note: the ICLR index lists this under the working title "Action Chunking and Data Augmentation Yield Exponential Improvements..."; the confirmed arXiv title uses "Exploratory Data Collection" and "Continuous Control."

Approach diagram

flowchart LR
  BC[Behavior cloning<br/>continuous control] --> Comp[Error compounds<br/>EXPONENTIALLY with horizon H]
  Comp --> I1[Intervention 1:<br/>Action chunking<br/>open-loop action sequences]
  Comp --> I2[Intervention 2:<br/>Exploratory data collection<br/>augment expert demos]
  I1 --> Stab{Control-theoretic<br/>stability}
  I2 --> Stab
  Stab --> Poly[Compounding error<br/>circumvented in different regimes]
  Poly --> Bounds[Tighter statistical<br/>guarantees on IL error]
Loading

Problem

In continuous control, behavior cloning can suffer error that compounds exponentially with the task horizon β€” a worst-case that information-theoretic analyses capture but do not explain mechanistically. Two interventions are widely used in practice (action chunking; augmenting/exploring around expert demonstrations) and empirically help, but lacked a rigorous account of why and when.

Method

A theoretical analysis with a control-theoretic lens:

  • Action chunking = predicting sequences of actions executed open-loop. The paper shows this circumvents exponential compounding error in certain regimes.
  • Exploratory data collection = augmenting expert demonstrations with exploratory data. This similarly avoids exponential blow-up, but in a different operating regime.
  • The unifying mechanism identified is control-theoretic stability of the underlying dynamics; stability is what prevents the error cascade. This yields fine-grained insight into how compounding error arises and tighter statistical guarantees on imitation-learning error than information-theoretic bounds alone.

Results

Theoretical predictions are validated on popular robot-learning benchmarks, matching the regimes in which each intervention is predicted to help. (The work is primarily analytical; precise bound constants and benchmark scores are omitted here pending the full tables.)

Significance

Provides a principled, stability-based explanation for two of the most impactful empirical tricks in modern robot imitation learning (chunking as used in ACT/Ο€0-style policies; data augmentation/DAgger-like exploration). Reframes the compounding-error story from worst-case pessimism to a regime-dependent picture governed by control-theoretic stability β€” guidance for when chunking vs. exploratory data is the right lever.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally