Skip to content

ICLR 2026 Constrained Demonstrators

Heungwoo edited this page Jun 1, 2026 · 1 revision

Learning from Constrained Demonstrators β€” when a robot can beat its (constrained) teacher

Venue: ICLR 2026 Β· Authors: Xinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz, Erdem BΔ±yΔ±k (USC) Β· arXiv 2510.09096 Β· Category: RL for manipulation Β· Trend tag: imitation beyond the expert / suboptimal demonstrations

Approach diagram

flowchart LR
  E[Constrained expert<br/>e.g. joystick = 2D plane] --> D[Suboptimal demonstrations]
  D --> RW[Infer state-only reward<br/>measuring task progress]
  RW --> SL[Self-label reward for unknown states<br/>via temporal interpolation]
  SL --> EX[Agent explores shorter,<br/>more efficient trajectories]
  EX --> P[Policy better than the demonstrated one]
Loading

Problem

Demonstration interfaces β€” kinesthetic teaching, joystick control, sim-to-real β€” often constrain the expert from showing optimal behavior because of indirect control, setup restrictions, and hardware-safety limits. A joystick may move an arm only in a 2D plane even though the robot acts in a higher-dimensional space, so collected demonstrations are suboptimal. Key question: can a robot learn a better policy than the one its constrained expert demonstrated?

Method

Rather than directly imitating expert actions, the agent is allowed to go beyond imitation and explore shorter, more efficient trajectories:

  • Use demonstrations to infer a state-only reward signal that measures task progress.
  • Self-label the reward for unknown / unvisited states using temporal interpolation.

Because the reward is state-only (not action-imitating), the policy is free to find more efficient paths than the constrained demonstrator could execute.

Results

  • Outperforms common imitation learning in both sample efficiency and task completion time.
  • On a real WidowX robotic arm, completes the task in 12 seconds β€” 10x faster than behavioral cloning.

Significance

Reframes learning-from-demonstration for the regime where the robot is more capable than the demonstration interface: instead of capping performance at the (constrained) expert, it uses demonstrations only to define task progress and lets RL-style exploration exceed them. Practically relevant wherever cheap but limited teleoperation interfaces are the data source.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally