Skip to content

CoRL 2026 CHIP

hwoo.han edited this page Sep 28, 2026 · 1 revision

CoRL 2026 β€” CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation

Venue: CoRL 2026 (Austin, TX, Nov 9–12). Paper: arXiv 2512.14689. Representative of: learned compliance for humanoid manipulation β€” control not just where to move but how stiff/compliant the arms should be. Companions: Humanoid VLA Β· CoRL 2026 survey.

CHIP training-and-deployment overview: hindsight perturbation during RL yields a controllable-stiffness tracking policy on the Unitree G1 (figure from the authors, arXiv 2512.14689, Β© the authors)

1. Problem

Humanoid robots have become impressively agile at locomotion, but they remain weak at forceful manipulation β€” moving heavy objects, wiping a surface, pushing a cart, opening a door. RL-based motion-tracking controllers optimize for precise tracking of a reference motion, which makes the arms effectively stiff: they resist external contact instead of yielding to it. Classical impedance/compliance control gives that knob but integrates poorly with learned agile tracking. The goal is a controller that lets you dial end-effector stiffness up or down while still tracking dynamic reference motions well.

2. Method

CHIP is a plug-and-play module for keypoint-based motion-tracking policies that adds controllable end-effector stiffness β€” needing neither data augmentation nor extra reward tuning. The trick is hindsight perturbation: during training a random force f is applied at the end effector, and instead of editing the reference trajectory, CHIP edits the observed tracking goal to a hindsight target g βˆ’ (1/k)Β·f, while the tracking reward stays anchored to the original reference g. The coefficient controls the emulated stiffness k, so at deployment a continuous compliance coefficient sets how much the arm yields under contact. Force is estimated implicitly from proprioceptive history, and the module works with both local and global 3-point tracking policies. Trained on a Unitree G1.

3. Results

Reported on the Unitree G1 (baselines: FALCON force-perturbation and a no-force standard tracker):

  • Agile tracking is preserved β€” global position tracking error stays around 0.08 m.
  • End-effector displacement scales roughly linearly with the compliance coefficient, i.e. stiffness is genuinely controllable.
  • On multi-robot collaborative grasping/transport, CHIP reaches ~80% success vs. much lower baseline rates.
  • Downstream demos: VR teleoperation with on-the-fly compliance, cart pushing, door opening, and a VLA trained for autonomous wiping (60–80% success).

4. Why it matters

Compliance has been the missing knob for humanoid manipulation: agile trackers are stiff by construction, so contact-rich tasks either fail or fight the environment. CHIP shows you can graft a controllable-stiffness capability onto an existing tracking policy cheaply β€” no trajectory relabeling, no reward surgery β€” which makes it an attractive building block for humanoid VLAs and teleoperation.

Limitations (reviewer): stiffness is emulated via the observed-goal offset rather than measured force control, so behavior under large or fast contact forces depends on the quality of the implicit force estimate; results are single-embodiment (G1); reported manipulation success rates leave substantial headroom.

5. Links

← Back to CoRL 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally