-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 CoCo InEKF
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Perception and Estimation Β· paper #178 Authors: Michael Baumgartner, David MΓΌller, Agon Serifi, Ruben Grandia, Espen Knoop, Markus Gross, Moritz BΓ€cher arXiv: 2605.15122 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Left: the bipedal robot with IMU, actuator measurements, and predefined contact candidate points on the feet. Right: the pipeline β a learned contact module maps proprioception to per-candidate contact velocity covariances, which feed a differentiable Invariant EKF that fuses IMU and leg odometry into the state estimate.
Proprioceptive state estimation for legged robots traditionally hinges on binary contact detection plus a stationarity assumption, which breaks under partial contact and directional slippage during highly dynamic motion; prior learned contact detectors (e.g., Lin et al.) need labeled contact data and still output binary states. Pure end-to-end estimators, meanwhile, underperform on accuracy and lose filter consistency.
CoCo-InEKF (ETH Zurich / Disney Research) makes the contact-aided Invariant EKF differentiable by permanently maintaining all contact candidates in the state, and replaces binary contacts with continuous contact velocity covariances predicted per candidate by a lightweight MLP contact module. Training is end-to-end via backpropagation through time (BPTT, horizon H=20, buffer L=128) with a simple body-frame velocity L2 state-error loss β no heuristic contact labels; physical interpretability is not enforced, letting the filter exploit non-physical constraints. An automated farthest-point-sampling procedure selects contact candidates. Experiments run on Lima, a custom 0.84 m, 16.2 kg, 20-DoF bipedal robot with a 600 Hz control loop on an onboard Intel i7.
On simulated dancing (RL motion-tracking policy over 81 retargeted Reallusion sequences), CoCo-InEKF achieves the lowest linear-velocity ATE (RMSE 0.046 m/s vs. 0.121β0.123 for the hybrid learned-contact baselines and 2.675 for heuristic-contact InEKF, which diverges). On contact-rich ground motions (N=10 candidates), it reaches RMSE 0.099 m/s, matching much larger end-to-end SET models (0.096β0.107) at a fraction of the cost β 335K vs. 4.8M parameters and 0.18 ms vs. 3.06 ms network inference β and scaling to 18 candidates improves RMSE to 0.069. Automated candidate selection matches or beats hand-picked placements, NEES analysis shows improved filter consistency, and the estimator supports dancing and full-body ground-interaction motions on the physical robot.
A hybrid estimator design point: keep the InEKF's structure, invariance, and consistency, but learn a richer-than-binary contact representation end-to-end through the filter itself. Useful context for whole-body humanoid control stacks discussed in Review-Humanoid-VLA.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)