-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 EigenSafe
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 2 Β· paper #146 Authors: Inkyu Jang, Jonghae Park, Sihyun Cho, Chams Eddine Mballo, Claire Tomlin, H. Jin Kim arXiv: 2509.17750 Β· program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 7 traces the learned eigenfunction ΟΟ during the almond-pouring task. Left: the naive behavior-cloned policy β ΟΟ drops sharply (A) before the actual spillage (B), showing the critic predicts failure ahead of time. Right: the safety-filtered policy (n=50 candidates, k=10) keeps ΟΟ high, achieving a stable grasp (D) and safe pouring (E, F).
Real robotic systems are stochastic (sensor noise, disturbances), so binary safe/unsafe classification is unrealistic; safety is better measured as the probability of avoiding failure over a horizon. HJ-reachability and CBF-based critics learn surrogates (value functions, barrier scores) that are not calibrated to the actual closed-loop safety probability, yielding inconsistent or overly conservative behavior.
EigenSafe (SNU / UC Berkeley) observes that the dynamic-programming recursion for the safety probability is a repeated application of a linear, positive, non-expansive operator (a Koopman-operator restriction). By PerronβFrobenius/KreinβRutman theory its dominant eigenpair (Ξ³Ο, ΟΟ) captures long-horizon safety: Z(t,x) β cΒ·ΟΟ(x)Β·Ξ³Οα΅, so the eigenvalue Ξ³Ο β [0,1] is a global decay-rate metric for the closed-loop policy, and the eigenfunction ΟΟ(x,u) is a calibrated safety Q function over state-action pairs. The eigenpair is learned model-free from transition tuples with a power-iteration-inspired loss (eigen-equation residual plus a sup-norm normalization term), avoiding artificial discount factors. Two applications: (1) safe RL β SAC with the global constraint Ξ³Ο β₯ Ξ³0 converted to an equivalent local constraint E[ΟΟ(xβ²,uβ²)] β₯ Ξ³0Β·ΟΟ(x,u) and enforced with state-action-dependent Lagrange multipliers; (2) test-time safety filtering for imitation learning β sample n candidate actions from a flow-matching BC policy and execute the one with the k-th largest ΟΟ (kβn/5 avoids OOD actions).
In four Gym safe-RL environments (CheetahLow, HopperHigh, LunarLanderHard, AntBall), EigenSafe consistently sits in the upper-right of the reward-vs-time-to-failure plane compared to vanilla SAC, RESPO, and EFPPO. In the real-world food-preparation task (UR3 + Robotiq 2F-85, two RGB cameras with DINOv2 features forming a 776-d state, GELLO-teleoperated data: 150 successful + 203 failed demos), safety filtering with n=50, k=10 gives the best success/safety rates over 20 trials β exceeding both the naive BC policy and a BC policy trained only on successful demonstrations β and ΟΟ visibly drops before failures occur, demonstrating predictive capability.
Replaces optimal-control surrogates with a spectrally grounded critic that is provably tied to the actual safety probability β a principled, calibrated alternative for both safe RL and post-hoc filtering of generative IL policies. The candidate-ranking filter is directly applicable to any stochastic policy (diffusion/flow VLAs included). Related wiki threads: Review-LBM-Cotraining Β· RL.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)