Skip to content

RSS 2026 EigenSafe

hwoo.han edited this page Aug 9, 2026 · 2 revisions

EigenSafe: A Spectral Framework for Learning-Based Probabilistic Safety Assessment

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Imitation learning 2 Β· paper #146 Authors: Inkyu Jang, Jonghae Park, Sihyun Cho, Chams Eddine Mballo, Claire Tomlin, H. Jin Kim arXiv: 2509.17750 Β· program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Safety-filtered imitation learning on the UR3 pouring task (Figure 7 of arXiv 2509.17750, Β© the authors)

Figure 7 traces the learned eigenfunction ΟˆΟ€ during the almond-pouring task. Left: the naive behavior-cloned policy β€” ΟˆΟ€ drops sharply (A) before the actual spillage (B), showing the critic predicts failure ahead of time. Right: the safety-filtered policy (n=50 candidates, k=10) keeps ΟˆΟ€ high, achieving a stable grasp (D) and safe pouring (E, F).

Problem

Real robotic systems are stochastic (sensor noise, disturbances), so binary safe/unsafe classification is unrealistic; safety is better measured as the probability of avoiding failure over a horizon. HJ-reachability and CBF-based critics learn surrogates (value functions, barrier scores) that are not calibrated to the actual closed-loop safety probability, yielding inconsistent or overly conservative behavior.

Method

EigenSafe (SNU / UC Berkeley) observes that the dynamic-programming recursion for the safety probability is a repeated application of a linear, positive, non-expansive operator (a Koopman-operator restriction). By Perron–Frobenius/Krein–Rutman theory its dominant eigenpair (Ξ³Ο€, ΟˆΟ€) captures long-horizon safety: Z(t,x) β‰ˆ c·φπ(x)Β·Ξ³Ο€α΅—, so the eigenvalue Ξ³Ο€ ∈ [0,1] is a global decay-rate metric for the closed-loop policy, and the eigenfunction ΟˆΟ€(x,u) is a calibrated safety Q function over state-action pairs. The eigenpair is learned model-free from transition tuples with a power-iteration-inspired loss (eigen-equation residual plus a sup-norm normalization term), avoiding artificial discount factors. Two applications: (1) safe RL β€” SAC with the global constraint Ξ³Ο€ β‰₯ Ξ³0 converted to an equivalent local constraint E[ΟˆΟ€(xβ€²,uβ€²)] β‰₯ Ξ³0Β·ΟˆΟ€(x,u) and enforced with state-action-dependent Lagrange multipliers; (2) test-time safety filtering for imitation learning β€” sample n candidate actions from a flow-matching BC policy and execute the one with the k-th largest ΟˆΟ€ (kβ‰ˆn/5 avoids OOD actions).

Results

In four Gym safe-RL environments (CheetahLow, HopperHigh, LunarLanderHard, AntBall), EigenSafe consistently sits in the upper-right of the reward-vs-time-to-failure plane compared to vanilla SAC, RESPO, and EFPPO. In the real-world food-preparation task (UR3 + Robotiq 2F-85, two RGB cameras with DINOv2 features forming a 776-d state, GELLO-teleoperated data: 150 successful + 203 failed demos), safety filtering with n=50, k=10 gives the best success/safety rates over 20 trials β€” exceeding both the naive BC policy and a BC policy trained only on successful demonstrations β€” and ΟˆΟ€ visibly drops before failures occur, demonstrating predictive capability.

Significance

Replaces optimal-control surrogates with a spectrally grounded critic that is provably tied to the actual safety probability β€” a principled, calibrated alternative for both safe RL and post-hoc filtering of generative IL policies. The candidate-ranking filter is directly applicable to any stochastic policy (diffusion/flow VLAs included). Related wiki threads: Review-LBM-Cotraining Β· RL.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally