-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 TAIL Safe
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 3 Β· paper #207 Authors: Riad Ahmed, Momotaz Begum arXiv: 2605.01195 Β· program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1. Overview of TAIL-Safe. Top-left: a Gaussian Splatting pipeline builds a digital twin (~20 min: 5 min capture + 15 min reconstruction), aligned to the robot via Umeyama's algorithm. Top-right: the simulator generates safe and unsafe trajectories under perturbations to train WeightNet (score fusion) and Q-ValueNet (success prediction). Middle/bottom: at deployment TAIL-Safe monitors Q(s,a) in real time, staying inactive while Q > 0 and steering the system back to safety via gradient-based recovery as Q approaches zero.
Imitation-learning policies (flow-matching, diffusion) can fail even within their training distribution due to sensitivity to initial conditions and compounding drift, making field deployment unsafe. Safe deployment requires knowing, for a trained policy, the set of states from which it is guaranteed to complete the learned task.
TAIL-Safe learns a Lipschitz-continuous Q-value function mapping stateβaction pairs to a safety score built from three task-agnostic criteria β visibility, recognizability, and graspability. The zero-superlevel set of Q defines a Control Invariant Set; when the nominal policy proposes an action outside it, Nagumo's theorem is used to compute a recovery action by gradient ascent on Q, steering back to safety. Training data is gathered from a photorealistic Gaussian-Splatting digital twin (~20 min to build) that generates ~500 rollouts per task (~40% failures) without risking hardware, feeding a WeightNet (score fusion) and Q-ValueNet (success prediction). The monitor runs at 20 Hz, with recovery taking 3β5 iterations (mean 2.3).
On a Franka Emika robot across two tabletop tasks, flow-matching policies without safety succeed only ~20β25% of the time under run-time perturbations, but reach 100% success when guided by TAIL-Safe. Compared to baselines, TAIL-Safe's Q(s,a) attains AUROC 0.999 with 100% recovery success at 2.8 ms per step, versus a Learned CBF (AUROC 0.987 but only 6.9% recovery) and an ensemble (0% recovery, whose mean recovery is actively harmful, ΞQ = β0.86). A WeightNet score fusion is shown to outperform equal weighting, and recovery achieves 100% success with 99.3% state-level accuracy.
Provides a task-agnostic run-time safety watchdog with formal invariance guarantees for otherwise-brittle IL policies, addressing reliability concerns central to Review-Realtime-Execution and safe deployment of learned manipulation.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)