-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 RC NF
RC-NF β Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
Venue: CVPR 2026 Category: Anomaly Detection / Robustness Trend tag: Trend 6 Affiliations: Fudan ITEA + SMU Authors: Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang
flowchart LR
OBS["observation"] --> SAM2["SAM2 mask<br/>β point sets"]
SAM2 --> P1["dynamic-shape branch"]
SAM2 --> P2["positional-residual branch"]
P1 --> NF["robot-conditioned<br/>coupling flow (RCPQNet)"]
P2 --> NF
ROBOT["robot state<br/>(joints, gripper, pose)"] --> NF
TASK["task embedding (FiLM)"] --> NF
NF --> SCORE["anomaly score (neg log-density)"]
SCORE --> DEC["anomaly? <100 ms"]
Anomaly detection on top of VLA policies has been slow and policy-specific. A general, real-time detector that can run on top of any VLA (Ο0, OpenVLA, RDT-class) is missing.
A task-aware, robot-conditioned coupling normalizing flow (Glow-based, K=12 flow steps) with dual-branch point features:
- Use SAM2 to segment objects in the observation, then grid-sample masks into point sets. Two complementary branches encode them: a dynamic-shape branch that normalizes each frame to remove translation/scale, and a positional-residual branch that restores the absolute-position information lost in normalization.
- The core block is the Robot-Conditioned Point Query Network (RCPQNet), an affine coupling layer that decouples robot-state and object features while preserving their interaction: robot-state features (joint angles, gripper state, Cartesian pose) act as query tokens and object point features as memory tokens in cross-attention, with transformation parameters modulated by the task embedding via FiLM. Conditioning is on the robot's state, not robot identity β so anomaly scoring is grounded in whether robot and object motion stay consistent with the task.
- Trained unsupervised on positive (normal) samples only. Score new observations by their negative log-density under the flow; low density β anomaly, enabling state-level rollback or task-level replanning.
Evaluated on LIBERO-Anomaly-10, a new simulation benchmark introduced by the paper covering three manipulation-specific anomaly types: Gripper Open (gripper stays open while grasping), Gripper Slippage (zero friction β object slips), and Spatial Misalignment (robot moves toward the wrong compartment).
- SOTA: 0.9309 average AUC / 0.9494 average AP, vs. FailDetect at 0.7181 AUC / 0.7700 AP. Baselines also include VLM-based detectors (GPT-5, Gemini 2.5 Pro, Claude 4.5).
- Sub-100 ms detection latency on a consumer GPU (RTX 3090), running as a plug-and-play monitor atop existing Ο0 / OpenVLA-class policies.
The first general, real-time, plug-in anomaly detector for VLAs. Critical infrastructure for safe deployment, especially as long-horizon VLAs proliferate. Sister role to SafeVLA and Latent Policy Barrier β both prevent unsafe actions; RC-NF detects unsafe states.
- arXiv: 2603.11106
- Project:
heikaishuizz.github.io/RC-NF
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)