-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 RoboMD
Venue: ICLR 2026 Authors: Som Sagar Β· Jiafei Duan Β· Sreevishakh Vasudevan Β· Yifan Zhou Β· Heni Ben Amor Β· Dieter Fox Β· Ransalu Senanayake (Arizona State University Β· University of Washington Β· NVIDIA) arXiv: 2412.02818 Category: Robustness / security for VLA Trend tag: Vulnerability discovery Β· deep RL over VL embeddings Β· semantic potential fields
flowchart LR
subgraph EMB[Continuous vision-language embedding]
S[Success regions]
Fv[Vulnerable / failure regions]
end
DATA[Limited success-failure data] --> EMB
EMB -->|treat as potential field| PF[Semantic potential field]
PF --> RL[Deep RL vulnerability-prediction policy]
RL -->|attracted to| Fv
RL -->|repelled from| S
RL --> VR[Virtual runs in simulation]
VR --> OUT[Uncovered unique vulnerabilities<br/>up to +23% vs VL baselines]
Robot manipulation policies are highly vulnerable to external variations in the real world, but diagnosing these vulnerabilities is hard for two reasons: (i) the relevant variations to test against are often unknown a priori, and (ii) direct real-world testing is costly and unsafe. Heuristic testing tends to miss subtle failure modes.
RoboMD learns a separate deep reinforcement learning policy for vulnerability prediction, run virtually rather than on hardware:
- Build a continuous vision-language embedding trained from limited success-failure data β a space rich in semantic and visual variations.
- Treat that embedding as a potential field: the RL policy is attracted toward vulnerable (failure) regions and repelled from success regions.
- Explore the field via virtual runs, surfacing variations that break the target manipulation policy without risking the physical robot.
Across simulation benchmarks and a physical robot arm, RoboMD uncovers up to 23% more unique vulnerabilities than state-of-the-art vision-language baselines, revealing subtle failure modes overlooked by heuristic testing. The discovered vulnerabilities can guide targeted, data-efficient fine-tuning to improve manipulation robustness.
RoboMD turns vulnerability discovery into a search problem over a semantic embedding, casting safe, simulation-based red-teaming of robot policies as RL on a potential field. This complements perturbation-robustness training: RoboMD finds which semantic variations matter, while methods like RobustVLA harden policies against given perturbation sets.
- arXiv: https://arxiv.org/abs/2412.02818
- OpenReview: https://openreview.net/forum?id=Gsrw1vxq1G
- Code: https://github.com/somsagar07/RoboMD
- RobustVLA β hardening against perturbations once found
- Survey: VLA & Manipulation
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)