-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 RealVLG R1
RealVLG-R1 β Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
Venue: CVPR 2026 Category: Affordance / Grounding Benchmark + Model Trend tag: Trend 1 Affiliations: Tongji University (School of Computer Science and Technology) Authors: Linfei Li, Lin Zhang, Ying Shen arXiv: 2603.14880v1 (16 Mar 2026)
flowchart LR
DATA["RealVLG-11B<br/>~165k images, ~1.3M annotations,<br/>~11B grasp examples"] --> RL["R1-style RFT<br/>(Qwen2.5-VL 3B/7B)"]
RL --> MODEL["RealVLG-R1 model"]
IMG["real-world image"] --> MODEL
LANG["language query"] --> MODEL
MODEL --> MASK["object mask"]
MODEL --> BBOX["bounding box"]
MODEL --> GRASP["grasp pose"]
MODEL --> CONTACT["contact point"]
Manipulation grounding tasks (masks, bboxes, grasps, contact points) have historically been siloed β separate datasets, separate models. There has been no large-scale real-world dataset that joints all four under a single language-grounding interface.
- Build RealVLG-11B β ~165 000 real-world images over ~800 object instances, with ~1.3 M annotations (segmentation, detection, language) spanning masks, bboxes, grasp poses, and contact points, all with human-verified fine-grained language descriptions. The "11B" denotes ~11 billion enumerable grasp examples. Split into ~120K training images plus Seen / Similar / Novel evaluation sets (~15K images each).
- Train an R1-style RFT model (Reasoning-then-Response, with rule-based reward shaping) on the unified dataset, reinforcement-fine-tuning a pretrained Qwen2.5-VL backbone (3B and 7B variants).
The released model jointly predicts all four output modalities from natural language and supports zero-shot perception/manipulation in unseen real-world environments. On the Seen evaluation set, the 7B RealVLG-R1 reports bbox gIoU 89.0%, segmentation FΞ² 88.9%, grasp mIoU 33.6%, and contact-point gAcc 37.2%, establishing baselines for unified cross-task grounding.
RealVLG-R1 is the CVPR-2026 grounding equivalent of CVPR-2025's RoboSpatial β but with R1-style RFT training applied and four output modalities unified. Likely to become the standard grounding-benchmark dataset for VLA-adjacent perception, replacing per-task baselines.
- arXiv: 2603.14880
- Code/data: github.com/lif314/RealVLG-R1
- Embodied-R1 Β· VLM4VLA Β· RoboSpatial (predecessor)
- CVPR 2026 survey
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)