-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 RoboInter
Venue: ICLR 2026 Affiliation: USTC Β· Shanghai AI Lab Β· Beihang Β· NTU Β· Zhejiang Β· Tsinghua Β· CUHK Category: Data + Benchmark + Model β Intermediate Representations for Manipulation Trend tag: Plan-then-execute supervision Β· Flexible CoT Β· VLM Planner + VLA Executor
flowchart LR
subgraph TOOL[RoboInter-Tool β GUI annotation pipeline]
Vid[Manipulation videos<br/>DROID + RH20T + OXE] --> Pre[GPT pre-annotation<br/>15 primitive skills]
Pre --> Clip[Human clip+keyframe+lang+contact]
Clip --> SAM[SAM2 segmentation + tracking]
SAM --> Trace[2D gripper trace<br/>via calibration + point tracking]
Trace --> Post[Post-process:<br/>grasp affordance Β· placement Β· gripper bbox]
end
Post --> DATA[RoboInter-Data<br/>230k episodes Β· 571 scenes Β· 6 arms<br/>61M obj-bbox Β· 70M trace Β· 760k lang]
DATA --> VQA[RoboInter-VQA<br/>9 spatial + 20 temporal categories<br/>~2.2M training entries]
DATA --> PLAN[Planner β Qwen2.5-VL / LLaVA-OV<br/>autoregressive VQA]
VQA --> PLAN
PLAN --> VLAS[RoboInter-VLA family]
subgraph VLAS_DETAIL[Three variants: shared VLM backbone, DiT action head]
IC[IC-E2E: copy Planner VLM,<br/>no explicit CoT]
EC[EC-E2E: copy Planner VLM,<br/>joint CoT+action loss]
MOD[Modular: separate Planner+Executor,<br/>F-CoT bridge, async inference]
end
VLAS --> IC
VLAS --> EC
VLAS --> MOD
The plan-then-execute paradigm (e.g., ECoT, Οβ.5, ReKep, VLA-OS) requires fine-grained intermediate representations β sub-tasks, primitive skills, 2D traces, grasp poses, affordance boxes, object/gripper boxes, contact points β but existing datasets supply at most one or two of these. LLARVA has only auto-generated traces; ECoT has auto-generated CoT text and grounding; ShareRobot has human-verified data at only 51k videos; Robo2VLM samples frames and lacks temporal alignment with actions. None provide per-frame dense annotations across all categories with temporally-aligned actions and second-camera views at scale.
This data gap forces every plan-then-execute paper to invent its own ad-hoc IR labels, prevents pre-training planners at scale, and makes cross-paper comparisons impossible.
A lightweight GUI for semi-automatic per-frame annotation on third-person videos. The pipeline:
- Video check & primitive-skill segmentation. GPT pre-annotates language; human annotators split clips into 15 predefined primitive skills and add clip-level + video-level language. Contact frames recorded manually.
- SAM2 propagation. Annotator clicks on the manipulated object; SAM2 (Ravi et al. 2024) segments and tracks, returned asynchronously for review/correction.
- Camera calibration. When intrinsics are missing in raw datasets, a calibration matrix is estimated; gripper detection + point tracking fill in parameter-missing episodes.
- Post-processed annotations. Grasp affordance from 2D end-effector at contact frame; contact points from gripper key points; grasp pose from 6D end-effector at contact; placement from end-of-subtask object pose; gripper bbox via 3D-to-2D projection of anchor points.
Built on DROID + OXE (for in-the-wild diversity) and RH20T (for table-top skill diversity).
| Statistic | Value |
|---|---|
| Episodes | 230k |
| Scenes | 571 |
| Robot arms | 6 |
| Primitive skills | 15 (object sorting, transferring, pick, place, press, fold, push, pull, twist, β¦) |
| Object-grounding frames | ~61M |
| Gripper-trace frames | ~70M |
| Affordance/placement annotations | ~190k |
| Language clip annotations | ~760k |
| Camera views | 3rd-person + wrist |
| Video resolution | 640Γ360 |
Comparison vs. prior IR datasets (Table 1):
| Dataset | Videos | Scenes | Dense | Emb-VQA | E2E-Act | Curated-CoT | Subtask | Affordance | Contact | Gripper-bbox | Obj-bbox | Trace | Annotation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LLARVA | β | 311 | β | β | β | β | β | β | β | β | β | β | Auto |
| Hamster | 136k | β | β | β | β | β | β | β | β | β | β | β | Auto |
| RH20T-P | 38k | 7 | β | β | β | β | β | β | β | β | β | β | Human+Auto |
| ECoT | 60k | 12 | β | β | β | β | β | β | β | β | β | β | Auto |
| AgiBot-World | 1M | 106 | β | β | β | β | β | β | β | β | β | β | Human |
| VLA-OS | 10k | β | β | β | β | β | β | β | β | β | β | β | Auto |
| ShareRobot | 51k | 102 | β | β | β | β | β | β | β | β | β | β | Human+Auto |
| Robo2VLM | 176k | 463 | β | β | β | β | β | β | β | β | β | β | Auto |
| VeBrain | 12k | β | β | β | β | β | β | β | β | β | β | β | Human |
| RoboInter | 230k | 571 | β | β | β | β | β | β | β | β | β | β | Human+Auto |
RoboInter is the only dataset offering all checked properties simultaneously.
VQA tasks across two axes: spatial vs. temporal, understanding vs. generation.
| Bucket | Categories |
|---|---|
| Spatial generation (5) | Object bbox, grasp pose, placement proposal, key points, gripper bbox |
| Spatial understanding (3+1) | bbox/grasp-pose/scene multiple choice + contact T/F |
| Temporal generation | Trace generation, multi-step planning |
| Temporal understanding (5+4) | Move direction, traceβdescription match, subtask/primitive discrimination, stage ID, task success / next-step feasibility judgements |
Scale: ~1M spatial generation, 172k spatial understanding, 131k temporal generation, 935k temporal understanding training entries. 7,246 videos withheld for evaluation.
Planner. Qwen2.5-VL-3B / -7B / LLaVA-OneVision-7B with multi-image support. Autoregressive cross-entropy loss on VQA + grounding data.
Executor. Qwen2.5-VL backbone with a Diffusion Transformer (DiT) action head (Peebles & Xie 2023). Consumes multi-view RGB (third-person + wrist), language, and intermediate representations. An information aggregator compresses VLM hidden states + IR tokens into fixed-length conditioning. Outputs multi-step action chunks via diffusion loss.
Three plan-then-execute variants:
- IC-E2E (Implicitly-Conditioned End-to-End): inject the Planner's VLM weights into the Executor without explicit CoT. The VLM acts as a stronger feature extractor.
- EC-E2E (Explicitly-Conditioned End-to-End): init from the Planner, jointly optimise CoT generation and action β modality interference is the trade-off.
- Modular: Planner and Executor are independent. Training uses GT intermediates; inference uses Planner-predicted intermediates. Supports async two-process inference.
Flexible Chain-of-Thought (F-CoT). A composable CoT of multiple IRs (subtask, skill, object box, grasp pose, gripper box, affordance, trace). Two forms: textual F-CoT (RoboInter-Te-Modular) and visual-prompted F-CoT (RoboInter-Im-Modular).
At deployment, the EC-E2E variant uses a shortened CoT (subtask + affordance box + gripper box only) with a caching mechanism that refreshes slow components (subtask, affordance) every 30 control steps β limiting control frequency to under 10 Hz on an online A800 cluster.
| Model | Where2Place β | RoboRefIt β | RoboVQA β | RefCOCO-g-val β | RefCOCO+val β | RefCOCO val β | TextVQA β | COCO β | OCR β | MME β | MMVET β | POPE β |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| QwenVL2.5-3B | 11.8 | 68.9 | 37.6 | 85.2 | 82.4 | 89.1 | 79.3 | 15.7 | 797 | 2175 | 61.8 | 85.9 |
| QwenVL2.5-7B | 18.9 | 75.8 | 38.4 | 87.2 | 84.2 | 90.2 | 84.9 | 15.0 | 864 | 2306 | 67.1 | 85.9 |
| RoboBrain-2.0-3B | 59.8 | 30.9 | 30.6 | 55.0 | 51.5 | 50.9 | 81.0 | 27.2 | 811 | 2126 | 59.4 | 88.1 |
| RoboBrain-2.0-7B | 63.6 | 8.8 | 31.6 | 62.9 | 70.1 | 76.1 | 75.9 | 25.2 | 857 | 2076 | 61.4 | 86.2 |
| RoboInter-Qwen-3B | 58.3 | 80.0 | 43.3 | 87.9 | 85.8 | 89.5 | 78.9 | 15.6 | 787 | 2180 | 61.0 | 90.5 |
| RoboInter-Qwen-7B | 65.8 | 85.6 | 74.4 | 88.4 | 86.6 | 91.5 | 83.0 | 15.9 | 832 | 2281 | 62.3 | 91.4 |
| RoboInter-LLaVAOV-7B | 66.3 | 89.3 | 74.5 | 87.3 | 84.2 | 91.3 | 72.2 | 15.8 | 725 | 2217 | 61.4 | 90.4 |
RoboInter-Qwen-3B gains +49.1% on RoboRefIt over RoboBrain-2.0-3B; 7B variant gains +76.8% RoboRefIt and +42.8% RoboVQA over RoboBrain-2.0-7B. General benchmark performance is stable (slight degradation on TextVQA / MME), confirming that embodied training does not crash general VLM ability.
| Model | Obj G.D. β | Grasp Aff β | Place Aff β | Gripper G.D. β | Grasp Choice β | G.D. Choice β | Contact T/F β | Trace DTW β | Visual Trace β | Plan Choice β | Task Plan T/F β |
|---|---|---|---|---|---|---|---|---|---|---|---|
| QwenVL2.5-7B | 51.2 | 14.7 | 38.2 | 10.2 | 27.3 | 25.7 | 52.5 | 1702 | 22.4 | 39.0 | 60.5 |
| Gemini-2.5-Flash | 1.7 | β | 1.2 | β | 32.7 | 69.4 | 65.5 | β | β | 49.4 | β |
| RoboBrain-2.0-7B | β | β | β | 2.5 | 23.3 | 21.5 | 49.2 | 541 | 15.3 | 29.5 | 46.4 |
| RoboInter-Qwen-7B | 75.1 | 37.8 | 56.9 | 62.0 | 76.1 | 75.7 | 75.6 | 323 | 63.4 | 81.9 | 93.0 |
| RoboInter-LLaVAOV-7B | 82.9 | 46.3 | 55.1 | 70.1 | 74.1 | 79.7 | 76.3 | 299 | 62.7 | 81.9 | 83.9 |
General VLMs sit near random (β25%) on most spatial choice tasks; RoboInter triples or quadruples them on generation tasks.
OLS@0.1 means a per-step action is "correct" if within 0.1 L_β of GT; mOLS averages thresholds {0.1, 0.05, 0.03, 0.01}.
| Method | OLS@0.1 | @0.05 | @0.03 | @0.01 | mOLS |
|---|---|---|---|---|---|
| VLA-OS | 0.6180 | 0.3905 | 0.1928 | 0.0129 | 0.3035 |
| Vanilla (no Planner) | 0.6793 | 0.3608 | 0.1753 | 0.0189 | 0.3086 |
| RoboInter-IC-E2E | 0.6984 | 0.3810 | 0.1873 | 0.0204 | 0.3218 |
| RoboInter-EC-E2E | 0.7049 | 0.3930 | 0.2066 | 0.0314 | 0.3340 |
| QwenVL+Executor | 0.6749 | 0.3582 | 0.1777 | 0.0298 | 0.3102 |
| RoboInter-Te-Modular | 0.7124 | 0.4133 | 0.2332 | 0.0584 | 0.3543 |
| RoboInter-Im-Modular | 0.7056 | 0.4029 | 0.2240 | 0.0430 | 0.3439 |
| Oracle+VLA-OS | 0.7260 | 0.4928 | 0.2734 | 0.0200 | 0.3780 |
| Oracle+Executor | 0.7511 | 0.4640 | 0.2705 | 0.0587 | 0.3861 |
Explicit > implicit, modular > end-to-end (in open-loop). Te-Modular > Im-Modular because dense images dilute IR signal.
| Model | Object Coll. (ID/OOD) | Cup Stack | Towel Fold | Clutter Clean | IDβOOD drop |
|---|---|---|---|---|---|
| OpenVLA | 53.3 / 20.0 | 66.7 / 33.3 | 26.7 / 6.7 | 33.3 / 33.3 | β21.7 |
| Οβ | 73.3 / 46.7 | 80.0 / 53.3 | 53.3 / 40.0 | 46.7 / 40.0 | β18.3 |
| Vanilla | 66.7 / 33.3 | 80.0 / 46.7 | 46.7 / 20.0 | 66.7 / 53.3 | β26.7 |
| RoboInter-IC-E2E | 86.7 / 53.3 | 86.7 / 60.0 | 60.0 / 46.7 | 73.3 / 73.3 | β19.0 |
| RoboInter-EC-E2E | 73.3 / 60.0 | 80.0 / 73.3 | 46.7 / 40.0 | 73.3 / 66.7 | β8.3 |
| RoboInter-Modular | 66.7 / 53.3 | 86.7 / 73.3 | 53.3 / 40.0 | 80.0 / 73.3 | β11.7 |
- IC-E2E wins ID (77.3% avg) but degrades 19 points to OOD (58.3%).
- EC-E2E is lower ID (68.3%) but the tightest IDβOOD drop (β8.3) and the highest OOD average. Conclusion: explicit CoT trades a bit of ID accuracy for substantial OOD generalisation.
- Modular is competitive but slightly more sensitive to distribution shift due to async two-process inference.
| Variant | OLS@0.1 | @0.05 | @0.03 | @0.01 | mOLS |
|---|---|---|---|---|---|
| Vanilla | 0.6793 | 0.3608 | 0.1753 | 0.0189 | 0.3086 |
| + Subtask | 0.6965 | 0.3676 | 0.1770 | 0.0171 | 0.3146 |
| + Primitive Skill | 0.6983 | 0.3681 | 0.1779 | 0.0194 | 0.3159 |
| + Object Box | 0.7025 | 0.3849 | 0.1988 | 0.0294 | 0.3289 |
| + Gripper Box | 0.7212 | 0.4032 | 0.2048 | 0.0272 | 0.3391 |
| + Affordance | 0.7245 | 0.4083 | 0.2114 | 0.0297 | 0.3435 |
| + Trace | 0.7511 | 0.4640 | 0.2705 | 0.0587 | 0.3861 |
Coarse-grained labels (subtask, primitive skill) help marginally; fine-grained spatial labels (boxes, affordance) drive the bulk of the gain; trace is the single biggest contributor because it supplies dense temporally-grounded guidance.
- F-CoT scaling (Table 12): scaling F-CoT data continues to improve Modular variants β IR data is not yet saturated.
- VQA scaling (Tables 9-11): spatial-generation, understanding, and planning subsets each scale near-monotonically up to the full corpus.
- Single IR breakdown (Table 13): Trace alone gives the strongest single-representation gain, confirming Table 5's marginal contributions.
- Single-arm, fixed-base only. Annotations come from DROID + OXE + RH20T; mobile-base and dual-arm coverage is missing. Cross-embodiment generalisation is therefore constrained.
- Planner generalisation to unseen scenarios remains an open challenge despite the strong VQA gains.
- Heavy inference cost of long-chain reasoning. The Executor depends on long F-CoT contexts and is limited to <10 Hz on an A800 cluster with caching. Future work points to lighter architectures and acceleration strategies (Tables 7-8 show the inference-latency / accuracy trade-off explicitly).
- Open-loop OLS is closer to OOD than to ID closed-loop β the metric and real-world correlation are not perfectly aligned (EC-E2E excels open-loop but is bested by IC-E2E in ID closed-loop).
RoboInter elevates intermediate representations from per-paper ad-hoc labels to a first-class data product. Compared with:
- LLARVA / Hamster: trace-only, automatic. RoboInter adds 9 other IR categories and human verification.
- ECoT: auto-generated CoT via Gemini, no second-person view, no aligned actions. RoboInter has dense per-frame human-verified annotations aligned with executed actions.
- AgiBot-World: 1M videos but only subtask labels and only one IR category.
- VLA-OS: has many IR categories but only 10k videos auto-annotated.
- RoboBrain-2.0: an embodied VLM Planner trained on a smaller embodied corpus; RoboInter-Qwen beats RoboBrain-2.0 by 49-77% on RoboRefIt.
- From Seeing to Doing β spatial-CoT VLA that implicitly needs the kind of supervision RoboInter provides.
- Embodied-R1 β pointing-style supervision; RoboInter generalises this to 10+ representation types.
- OmniEVA β adjacent embodied-VLM effort.
- Ο0.6 β uses VLM reasoning implicitly; RoboInter offers the explicit CoT data needed to reproduce or improve such systems.
The empirical contribution of the three-variant comparison (IC-E2E > Vanilla; EC-E2E generalises best to OOD; Modular bests ID under perfect IRs) provides a methodological map for plan-then-execute VLA design choices that previously was anecdotal.
- OpenReview: https://openreview.net/forum?id=PGUC3mmMoi
- Project: https://lihaohn.github.io/RoboInter.github.io
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)