Skip to content

ICLR 2026 RoboInter

Heungwoo edited this page Jun 1, 2026 · 3 revisions

RoboInter β€” Holistic Intermediate Representation Suite

Venue: ICLR 2026 Affiliation: USTC Β· Shanghai AI Lab Β· Beihang Β· NTU Β· Zhejiang Β· Tsinghua Β· CUHK Category: Data + Benchmark + Model β€” Intermediate Representations for Manipulation Trend tag: Plan-then-execute supervision Β· Flexible CoT Β· VLM Planner + VLA Executor

Approach diagram

flowchart LR
  subgraph TOOL[RoboInter-Tool β€” GUI annotation pipeline]
    Vid[Manipulation videos<br/>DROID + RH20T + OXE] --> Pre[GPT pre-annotation<br/>15 primitive skills]
    Pre --> Clip[Human clip+keyframe+lang+contact]
    Clip --> SAM[SAM2 segmentation + tracking]
    SAM --> Trace[2D gripper trace<br/>via calibration + point tracking]
    Trace --> Post[Post-process:<br/>grasp affordance Β· placement Β· gripper bbox]
  end
  Post --> DATA[RoboInter-Data<br/>230k episodes Β· 571 scenes Β· 6 arms<br/>61M obj-bbox Β· 70M trace Β· 760k lang]
  DATA --> VQA[RoboInter-VQA<br/>9 spatial + 20 temporal categories<br/>~2.2M training entries]
  DATA --> PLAN[Planner β€” Qwen2.5-VL / LLaVA-OV<br/>autoregressive VQA]
  VQA --> PLAN
  PLAN --> VLAS[RoboInter-VLA family]
  subgraph VLAS_DETAIL[Three variants: shared VLM backbone, DiT action head]
    IC[IC-E2E: copy Planner VLM,<br/>no explicit CoT]
    EC[EC-E2E: copy Planner VLM,<br/>joint CoT+action loss]
    MOD[Modular: separate Planner+Executor,<br/>F-CoT bridge, async inference]
  end
  VLAS --> IC
  VLAS --> EC
  VLAS --> MOD
Loading

Problem

The plan-then-execute paradigm (e.g., ECoT, Ο€β‚€.5, ReKep, VLA-OS) requires fine-grained intermediate representations β€” sub-tasks, primitive skills, 2D traces, grasp poses, affordance boxes, object/gripper boxes, contact points β€” but existing datasets supply at most one or two of these. LLARVA has only auto-generated traces; ECoT has auto-generated CoT text and grounding; ShareRobot has human-verified data at only 51k videos; Robo2VLM samples frames and lacks temporal alignment with actions. None provide per-frame dense annotations across all categories with temporally-aligned actions and second-camera views at scale.

This data gap forces every plan-then-execute paper to invent its own ad-hoc IR labels, prevents pre-training planners at scale, and makes cross-paper comparisons impossible.

Method (detailed)

Component 1 β€” RoboInter-Tool (GUI annotation pipeline)

A lightweight GUI for semi-automatic per-frame annotation on third-person videos. The pipeline:

  1. Video check & primitive-skill segmentation. GPT pre-annotates language; human annotators split clips into 15 predefined primitive skills and add clip-level + video-level language. Contact frames recorded manually.
  2. SAM2 propagation. Annotator clicks on the manipulated object; SAM2 (Ravi et al. 2024) segments and tracks, returned asynchronously for review/correction.
  3. Camera calibration. When intrinsics are missing in raw datasets, a calibration matrix is estimated; gripper detection + point tracking fill in parameter-missing episodes.
  4. Post-processed annotations. Grasp affordance from 2D end-effector at contact frame; contact points from gripper key points; grasp pose from 6D end-effector at contact; placement from end-of-subtask object pose; gripper bbox via 3D-to-2D projection of anchor points.

Component 2 β€” RoboInter-Data

Built on DROID + OXE (for in-the-wild diversity) and RH20T (for table-top skill diversity).

Statistic Value
Episodes 230k
Scenes 571
Robot arms 6
Primitive skills 15 (object sorting, transferring, pick, place, press, fold, push, pull, twist, …)
Object-grounding frames ~61M
Gripper-trace frames ~70M
Affordance/placement annotations ~190k
Language clip annotations ~760k
Camera views 3rd-person + wrist
Video resolution 640Γ—360

Comparison vs. prior IR datasets (Table 1):

Dataset Videos Scenes Dense Emb-VQA E2E-Act Curated-CoT Subtask Affordance Contact Gripper-bbox Obj-bbox Trace Annotation
LLARVA – 311 βœ— βœ— βœ“ βœ— βœ— βœ— βœ— βœ— βœ— βœ“ Auto
Hamster 136k – βœ“ βœ“ βœ“ βœ— βœ— βœ— βœ— βœ— βœ— βœ“ Auto
RH20T-P 38k 7 βœ— βœ“ βœ— βœ— βœ“ βœ— βœ“ βœ— βœ— βœ— Human+Auto
ECoT 60k 12 βœ“ βœ— βœ“ βœ“ βœ“ βœ— βœ— βœ“ βœ“ βœ— Auto
AgiBot-World 1M 106 βœ“ βœ— βœ“ βœ— βœ“ βœ— βœ— βœ— βœ— βœ— Human
VLA-OS 10k – βœ“ βœ— βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ Auto
ShareRobot 51k 102 βœ— βœ“ βœ— βœ— βœ“ βœ“ βœ“ βœ— βœ— βœ“ Human+Auto
Robo2VLM 176k 463 βœ— βœ“ βœ— βœ— βœ“ βœ— βœ“ βœ— βœ— βœ“ Auto
VeBrain 12k – βœ“ βœ“ βœ— βœ— βœ“ βœ— βœ“ βœ— βœ— βœ— Human
RoboInter 230k 571 βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ Human+Auto

RoboInter is the only dataset offering all checked properties simultaneously.

Component 3 β€” RoboInter-VQA

VQA tasks across two axes: spatial vs. temporal, understanding vs. generation.

Bucket Categories
Spatial generation (5) Object bbox, grasp pose, placement proposal, key points, gripper bbox
Spatial understanding (3+1) bbox/grasp-pose/scene multiple choice + contact T/F
Temporal generation Trace generation, multi-step planning
Temporal understanding (5+4) Move direction, trace–description match, subtask/primitive discrimination, stage ID, task success / next-step feasibility judgements

Scale: ~1M spatial generation, 172k spatial understanding, 131k temporal generation, 935k temporal understanding training entries. 7,246 videos withheld for evaluation.

Component 4 β€” RoboInter-VLA (Planner + Executor)

Planner. Qwen2.5-VL-3B / -7B / LLaVA-OneVision-7B with multi-image support. Autoregressive cross-entropy loss on VQA + grounding data.

Executor. Qwen2.5-VL backbone with a Diffusion Transformer (DiT) action head (Peebles & Xie 2023). Consumes multi-view RGB (third-person + wrist), language, and intermediate representations. An information aggregator compresses VLM hidden states + IR tokens into fixed-length conditioning. Outputs multi-step action chunks via diffusion loss.

Three plan-then-execute variants:

  • IC-E2E (Implicitly-Conditioned End-to-End): inject the Planner's VLM weights into the Executor without explicit CoT. The VLM acts as a stronger feature extractor.
  • EC-E2E (Explicitly-Conditioned End-to-End): init from the Planner, jointly optimise CoT generation and action β€” modality interference is the trade-off.
  • Modular: Planner and Executor are independent. Training uses GT intermediates; inference uses Planner-predicted intermediates. Supports async two-process inference.

Flexible Chain-of-Thought (F-CoT). A composable CoT of multiple IRs (subtask, skill, object box, grasp pose, gripper box, affordance, trace). Two forms: textual F-CoT (RoboInter-Te-Modular) and visual-prompted F-CoT (RoboInter-Im-Modular).

At deployment, the EC-E2E variant uses a shortened CoT (subtask + affordance box + gripper box only) with a caching mechanism that refreshes slow components (subtask, affordance) every 30 control steps β€” limiting control frequency to under 10 Hz on an online A800 cluster.

Comprehensive Results

Third-party VLM benchmarks (Table 2)

Model Where2Place ↑ RoboRefIt ↑ RoboVQA ↑ RefCOCO-g-val ↑ RefCOCO+val ↑ RefCOCO val ↑ TextVQA ↑ COCO ↑ OCR ↑ MME ↑ MMVET ↑ POPE ↑
QwenVL2.5-3B 11.8 68.9 37.6 85.2 82.4 89.1 79.3 15.7 797 2175 61.8 85.9
QwenVL2.5-7B 18.9 75.8 38.4 87.2 84.2 90.2 84.9 15.0 864 2306 67.1 85.9
RoboBrain-2.0-3B 59.8 30.9 30.6 55.0 51.5 50.9 81.0 27.2 811 2126 59.4 88.1
RoboBrain-2.0-7B 63.6 8.8 31.6 62.9 70.1 76.1 75.9 25.2 857 2076 61.4 86.2
RoboInter-Qwen-3B 58.3 80.0 43.3 87.9 85.8 89.5 78.9 15.6 787 2180 61.0 90.5
RoboInter-Qwen-7B 65.8 85.6 74.4 88.4 86.6 91.5 83.0 15.9 832 2281 62.3 91.4
RoboInter-LLaVAOV-7B 66.3 89.3 74.5 87.3 84.2 91.3 72.2 15.8 725 2217 61.4 90.4

RoboInter-Qwen-3B gains +49.1% on RoboRefIt over RoboBrain-2.0-3B; 7B variant gains +76.8% RoboRefIt and +42.8% RoboVQA over RoboBrain-2.0-7B. General benchmark performance is stable (slight degradation on TextVQA / MME), confirming that embodied training does not crash general VLM ability.

RoboInter-VQA benchmark (Table 3; ACC unless stated)

Model Obj G.D. ↑ Grasp Aff ↑ Place Aff ↑ Gripper G.D. ↑ Grasp Choice ↑ G.D. Choice ↑ Contact T/F ↑ Trace DTW ↓ Visual Trace ↑ Plan Choice ↑ Task Plan T/F ↑
QwenVL2.5-7B 51.2 14.7 38.2 10.2 27.3 25.7 52.5 1702 22.4 39.0 60.5
Gemini-2.5-Flash 1.7 – 1.2 – 32.7 69.4 65.5 – – 49.4 –
RoboBrain-2.0-7B – – – 2.5 23.3 21.5 49.2 541 15.3 29.5 46.4
RoboInter-Qwen-7B 75.1 37.8 56.9 62.0 76.1 75.7 75.6 323 63.4 81.9 93.0
RoboInter-LLaVAOV-7B 82.9 46.3 55.1 70.1 74.1 79.7 76.3 299 62.7 81.9 83.9

General VLMs sit near random (β‰ˆ25%) on most spatial choice tasks; RoboInter triples or quadruples them on generation tasks.

Open-Loop Score (OLS) on Executor β€” In-the-Wild (Table 4)

OLS@0.1 means a per-step action is "correct" if within 0.1 L_∞ of GT; mOLS averages thresholds {0.1, 0.05, 0.03, 0.01}.

Method OLS@0.1 @0.05 @0.03 @0.01 mOLS
VLA-OS 0.6180 0.3905 0.1928 0.0129 0.3035
Vanilla (no Planner) 0.6793 0.3608 0.1753 0.0189 0.3086
RoboInter-IC-E2E 0.6984 0.3810 0.1873 0.0204 0.3218
RoboInter-EC-E2E 0.7049 0.3930 0.2066 0.0314 0.3340
QwenVL+Executor 0.6749 0.3582 0.1777 0.0298 0.3102
RoboInter-Te-Modular 0.7124 0.4133 0.2332 0.0584 0.3543
RoboInter-Im-Modular 0.7056 0.4029 0.2240 0.0430 0.3439
Oracle+VLA-OS 0.7260 0.4928 0.2734 0.0200 0.3780
Oracle+Executor 0.7511 0.4640 0.2705 0.0587 0.3861

Explicit > implicit, modular > end-to-end (in open-loop). Te-Modular > Im-Modular because dense images dilute IR signal.

Closed-loop real-world (Franka Research 3, Table 6, 15 ID + 15 OOD trials per task)

Model Object Coll. (ID/OOD) Cup Stack Towel Fold Clutter Clean ID→OOD drop
OpenVLA 53.3 / 20.0 66.7 / 33.3 26.7 / 6.7 33.3 / 33.3 βˆ’21.7
Ο€β‚€ 73.3 / 46.7 80.0 / 53.3 53.3 / 40.0 46.7 / 40.0 βˆ’18.3
Vanilla 66.7 / 33.3 80.0 / 46.7 46.7 / 20.0 66.7 / 53.3 βˆ’26.7
RoboInter-IC-E2E 86.7 / 53.3 86.7 / 60.0 60.0 / 46.7 73.3 / 73.3 βˆ’19.0
RoboInter-EC-E2E 73.3 / 60.0 80.0 / 73.3 46.7 / 40.0 73.3 / 66.7 βˆ’8.3
RoboInter-Modular 66.7 / 53.3 86.7 / 73.3 53.3 / 40.0 80.0 / 73.3 βˆ’11.7
  • IC-E2E wins ID (77.3% avg) but degrades 19 points to OOD (58.3%).
  • EC-E2E is lower ID (68.3%) but the tightest IDβ†’OOD drop (βˆ’8.3) and the highest OOD average. Conclusion: explicit CoT trades a bit of ID accuracy for substantial OOD generalisation.
  • Modular is competitive but slightly more sensitive to distribution shift due to async two-process inference.

Ablation Studies

IR composition (Table 5, Oracle+Executor on In-the-Wild)

Variant OLS@0.1 @0.05 @0.03 @0.01 mOLS
Vanilla 0.6793 0.3608 0.1753 0.0189 0.3086
+ Subtask 0.6965 0.3676 0.1770 0.0171 0.3146
+ Primitive Skill 0.6983 0.3681 0.1779 0.0194 0.3159
+ Object Box 0.7025 0.3849 0.1988 0.0294 0.3289
+ Gripper Box 0.7212 0.4032 0.2048 0.0272 0.3391
+ Affordance 0.7245 0.4083 0.2114 0.0297 0.3435
+ Trace 0.7511 0.4640 0.2705 0.0587 0.3861

Coarse-grained labels (subtask, primitive skill) help marginally; fine-grained spatial labels (boxes, affordance) drive the bulk of the gain; trace is the single biggest contributor because it supplies dense temporally-grounded guidance.

Other ablations

  • F-CoT scaling (Table 12): scaling F-CoT data continues to improve Modular variants β€” IR data is not yet saturated.
  • VQA scaling (Tables 9-11): spatial-generation, understanding, and planning subsets each scale near-monotonically up to the full corpus.
  • Single IR breakdown (Table 13): Trace alone gives the strongest single-representation gain, confirming Table 5's marginal contributions.

Limitations (stated by authors, Appendix A.8)

  • Single-arm, fixed-base only. Annotations come from DROID + OXE + RH20T; mobile-base and dual-arm coverage is missing. Cross-embodiment generalisation is therefore constrained.
  • Planner generalisation to unseen scenarios remains an open challenge despite the strong VQA gains.
  • Heavy inference cost of long-chain reasoning. The Executor depends on long F-CoT contexts and is limited to <10 Hz on an A800 cluster with caching. Future work points to lighter architectures and acceleration strategies (Tables 7-8 show the inference-latency / accuracy trade-off explicitly).
  • Open-loop OLS is closer to OOD than to ID closed-loop β€” the metric and real-world correlation are not perfectly aligned (EC-E2E excels open-loop but is bested by IC-E2E in ID closed-loop).

Significance & Positioning

RoboInter elevates intermediate representations from per-paper ad-hoc labels to a first-class data product. Compared with:

  • LLARVA / Hamster: trace-only, automatic. RoboInter adds 9 other IR categories and human verification.
  • ECoT: auto-generated CoT via Gemini, no second-person view, no aligned actions. RoboInter has dense per-frame human-verified annotations aligned with executed actions.
  • AgiBot-World: 1M videos but only subtask labels and only one IR category.
  • VLA-OS: has many IR categories but only 10k videos auto-annotated.
  • RoboBrain-2.0: an embodied VLM Planner trained on a smaller embodied corpus; RoboInter-Qwen beats RoboBrain-2.0 by 49-77% on RoboRefIt.
  • From Seeing to Doing β€” spatial-CoT VLA that implicitly needs the kind of supervision RoboInter provides.
  • Embodied-R1 β€” pointing-style supervision; RoboInter generalises this to 10+ representation types.
  • OmniEVA β€” adjacent embodied-VLM effort.
  • Ο€0.6 β€” uses VLM reasoning implicitly; RoboInter offers the explicit CoT data needed to reproduce or improve such systems.

The empirical contribution of the three-variant comparison (IC-E2E > Vanilla; EC-E2E generalises best to OOD; Modular bests ID under perfect IRs) provides a methodological map for plan-then-execute VLA design choices that previously was anecdotal.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally