Skip to content

ICLR 2026 Embodied R1

Heungwoo edited this page Jun 1, 2026 · 3 revisions

Embodied-R1 β€” Reinforced Embodied Reasoning for General Robotic Manipulation

Venue: ICLR 2026 Category: Training / RL Trend tag: Trends 2 + 3

Approach diagram

flowchart LR
  Img[Image] --> VLM[Pointing VLM]
  Inst[Instruction] --> VLM
  VLM --> R{Pointing primitives}
  R --> REG[REG: refer-to-object]
  R --> RRG[RRG: refer-to-relation]
  R --> OFG[OFG: refer-to-functional-part]
  R --> VTG[VTG: visual trace]
  R --> RFT[Two-stage Reinforced Fine-Tuning<br/>on Embodied-Points-200K]
  RFT --> Pol[Policy with embodiment-agnostic<br/>intermediate representation]
Loading

Problem

Chain-of-thought reasoning is typically learned via supervised fine-tuning on human-written reasoning traces β€” expensive and brittle. Reinforcement-based reasoning (as in R1-style LLMs) may work better but had not been applied to embodied tasks.

Method

A 3B pointing VLM (built on Qwen2.5-VL-3B-Instruct) trained via Reinforced Fine-Tuning (RFT). Intermediate representations are four pointing abilities:

  • REG β€” Referring Expression Grounding (localize an object from a linguistic description)
  • RRG β€” Region Referring Grounding (identify a spatial region from relational language)
  • OFG β€” Object Functional Grounding (identify a functional part / affordance, e.g., handle)
  • VTG β€” Visual Trace Generation (output an ordered point sequence forming a manipulation trajectory)

These primitives are embodiment-agnostic. Training uses a two-stage RFT curriculum with a multi-task reward design on the new Embodied-Points-200K dataset (~200k samples across the four pointing capabilities).

Results

  • State-of-the-art on 11 embodied spatial and pointing benchmarks.
  • Zero-shot generalization (no task-specific fine-tuning): 56.2% average success in SimplerEnv and 87.5% across 8 real-world XArm tasks, a 62% improvement over strong baselines on real-world manipulation.

Significance

Establishes that R1-style RL can be applied to embodied reasoning, not just textual reasoning. The pointing-primitive vocabulary is a promising candidate for the transferable interface that Trend 2 calls for.

Links

  • ICLR 2026 listing

Related pages

← Back to ICLR-2026 Β· Topic: RL

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally