Skip to content

CoRL 2025

hwoo.han edited this page Jun 11, 2026 · 3 revisions

CoRL 2025

Conference on Robot Learning 2025 β€” Seoul, Sept 27–30, 2025.

263 accepted papers (42 orals + 221 posters). VLA/manipulation was the dominant thread; NVIDIA launched Isaac GR00T N1.6 + the Newton physics engine at the venue.

Surveys hosted in this wiki

Awards

Award Paper
Best Paper Fabrica β€” dual-arm general multi-part assembly
Best Paper UniFP β€” unified position+force policy for legged loco-manipulation
Best Student Paper Visual Imitation β†’ Humanoid β€” contextual humanoid control from internet video
Best Paper Finalist DexUMI, DSRL, LocoFormer, The Sound of Simulation, Ο€0.5

Quick paper index (by category)

Flagship baseline

  • Ο€0.5 (Oral) β€” Physical Intelligence's hierarchical VLA with co-training; sets up the entire 2025β†’2026 Ο€ series. See Ο€ series evolution.

VLA architecture

  • DexVLA β€” plug-in ~1B diffusion action expert atop a VLM
  • TA-VLA β€” single torque-history token in the decoder for contact-rich tasks
  • Streaming Flow Policy (Oral) β€” action chunk = point on a longer flow

Training & inference recipes

  • ECoT-Lite β€” which parts of embodied chain-of-thought actually matter
  • RoboMonkey β€” best-of-N VLA sampling with a learned verifier

Dexterous / humanoid / whole-body

  • DexUMI (Finalist) β€” wearable "universal manipulation interface"
  • ClutterDexGrasp (Oral) β€” zero-shot sim-to-real closed-loop dex grasping in clutter
  • UniFP (Best Paper) β€” unified position+force legged loco-manipulation
  • Visual Imitation β†’ Humanoid (Best Student Paper) β€” everyday video β†’ humanoid skills
  • Fabrica (Best Paper) β€” dual-arm multi-part assembly

Cross-embodiment from human video

  • X-Sim (Oral) β€” real-to-sim-to-real via object motion

Diffusion / flow policies + RL

  • DSRL (Oral, Finalist) β€” RL in the initial-noise latent space of a frozen diffusion policy

World models for policy

  • DreamGen β€” policy training inside a video world model (NVIDIA)

Data & benchmarks

  • ManipBench β€” first VLM benchmark targeting low-level manipulation reasoning

Trends (6)

  1. Hierarchical VLAs beat monolithic VLAs for open-world generalization (Ο€0.5, OneTwoVLA, Long-VLA).
  2. Test-time scaling + trajectory streaming are the cheap latency/quality levers (RoboMonkey, Streaming Flow Policy, DemoSpeedup, SAIL).
  3. Human video is the default cross-embodiment data source (DexUMI, Visual Imitation β†’ Humanoid, UniSkill, ImMimic, X-Sim).
  4. Diffusion / flow policies get RL-ified without log-probs (DSRL, DiWA) β€” prefigures ICLR 2026's RECAP, RL Tokens, VLA-RFT, SimpleVLA-RL.
  5. Contact / force is finally modeled inside VLAs (TA-VLA, DexSkin, UniFP, KineSoft, Tactile Beyond Pixels).
  6. World models move inside (not adjacent to) training pipelines (DreamGen, LaDi-WM, FLARE, ParticleFormer) β€” with NVIDIA GR00T N1.6 + Newton as the ecosystem signal.

← Back to CoRL Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally