-
Notifications
You must be signed in to change notification settings - Fork 0
CoRL 2025
hwoo.han edited this page Jun 11, 2026
·
3 revisions
Conference on Robot Learning 2025 β Seoul, Sept 27β30, 2025.
263 accepted papers (42 orals + 221 posters). VLA/manipulation was the dominant thread; NVIDIA launched Isaac GR00T N1.6 + the Newton physics engine at the venue.
- VLA & Manipulation Survey (CoRL 2025) β categorized summary of CoRL 2025's VLA/manipulation papers plus trends and a CoRL 2025 β ICLR 2026 lineage map.
| Award | Paper |
|---|---|
| Best Paper | Fabrica β dual-arm general multi-part assembly |
| Best Paper | UniFP β unified position+force policy for legged loco-manipulation |
| Best Student Paper | Visual Imitation β Humanoid β contextual humanoid control from internet video |
| Best Paper Finalist | DexUMI, DSRL, LocoFormer, The Sound of Simulation, Ο0.5 |
- Ο0.5 (Oral) β Physical Intelligence's hierarchical VLA with co-training; sets up the entire 2025β2026 Ο series. See Ο series evolution.
- DexVLA β plug-in ~1B diffusion action expert atop a VLM
- TA-VLA β single torque-history token in the decoder for contact-rich tasks
- Streaming Flow Policy (Oral) β action chunk = point on a longer flow
- ECoT-Lite β which parts of embodied chain-of-thought actually matter
- RoboMonkey β best-of-N VLA sampling with a learned verifier
- DexUMI (Finalist) β wearable "universal manipulation interface"
- ClutterDexGrasp (Oral) β zero-shot sim-to-real closed-loop dex grasping in clutter
- UniFP (Best Paper) β unified position+force legged loco-manipulation
- Visual Imitation β Humanoid (Best Student Paper) β everyday video β humanoid skills
- Fabrica (Best Paper) β dual-arm multi-part assembly
- X-Sim (Oral) β real-to-sim-to-real via object motion
- DSRL (Oral, Finalist) β RL in the initial-noise latent space of a frozen diffusion policy
- DreamGen β policy training inside a video world model (NVIDIA)
- ManipBench β first VLM benchmark targeting low-level manipulation reasoning
- Hierarchical VLAs beat monolithic VLAs for open-world generalization (Ο0.5, OneTwoVLA, Long-VLA).
- Test-time scaling + trajectory streaming are the cheap latency/quality levers (RoboMonkey, Streaming Flow Policy, DemoSpeedup, SAIL).
- Human video is the default cross-embodiment data source (DexUMI, Visual Imitation β Humanoid, UniSkill, ImMimic, X-Sim).
- Diffusion / flow policies get RL-ified without log-probs (DSRL, DiWA) β prefigures ICLR 2026's RECAP, RL Tokens, VLA-RFT, SimpleVLA-RL.
- Contact / force is finally modeled inside VLAs (TA-VLA, DexSkin, UniFP, KineSoft, Tactile Beyond Pixels).
- World models move inside (not adjacent to) training pipelines (DreamGen, LaDi-WM, FLARE, ParticleFormer) β with NVIDIA GR00T N1.6 + Newton as the ecosystem signal.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)