-
Notifications
You must be signed in to change notification settings - Fork 0
ICRA 2026 Topic Planning
Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1β5, 2026 Compiled against the official ICRA 2026 PaperCept program. Paper IDs (e.g.
ThI1I.72) are program session/slot codes. arXiv IDs were confirmed individually where listed; entries without an arXiv link were not located and are left blank rather than guessed.
This page covers the 50 papers grouped under Manipulation Planning & Task-and-Motion Planning in the ICRA 2026 program. The cluster spans the full planning stack: classical sampling/optimization-based motion planning (RRT-Connect variants, tensor/GPU-batched planners, screw-theoretic IK), learning-augmented planning (diffusion/consistency/flow planners, RL-tuned classical planners), LLM/VLM-driven task planning (failure detection, contact-acceptability reasoning, assembly-manual parsing), and long-horizon symbolic + skill planning (TAMP solvers, skill libraries, symbol/skill co-invention). Where the VLA survey tracks the policy frontier, this cluster tracks the deliberative frontier β how a robot decides what to do and in what order, and how to certify the resulting motions are feasible, optimal, or safe.
A recurring theme: foundation models are no longer only the policy. They increasingly sit above a classical planner β as a sequence proposer, a constraint/cost writer, a failure detector, or a demonstration generator β while the geometric/optimization machinery underneath retains the feasibility and optimality guarantees that learned policies still lack.
The largest semantic shift is using VLMs as the deliberative layer rather than the action layer. Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images (ThAT1.6) argues that LLM/VLM failure detectors over-assume full scene understanding and grounds detection in an explicit multi-view scene graph. IMPACT (WeI1I.64, arXiv 2503.10110) uses a VLM to infer which surfaces tolerate contact, emitting an anisotropic cost map for a contact-aware A* β relaxing the collision-free assumption that makes classical planning brittle in clutter. Manual2Skill++ (WeI2I.159, arXiv 2510.16344) parses assembly instruction manuals into connector-aware hierarchical graphs, elevating connectors (screws, pegs) to first-class planning primitives. AdaptPNP (WeI2I.184, arXiv 2511.11052) has a VLM emit a prehensile/non-prehensile plan skeleton refined against a digital-twin object-pose predictor. Seeing Farther and Smarter (ThI2I.132) adds value-guided multi-path reflection to VLM policy optimization for long-horizon reasoning, and TARAD (TuI1I.415) uses LLM-generated demonstrations to bootstrap an affordance-centric diffusion policy.
Classic TAMP work targets the combinatorial blowup of long horizons. Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning (ThBT2.7) attacks the exponential growth of TAMP solve time with object count via learned decomposition for fast replanning in dynamic scenes. SymSkill (WeBT1.8, arXiv 2510.01661) co-invents predicates, operators and skills from unsegmented play data β bridging IL's reactivity with TAMP's compositional generalization (85% single-step in RoboCasa; 11 operators from 5 min of Franka play). From CAD to POMDP (TuI1I.114) casts robotic disassembly sequence planning as a POMDP to handle uncertain, partially observable end-of-life products. MOASIC (WeI1I.256) and Uni-Skill (ThI1I.86) attack long-horizon planning over predefined / self-evolving skill libraries with physics simulation in the loop.
Diffusion/flow/consistency models are increasingly used as fast, multi-modal trajectory generators inside otherwise classical pipelines. Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (ThI2I.183, arXiv 2510.14615; "CAMPD") conditions a classifier-free diffusion U-Net on arbitrary context for 7-DoF planning at a fraction of baseline time. CAPE (ThI2I.63) expands diffusion-policy modes for collision avoidance; ConsistencyPlanner (TuI1I.298) uses fast-sampling consistency models for real-time closed-loop planning. Enhancing Classical Motion Planners Using RL with Safety Guarantees (TuI2I.143) keeps a classical planner's safety while RL-tuning its parameters online. DynDLO (TuI2I.201) learns trajectory planning for dynamic deformable-linear-object manipulation, and KAN Policy (TuI1I.47) uses KolmogorovβArnold networks for smooth trajectories.
A strong "fast + provably good" thread persists. AORRTC (WeI1I.339, arXiv 2505.10542) applies the AO-x meta-algorithm to RRT-Connect, getting RRT-Connect-speed initial solutions and almost-sure asymptotic optimality β solving hard high-DoF problems in milliseconds (Panda 7-DoF, Fetch 8-DoF on MotionBenchMaker). Global Tensor Motion Planning (WeI1I.68, arXiv 2411.19393) reformulates sampling-based planning as pure tensor ops over a random multipartite graph for GPU/TPU batch planning. GeoFIK (ThI2I.300, arXiv 2503.03992) is an analytical screw-theory IK solver for the 7-DoF Franka that enumerates redundancy solutions with free Jacobian. Safety-Critical Dynamic Motion Generation (TuI1I.244) uses differentiable configuration-space distance fields with CBFs; Optimal Dexterity Path Planning (WeI1I.187) maximizes workspace-density dexterity inside a sampling planner.
Echoing ReKep's "VLM-as-constraint-writer" paradigm, several papers plan over geometric/contact constraints rather than dense trajectories. A Closed-Chain Approach to Generating Affordance Joint Trajectories (ThI1I.358) extends screw-based affordance planning while avoiding singular/undesirable configurations. Screw Geometry Meets Bandits (WeI2I.259) incrementally acquires kinesthetic demonstrations (bandit-driven) to build screw-geometry manipulation plans. IMPACT (above) and the affordance/connector framing of Manual2Skill++ also fit this constraint-first lineage, as does A Contact-Driven Framework for Manipulating in the Blind (ThI2I.295), which plans from contact feedback when vision is inadequate.
Beyond pick-and-place, planners increasingly reason about pushing/sliding and contact-mode switches. H-MaP (TuI2I.15, arXiv 2403.10436) is a hybrid sequential planner decoupling object-trajectory from manipulation planning, handling tool use and contact-mode switches. AdaptPNP (above) unifies prehensile + non-prehensile skill selection. Robustness-Aware Tool Selection and Manipulation Planning (TuI2I.88) jointly picks tools and plans contact-rich motions under learned energy-informed robustness guidance. Pack It In (TuI1I.211) plans packing into partially filled containers through contact, and Peg-in-Hole (TuI2I.12) uses passive compliance for error-tolerant insertion.
Rearrangement is its own hard combinatorial planning problem. MO-SeGMan (TuI2I.105, arXiv 2511.01476) is a multi-objective sequential/guided rearrangement planner with a Selective Guided Forward Search for non-monotone, cluttered scenes (feasible on all 9 benchmark tasks). Tidiness Score-Guided MCTS (ThI1I.29) plans tabletop tidying from RGB-D via a learned tidiness score guiding Monte Carlo tree search. Placeit! (WeI1I.124) learns object-placement skills with auto-generated training data.
Belief-space and subgoal planning close the loop with partial observability. Planning Using Belief Summaries (TuI1I.126) does goal-directed articulated-object manipulation from force/proprioception under belief uncertainty. From CAD to POMDP (above) is the disassembly instance. Not Throwing Away My Shot (ThI1I.72) plans long-horizon manipulation with dual subgoals (short-horizon + low-variance) to pick informative subgoals; Learning Composable Skills ("STACK", TuI2I.142) discovers spatial/temporal structure from foundation models for skill composition.
AORRTC β WeI1I.339 Β· arXiv 2505.10542
The cleanest "classical planning still wins on guarantees" result in the cluster. By wrapping RRT-Connect in the AO-x meta-algorithm, AORRTC matches RRT-Connect's initial-solution speed yet converges almost-surely to the optimum in an anytime fashion. On MotionBenchMaker with the Panda (7-DoF) and Fetch (8-DoF), it finds solutions to hard high-DoF instances in milliseconds where prior a.s.a.o. planners couldn't reliably solve in seconds. A reminder that the bar learned planners must clear is high.
SymSkill β WeBT1.8 Β· arXiv 2510.01661
A genuine TAMP-meets-IL synthesis (UPenn GRASP): jointly co-invents predicates, operators, and skills from unlabeled, unsegmented demonstrations, getting TAMP's compositional generalization with IL's real-time reactivity and recovery. 85% single-step success in RoboCasa, composing to multi-step tasks with no extra data; on a real Franka it learns 11 operators from 5 minutes of play data and hits user-specified symbolic goals in real time. Directly addresses the symbol-grounding bottleneck that has limited classical TAMP.
IMPACT β WeI1I.64 Β· arXiv 2503.10110
A ReKep-style relaxation of the collision-free dogma. A VLM infers per-region contact tolerance from object semantics, producing an anisotropic 3D cost map encoding directional push safety; a contact-aware A* then plans semantically-acceptable contact-rich paths through clutter that pure collision-free planners cannot traverse. Builds conceptually on the VLM-as-cost/constraint-writer paradigm of ReKep. (USC LIRA Lab.)
MO-SeGMan β TuI2I.105 Β· arXiv 2511.01476
State-of-the-art constrained multi-object rearrangement (TU Munich / Toussaint & Oguz). A Selective Guided Forward Search relocates only critical obstacles, plus adaptive subgoal refinement removes redundant pick-and-place; lazy evaluation jointly minimizes per-object replanning and robot travel. Generates feasible plans on all 9 benchmark rearrangement tasks with faster solve times and better quality than baselines β important for non-monotone, highly cluttered scenes.
Manual2Skill++ β WeI2I.159 Β· arXiv 2510.16344
Treats connectors as first-class primitives: a VLM extracts structured hierarchical connection graphs (connector type, spec, quantity, placement) from assembly manuals, enabling millimeter-level pose alignment for robust execution. Ships a connector-annotated dataset and a multi-connector-modality simulation benchmark. A concrete instance of LLM/VLM-driven task planning grounded in real document structure rather than free-form prompting.
Accelerated Multi-Modal Motion Planning (CAMPD) β ThI2I.183 Β· arXiv 2510.14615
Representative of the learned-planner-as-generator thread. A classifier-free denoising diffusion U-Net with an attention mechanism conditions on an arbitrary number of sensor-agnostic context parameters, generalizing to unseen environments and producing high-quality multi-modal 7-DoF trajectories at a fraction of the time of state-of-the-art baselines β useful both for deployment and for generating diverse trajectory datasets.
| ID | Title | arXiv |
|---|---|---|
| ThAT1.6 | Robust Task Planning via Failure Detection Using Scene Graph from Multi-View Images | |
| ThBT2.7 | Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning | 2408.06843 |
| ThI1I.242 | Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation | 2505.16547 |
| ThI1I.29 | Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement | 2502.17235 |
| ThI1I.358 | A Closed-Chain Approach to Generating Affordance Joint Trajectories for Robotic Manipulators | |
| ThI1I.397 | Whole-Body Integrated Motion Planning for Aerial Manipulators | 2501.06493 |
| ThI1I.72 | Not Throwing Away My Shot: Planning Ahead with Dual Subgoals in Long-Horizon Robot Manipulation Tasks | |
| ThI1I.86 | Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation | 2603.02623 |
| ThI2I.132 | Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization | 2602.19372 |
| ThI2I.175 | The iMETRO Dynamic Simulation: An Open-Source Simulator for Intravehicular Space Robotics Research | |
| ThI2I.183 | Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models (CAMPD) | 2510.14615 |
| ThI2I.295 | A Contact-Driven Framework for Manipulating in the Blind | 2510.20177 |
| ThI2I.300 | GeoFIK: A Fast and Reliable Geometric Solver for the IK of the Franka Arm Based on Screw Theory | 2503.03992 |
| ThI2I.54 | Distracted Robot: How Visual Clutter Undermine Robotic Manipulation | 2511.22780 |
| ThI2I.63 | CAPE: Context-Aware Diffusion Policy via Proximal Mode Expansion for Collision Avoidance | 2511.22773 |
| TuAT3.4 | DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation | 2510.06199 |
| TuI1I.114 | From CAD to POMDP: Probabilistic Planning for Robotic Disassembly of End-Of-Life Products | 2511.23407 |
| TuI1I.126 | Planning Using Belief Summaries for Goal-Directed Manipulation of Articulated Objects with Force and Proprioception | |
| TuI1I.211 | Pack It In: Packing into Partially Filled Containers through Contact | 2602.12095 |
| TuI1I.244 | Safety-Critical Dynamic Motion Generation for Manipulators Using Differentiable Distance Fields in Configuration Space | 2412.16456 |
| TuI1I.298 | ConsistencyPlanner: Real-Time Planning with Fast-Sampling Consistency Models | |
| TuI1I.415 | TARAD: Task-Aware Robot Affordance-Centric Diffusion Policy Learned from LLM-Generated Demonstrations | |
| TuI1I.47 | KAN Policy: Learning Efficient and Smooth Robotic Trajectories via Kolmogorov-Arnold Networks | |
| TuI2I.105 | MO-SeGMan: Rearrangement Planning Framework for Multi-Objective Sequential and Guided Manipulation in Constrained Environments | 2511.01476 |
| TuI2I.12 | Robust and Error-Tolerant Peg-In-Hole Assembly Using Simple Control | |
| TuI2I.142 | Learning Composable Skills by Discovering Spatial and Temporal Structure with Foundation Models (STACK) | |
| TuI2I.143 | Enhancing Classical Motion Planners Using RL with Safety Guarantees | 2403.18524 |
| TuI2I.15 | H-MaP: An Iterative and Hybrid Sequential Manipulation Planner | 2403.10436 |
| TuI2I.201 | DynDLO: Learning-Based Trajectory Planning for Dynamic Robotic Manipulation of Deformable Linear Objects | |
| TuI2I.272 | Run-Time Optimization of Overall Energy Consumption in Lightweight Collaborative Arms for Repetitive Tasks | |
| TuI2I.303 | Learning to Drive by Imitating Surrounding Vehicles | 2503.05997 |
| TuI2I.407 | A Differential Dynamic Programming Framework for Inverse Reinforcement Learning | 2407.19902 |
| TuI2I.88 | Robustness-Aware Tool Selection and Manipulation Planning with Learned Energy-Informed Guidance | 2506.03362 |
| WeBT1.8 | SymSkill: Symbol and Skill Co-Invention for Data-Efficient and Reactive Long-Horizon Manipulation | 2510.01661 |
| WeBT2.2 | Human2Nav: Learning Crowd Navigation from Human Videos across Robots via Feasibility-Guided Flow Matching | |
| WeBT2.4 | Shifted Flow Policy: Uncertainty-Aware Time Reparameterization for Visuomotor Learning | |
| WeBT2.5 | Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy (DCDP) | 2603.01953 |
| WeI1I.124 | Placeit! A Framework for Learning Robot Object Placement Skills | 2510.09267 |
| WeI1I.179 | Task Generalization with Pathwise Conditioning of Gaussian Process for Learning from Demonstration | |
| WeI1I.187 | Optimal Dexterity Path Planning for Robotic Manipulators Using Rapid Workspace Density Approximation | |
| WeI1I.221 | 3DFacePolicy: Speech-Driven 3D Facial Animation Based on Diffusion Policy | 2409.10848 |
| WeI1I.256 | MOASIC: Skill-Centric Manipulation Planning with Physics Simulation | 2504.16738 |
| WeI1I.339 | AORRTC: Almost-Surely Asymptotically Optimal Planning with RRT-Connect | 2505.10542 |
| WeI1I.64 | IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models | 2503.10110 |
| WeI1I.68 | Global Tensor Motion Planning | 2411.19393 |
| WeI2I.116 | MetaDP: Meta-Manipulation Diffusion Policy for Robotic Manipulation | |
| WeI2I.159 | Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via VisionβLanguage Models | 2510.16344 |
| WeI2I.184 | AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation | 2511.11052 |
| WeI2I.259 | Screw Geometry Meets Bandits: Incremental Acquisition of Demonstrations to Generate Manipulation Plans | 2410.18275 |
| WeI2I.300 | GPU-Accelerated Continuous-Time Successive Convexification for Contact-Implicit Legged Locomotion | 2604.09993 |
- ICRA 2026 Survey
- ReKep β VLM-as-constraint-writer lineage referenced by IMPACT / affordance-constraint planners.
β Back to ICRA-2026-VLA-Manipulation-Survey
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)