Skip to content

ICRA 2026 Topic Perception

Heungwoo edited this page Jun 1, 2026 · 2 revisions

ICRA 2026 β€” Perception for Manipulation (Topic Analysis)

Venue: IEEE ICRA 2026 Β· Vienna, Austria Β· June 1–5, 2026 Compiled against the official ICRA 2026 PaperCept program. Numbers below are taken from the official program and from authors' arXiv/project pages; where a number could not be confirmed against an authoritative source it is omitted rather than guessed.

This page covers the 82 papers in the ICRA 2026 "Perception for Manipulation" cluster β€” work whose primary contribution is seeing well enough to act, rather than the action policy itself. Across modern robot learning, the policy architecture has commoditized faster than the perception that feeds it: diffusion/flow policies and VLAs are increasingly off-the-shelf, but they remain only as good as the object pose, 3D geometry, affordance, or correspondence handed to them. Perception is the manipulation bottleneck, and this cluster is where ICRA 2026 attacks it. Three throughlines dominate: (1) 3D and 4D scene representation (point clouds, Gaussian splatting, neural fields, meshes) as the substrate manipulation policies want; (2) pose, shape, and affordance estimation for known, category-level, and fully novel objects, increasingly with explicit uncertainty; and (3) foundation-model perception (SAM, DINO, video-diffusion, MLLMs) repurposed as zero-shot front-ends for grasping and flow. A recurring sub-current is hard-case perception β€” transparent/specular glassware, deformable linear objects (cables/cloth), articulated mechanisms, and heavily occluded clutter β€” exactly the cases where commodity RGB-D and discriminative depth break.

Sub-trends

1. 6-DoF object pose estimation & tracking

Classical 6-DoF pose remains a live research target, now pushed toward distributions and dynamics rather than single point estimates. SE(3)-PoseFlow (WeI1I.194; arXiv 2511.01501) does flow-matching on the SE(3) manifold to produce a full sample-based pose distribution, explicitly modeling multi-modality from symmetry and occlusion and reporting SOTA on Real275/YCB-V/LM-O. MGS-Track (TuI2I.316) tracks monocular 6-DoF pose via a masked 3D prior plus online Gaussian splatting, targeting depth-sensor-free deployment. Learning 6D Object Pose Estimation with Event Cameras (TuI1I.397) attacks high-speed and adverse-lighting regimes where RGB/RGB-D fail, training on synthetic data with domain randomization. PartPose (ThI2I.337) reframes 6D pose for multi-part deformable objects (cable-attached appliances, pouch drinks) by attending to graspable parts. DynOPETs (WeI2I.33; arXiv 2503.19625) is the supporting benchmark: 175 object instances under simultaneous camera and object motion with synchronized 6-DoF annotations, evaluating 18 methods β€” squarely targeting the moving-camera/moving-object regime that static pose datasets ignore.

2. Category-level & novel-object pose / detection

Where instance-level pose assumes a CAD model, this thread drops that assumption. Category-Level Object Shape and Pose Estimation in Less Than a Millisecond (TuI2I.250; arXiv 2509.18979, MIT-SPARK) is the standout: a self-consistent-field solver over a linear active-shape model that runs one iteration in ~100 microseconds with a certificate of global optimality β€” fast enough to be used as an outlier-rejection inner loop. GFreeDet2 (WeAT3.1; building on GFreeDet, arXiv 2412.01552, BOP-Challenge-2024 winner) goes fully model-free: reconstruct 3D Gaussian object models from multi-view RGB references, then do 2D + 6D detection of unseen objects via SAM/DINOv2 mask matching. PIRATR (WeI1I.214) does parametric 3D object inference with transformers directly in point clouds, jointly estimating multi-class 6-DoF poses and class-specific parameters. Plug-And-Play Shape Matching (TuI1I.353) refines grasps on unknown objects with a mesh-free, training-free geometric module.

3. 3D / point-cloud / 4D representations for manipulation

The largest cluster: what representation should feed a policy. FP3 (TuAT1.3; arXiv 2503.08950) is a 3D foundation policy β€” a DiT pre-trained on 60k point-cloud trajectories with a Uni3D encoder, learning new tasks at >90% success from only 80 demos. GP3 (ThI1I.52) builds a geometry-aware policy from multi-view images; VO-DP (WeI1I.252) argues for vision-only semantic-geometric features as an alternative to point clouds; 3D Dynamics-Aware Manipulation (TuI2I.323) adds explicit depth-wise 3D foresight to world-model policies. Gaussian splatting recurs as the 3D substrate: GAF (ThI1I.185; arXiv 2506.14135) makes a Gaussian Action Field β€” a 4D representation that jointly reconstructs the current scene, predicts future frames, and estimates action via Gaussian motion (a "Vision-to-4D-to-Action" paradigm); Informative Object-Centric Next Best View (ThI2I.51) and Real-To-Sim with Gaussian Splatting of Soft-Body Interactions (ThI2I.288) use 3DGS for active sensing and policy evaluation respectively. Meshes and latent maps also appear: Subsecond 3D Mesh Generation (WeI2I.230; arXiv 2512.24428) produces a manipulation-ready mesh from a single RGB-D image in <1s (92% pick-and-place success), and Seeing the Bigger Picture (WeI1I.89) shows a 3D latent map beats image-only policies for mobile manipulation.

4. Affordance detection & grounding

Affordances bridge perception and action β€” where and how to interact. Coupled Particle Filters for Robust Affordance Estimation (TuI1I.225) disambiguates graspable vs. movable regions with two coupled recursive estimators. Visual Category-Guided One-Shot Open Affordance Grounding (TuI2I.345) leverages visual foundation models for one-shot open-vocabulary grounding. RoboPCA (WeI2I.198) learns pose-centered affordances (contact regions + contact poses) from human demos. NaturalVLM (WeI2I.5) and T-FunS3D (ThI2I.137) push toward language-conditioned/3D functionality grounding β€” the latter doing task-driven hierarchical open-vocabulary 3D functionality segmentation. AdapGrasp (TuI1I.86) pairs a stiffness-and-affordance dataset with a transformer grasp model, while RoboHitch (TuI1I.97) learns visual affordance from disordered keypoints for knot tying.

5. Articulated-object perception

Understanding kinematic structure is a perception problem in its own right. PokeNet (TuI1I.242) learns joint parameters of articulated objects from human observations without CAD priors. Kinematify (ThI1I.333; arXiv 2511.01294) synthesizes high-DoF articulated objects (exported as URDF) from a single RGB image or text via MCTS structural inference plus geometry-driven joint optimization β€” no motion data required. UniDoorManip (ThI1I.330; arXiv 2403.02604) builds a large-scale door environment (6 categories, thousands of instances) and learns a universal door policy from partial/occluded point clouds, a canonical articulated-perception-for-control setting.

6. Transparent / specular & deformable perception (the hard cases)

RGB-D's failure modes get dedicated treatment. Diffusion Knows Transparency (TuI2I.90) repurposes a video-diffusion prior for transparent-object depth and normals where stereo/ToF break. TORM (WeI1I.362) reconstructs and manipulates transparent objects via multi-view segmentation; SilRef (ThI2I.406) jointly optimizes visual silhouette and tactile pose for transparent manipulation; SPILL (ThI1I.383) estimates size, pose, and internal liquid level of transparent glassware for bartending. Liquid recurs in Toward Multimodal Liquid-Level Estimation (TuI2LB.16). Deformables form a parallel hard-case cluster: CloSE (TuI2I.219; arXiv 2504.05033) is a compact shape-/orientation-agnostic cloth-state representation built on a topological dGLI disk; CVF-DLO (WeI1I.16) and Interactive Robotic Moving Cable Segmentation (ThI2I.364) tackle tangled/branched cable routing and motion-correlation segmentation; Uncertainty-Aware Stereo Grasp Point Selection (TuI1LB.8) adds prediction-reliability to cable grasping.

7. Visual servoing & active/interactive perception

Closed-loop perception-for-control. Perception-Control Coupled Visual Servoing for Textureless Objects (TuI1I.68) uses a keypoint-based EKF to servo on feature-poor surfaces; An Autonomous Hardware-Agnostic Vision-Servoed System (TuI1I.325) servos nanoliter microdevice injection without custom calibration. Active/interactive perception threads through RUMI (ThI2I.374; arXiv 2408.10450), which plans contact-rich rummaging via mutual information between object-pose belief and trajectory in occluded bins; COMPASS (TuI1I.294), active sensing in confined spaces; Active-Perceptive Language-Oriented Grasp (TuI1I.431) for heavily cluttered scenes; and SceneComplete (TuAT3.5; arXiv 2410.23643), which composes pretrained perception modules into a full segmented 3D scene from a single view for robust grasping.

8. Keypoint / dense-correspondence & flow

Correspondence as the perceptual interface to action. Sparse Meets Dense (WeI1I.228) fuses sparse and dense correspondence for rigid-deformable interactions (hanging clothes, dressing). NovaFlow (ThI1I.134; arXiv 2510.08568) distills generated videos into 3D actionable object flow, transferring zero-shot across a Franka and a Spot; Dream2Flow (TuI2I.123; arXiv 2512.24766) likewise reconstructs 3D object flow from generated video as an embodiment-agnostic interface for open-world manipulation; Actron3D (ThI2I.114) learns actionable neural functions from uncalibrated RGB-only human video for transferable 6-DoF skills.

9. Foundation-model perception (SAM / DINO / MLLM / video-diffusion) for manipulation

A cross-cutting enabler rather than a single application. SAM/DINO underpin GFreeDet2 (WeAT3.1) and Subsecond 3D Mesh Generation (Florence-2 + SAM2 + Depth-Anything-v2). MLLMs are distilled into 3D for grasping in Point2Act (WeI1I.92; arXiv 2508.03099), which builds 3D relevancy fields and produces a grasp in <20s. Video-diffusion priors drive Diffusion Knows Transparency (TuI2I.90), NovaFlow, and Dream2Flow. VERM (ThI1I.407) leverages foundation models to construct a "virtual eye" that prunes multi-camera redundancy for 3D manipulation, and Clutt3R-Seg (WeI2I.181) does sparse-view 3D instance segmentation for language-grounded grasping.

Standout deep-dives

These papers were cross-checked against an authoritative source (arXiv + project/GitHub) for both identity and a concrete number.

  • Category-Level Object Shape & Pose in <1 ms (TuI2I.250) β€” arXiv 2509.18979 (MIT-SPARK). Self-consistent-field iteration over a linear active-shape model; ~100 Β΅s per iteration, with a global-optimality certificate. Reframes category-level pose as a tiny eigenproblem fast enough to use for outlier rejection. Code.
  • FP3: A 3D Foundation Policy (TuAT1.3) β€” arXiv 2503.08950. DiT pre-trained on 60k point-cloud trajectories (Uni3D encoder); fine-tunes to new tasks at >90% success from 80 demos in novel environments with unseen objects. A 3D answer to image-only foundation policies.
  • SE(3)-PoseFlow (WeI1I.194) β€” arXiv 2511.01501. Flow-matching on SE(3) yields a full pose distribution (not a point estimate); SOTA on Real275, YCB-V, LM-O, with downstream active-perception and uncertainty-aware grasp synthesis.
  • GAF: Gaussian Action Field (ThI1I.185) β€” arXiv 2506.14135. 4D "Vision-to-4D-to-Action" representation extending 3DGS with learnable motion; +15.7% success over Diffusion Policy and +7.3% over Act3D, real-time on a single GPU.
  • NovaFlow (ThI1I.134) β€” arXiv 2510.08568 (RAI Institute / Brown). Synthesizes a video from language, distills 3D actionable object flow, and executes zero-shot across rigid/articulated/deformable objects on a Franka and a Spot β€” no demonstrations, no embodiment-matched data.
  • GFreeDet2 (WeAT3.1) β€” built on GFreeDet, arXiv 2412.01552, best-overall + best-fast in the model-free 2D-detection track of BOP Challenge 2024. Gaussian-splatting object models from RGB references + SAM/DINOv2 enable model-free 2D+6D detection of unseen objects.
  • Subsecond 3D Mesh Generation (WeI2I.230) β€” arXiv 2512.24428 (Yale). Single RGB-D image β†’ manipulation-ready mesh in <1s (Florence-2 + SAM2 + Depth-Anything-v2 + FlashVDM-distilled Hunyuan3D 2.0); 92% pick-and-place success.

Complete paper list (82)

Code Title arXiv
ThI1I.114 Latent Representations for Visual Proprioception in Inexpensive Robots 2504.14634
ThI1I.134 NovaFlow: Zero-Shot Manipulation Via Actionable Flow from Generated Videos 2510.08568
ThI1I.185 GAF: Gaussian Action Field As a 4D Representation for Dynamic World Modeling in Robotic Manipulation 2506.14135
ThI1I.187 Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos 2505.18899
ThI1I.195 Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery 2603.03181
ThI1I.198 OHMM-PA: A Learning from Demonstration Approach Using Online Hidden Markov Models with Path Planning β€”
ThI1I.212 CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion 2512.16023
ThI1I.25 Haptic Stiffness Perception Using Hand Exoskeletons in Tactile Robotic Telemanipulation 2412.02613
ThI1I.330 UniDoorManip: Learning Universal Door Manipulation Policy Over Large-Scale and Diverse Door Manipulation Environments 2403.02604
ThI1I.333 Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects 2511.01294
ThI1I.383 SPILL: Size, Pose, and Internal Liquid Level Estimation of Transparent Glassware for Robotic Bartending β€”
ThI1I.407 VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation 2512.16724
ThI1I.52 GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation 2509.15733
ThI1I.99 Mash, Spread, Slice! Learning to Manipulate Object States Via Visual Spatial Progress 2509.24129
ThI2I.114 Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation 2510.12971
ThI2I.121 Improving Robotic Manipulation Robustness Via NICE Scene Surgery 2511.22777
ThI2I.137 T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation β€”
ThI2I.182 Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion 2605.23847
ThI2I.197 OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-To-Robot Action Transfer 2603.14401
ThI2I.258 VistaBot: View-Robust Robot Manipulation Via Spatiotemporal-Aware View Synthesis 2604.21914
ThI2I.288 Real-To-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions 2511.04665
ThI2I.337 PartPose: Attentive 6D Pose Estimation by Focusing on Graspable Parts of Multi-Part Deformable Objects β€”
ThI2I.364 Interactive Robotic Moving Cable Segmentation by Motion Correlation β€”
ThI2I.374 RUMI: Rummaging Using Mutual Information 2408.10450
ThI2I.406 SilRef: Joint Visual Silhouette and Tactile Pose Optimization for Transparent Object Manipulation β€”
ThI2I.51 Informative Object-Centric Next Best View for Object-Aware 3D Gaussian Splatting in Cluttered Scenes β€”
TuAT1.3 FP3: A 3D Foundation Policy for Robotic Manipulation 2503.08950
TuAT3.5 SceneComplete: Open-World 3D Scene Completion in Cluttered Real World Environments for Robot Manipulation 2410.23643
TuAT3.6 Robust Bayesian Scene Reconstruction with Retrieval-Augmented Priors for Precise Grasping and Planning 2411.19461
TuI1I.12 Distributional Treatment of Real2Sim2Real for Object-Centric Agent Adaptation in Vision-Driven DLO Manipulation 2502.18615
TuI1I.225 Coupled Particle Filters for Robust Affordance Estimation 2603.15223
TuI1I.242 PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations 2602.02741
TuI1I.294 COMPASS: Confined-Space Manipulation Planning with Active Sensing Strategy 2509.14787
TuI1I.325 An Autonomous and Hardware-Agnostic Vision-Servoed System for Microdevice Injection β€”
TuI1I.353 Plug-And-Play Shape Matching Module for Zero-Shot Mesh-Free Grasp Refinement on Unknown Objects β€”
TuI1I.364 Fine-Grained Classification for Depth Estimation from Monocular Microscopy for Robotic Micromanipulation of Motile Cells β€”
TuI1I.397 Learning 6D Object Pose Estimation with Event Cameras Using Synthetic Data and Domain Randomization β€”
TuI1I.431 Active-Perceptive Language-Oriented Grasp Policy for Heavily Cluttered Scenes β€”
TuI1I.68 Perception-Control Coupled Visual Servoing for Textureless Objects Using Keypoint-Based EKF 2602.06834
TuI1I.86 AdapGrasp: A Stiffness and Grasp Affordance Dataset with a Transformer-Based Adaptive Grasp Model β€”
TuI1I.97 RoboHitch: Learning Visual Affordance from Disordered Keypoints for Hitch Knots Tying 2605.24394
TuI1LB.8 Uncertainty-Aware Stereo Grasp Point Selection for Deformable Linear Objects β€”
TuI2I.123 Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow 2512.24766
TuI2I.198 Tactile Memory for Continuous Policy Blending in Unified Force-Impedance Control β€”
TuI2I.219 CloSE: A Geometric Shape-Agnostic Cloth State Representation 2504.05033
TuI2I.250 Category-Level Object Shape and Pose Estimation in Less Than a Millisecond 2509.18979
TuI2I.316 MGS-Track: Monocular 6DoF Pose Tracking Via Masked 3D Prior and Online Gaussian Splatting β€”
TuI2I.323 3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight 2502.10028
TuI2I.336 The Price Is Not Right: Neuro-Symbolic Methods Outperform VLAs on Structured Long-Horizon Manipulation Tasks with Significantly Lower Energy Consumption 2602.19260
TuI2I.345 Visual Category-Guided One-Shot Open Affordance Grounding β€”
TuI2I.364 ILeSiA: Interactive Learning of Robot Situational Awareness from Camera Input 2409.20173
TuI2I.436 Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies towards Visual Robustness 2505.08627
TuI2I.90 Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation 2512.23705
TuI2LB.16 Toward Multimodal Liquid-Level Estimation for Closed-Loop Robotic Pouring β€”
WeAT3.1 GFreeDet2: Exploiting Gaussian Splatting and Foundation Models for RGB-Based Model-Free 2D and 6D Detection of Unseen Objects 2412.01552
WeI1I.131 Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation 2509.17125
WeI1I.159 Bi-Manual Joint Camera Calibration and Scene Representation 2505.24819
WeI1I.16 CVF-DLO: Cross-Visual-Field Branched Deformable Linear Objects Route Estimation β€”
WeI1I.194 SE(3)-PoseFlow: Estimating 6D Pose Distributions for Uncertainty-Aware Robotic Manipulation 2511.01501
WeI1I.214 PIRATR: Parametric Object Inference for Robotic Applications with Transformers in 3D Point Clouds 2602.05557
WeI1I.228 Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions β€”
WeI1I.252 VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation 2510.15530
WeI1I.268 Ego-Vision World Model for Humanoid Contact Planning 2510.11682
WeI1I.271 GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning 2602.04231
WeI1I.362 TORM: Transparent Objects Reconstruction and Manipulation with Multi-View Segmentation β€”
WeI1I.382 Fixture-Free Automated Sewing System Using Dual-Arm Manipulator and High-Speed Fabric Edge Detection β€”
WeI1I.88 EdgeGrasp: Enhancing Edge Perception for 7-DoF Grasping Pose Estimation in Cluttered Scenes β€”
WeI1I.89 Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning 2510.03885
WeI1I.92 Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping 2508.03099
WeI2I.131 Learning to Grasp by Integrating Human Preferences and Success Feedback β€”
WeI2I.181 Clutt3R-Seg: Sparse-View 3D Instance Segmentation for Language-Grounded Grasping in Cluttered Scenes 2602.11660
WeI2I.196 Visual-Auditory Extrinsic Contact Estimation 2409.14608
WeI2I.198 RoboPCA: Pose-Centered Affordance Learning from Human Demonstrations for Robot Manipulation 2603.07691
WeI2I.223 From Swept Contact to Pose: Probe-Aware Registration Via Complementary-Shape Docking β€”
WeI2I.230 Subsecond 3D Mesh Generation for Robot Manipulation 2512.24428
WeI2I.254 Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets 2601.09605
WeI2I.289 Beyond the Patch: Exploring Vulnerabilities of Visuomotor Policies Via Viewpoint-Consistent 3D Adversarial Object 2603.04913
WeI2I.312 CAVER: Curious AudioVisual Exploring Robot 2511.07619
WeI2I.327 GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-Trained Robot Policy Enhancement 2511.03400
WeI2I.33 DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios 2503.19625
WeI2I.331 OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics 2509.07500
WeI2I.5 NaturalVLM: Leveraging Fine-Grained Natural Language for Affordance-Guided Visual Manipulation 2403.08355

Related

← Back to ICRA-2026-VLA-Manipulation-Survey

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally