-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026
International Conference on Machine Learning 2026 β COEX Convention & Exhibition Center, Seoul, South Korea, July 6β11, 2026 (Jul 6 tutorials/expo Β· Jul 7β9 main conference Β· Jul 10β11 workshops).
Scale: 23,918 submissions β ~6,352 accepted (β26.6%) per the public submission statistics; the official virtual proceedings list 6,636 distinct accepted papers (6,804 schedule events, of which 168 are Orals and the rest Posters). Spotlight/Oral tier β the top 0.7% of submissions.
Robotics / manipulation share: a keyword sweep of the official accepted list surfaced ~220 robotics / embodied-AI candidates, of which 99 are genuinely about robot manipulation (the rest are autonomous-driving VLAs, pure navigation, world-model/RL learning theory, or non-robotic uses of "manipulation"). That makes manipulation β 1.5% of all ICML 2026 papers β smaller than at CVPR/ICLR by share, but the largest manipulation presence ICML has ever had, and notably ML-methodology-heavy (systematic studies, recipes, theory) rather than systems-paper-heavy.
Source & method. This index was built against the official virtual data file
https://icml.cc/static/virtual/data/icml-2026-orals-posters.json(2026-06 snapshot, 6,804 events). Every paper below was filtered by title+keyword match, then each abstract was fetched from itsicml.cc/virtual/2026/poster/<id>page and summarized individually. Quantitative figures in back-ticks are quoted from the paper's own abstract; where an abstract states no number, none is shown (no estimated/invented metrics). OpenReview blocks guest access to ICML 2026 notes, so author institutions are not yet attached. Treat one-line summaries as abstract-derived, pending full-text review.
- Per-paper in-depth pages now exist for all 99 manipulation papers (linked from the category list below) β each with Problem Β· Method Β· Results Β· Significance and, where the paper has a preprint, the paper's own figures embedded. A few highest-value papers also have long-form reviews: Review-RoboMME and Action Space: EEF vs Joint.
- XR-1 β Unified Vision-Motion Codes (dual-branch VQ-VAE), 12,000+ real rollouts across 6 embodiments. (also tagged ICLR 2026 in our wiki β verify which venue is canonical)
- From Pixels to Tokens β systematic study of latent-action supervision for VLAs; discrete latent action tokens win.
- From Abstraction to Instantiation (BehaviorVLA) β causal Mamba behavior encoder + phase-conditioned decoder for robustness under shift.
- Pretrained VLAs are Surprisingly Resistant to Forgetting in Continual Learning β big pretrained VLAs barely forget; simple Experience Replay can hit zero forgetting.
- RoboMME β standardized benchmark for memory in robotic generalist policies (16 tasks, 14 memory-augmented Ο0.5 variants).
Synthesis of editorial judgment + quantitative signals (GitHub β + arXiv-preprint citations via Semantic Scholar + Oral status), snapshot 2026-06-09. Pure simulators / data-generators are excluded (RoboTwin 2.0, VLA-Arena, CaP-X, OXE-AugE, SoMA, DLO-Lab, FlatLab, ManiSoft, SafeLab, AIR-VLA) β but insight / evaluation papers are kept (RoboMME, From Pixels to Tokens, Demystifying Action Space, Pretrained-VLA-Forgetting). β /citation are early preprint-era signals that favor early code-releasers and undercount work with no public repo/arXiv.
| # | Paper | Category | β | cite | Oral | Affiliations |
|---|---|---|---|---|---|---|
| 1 | RDT2 ΒΉ | Foundation/scaling | 775 | 15 | Tsinghua University (THU-ML) | |
| 2 | DreamDojo | World model | 923 | 42 | NVIDIA; HKUST; UC Berkeley; UW; Stanford; KAIST | |
| 3 | XR-1 ΒΉ | Architecture | 174 | 11 | β | Beijing Innovation Center of Humanoid Robotics (X-Humanoid); Beihang; PKU |
| 4 | Discrete Diffusion VLA ΒΉ | Architecture | 65 | 64 | HKU; Shanghai AI Lab; SJTU; Huawei | |
| 5 | Being-H0 ΒΉ | Human-video pretrain | 48 | 69 | Peking University; Renmin University; BeingBeyond | |
| 6 | DexMachina | Dexterous | 225 | 31 | Stanford University; NVIDIA | |
| 7 | VLAC (Progress Critic) | RL for VLA | 306 | β | Shanghai AI Laboratory (InternRobotics) | |
| 8 | Latent Reasoning VLA | Reasoning | 67 | 7 | Tsinghua; PKU; USTC | |
| 9 | LangForce | Analysis/method | 65 | 10 | ZGC-EmbodyAI (Zhongguancun Academy) | |
| 10 | HALO | Reasoning | β | 2 | HKUST | |
| 11 | Dual-Stream Diffusion | World model | β | 13 | KAIST (RLWRLD) | |
| 12 | DECO | Dexterous/tactile | 28 | 0 | BAAI; TU Munich | |
| 13 | See What Matters | Efficiency | 62 | 1 | University of Sydney | |
| 14 | SpecPrune-VLA | Efficiency | β | 24 | SJTU; Infinigence-AI; Shanghai Innovation Institute | |
| 15 | BehaviorVLA | Representation | β | 0 | β | HIT (Shenzhen); Sun Yat-sen University |
| 16 | VLANeXt | Recipe/insight | 196 | 5 | S-Lab NTU; SYSU; ACE Robotics | |
| 17 | From Pixels to Tokens | Insight | 28 | 0 | β | Renmin University (KBReasoning) |
| 18 | Pretrained VLAs Resist Forgetting | Insight | β | 4 | β | UT Austin; KAIST; Microsoft |
| 19 | Demystifying Action Space | Insight | β | 1 | Tsinghua (IAIR/Wuxi) | |
| 20 | RoboMME | Memory insight | 111 | 6 | β | University of Michigan; Stanford; Figure AI |
ΒΉ Already has a wiki page under another 2026 venue (CVPR/ICLR); these titles are also in the official ICML 2026 accepted list, so venue-of-record needs reconciliation β link points to the existing page.
How to read it: the top tier (RDT2, DreamDojo, XR-1, DexMachina) is strong on both axes. Citations elevate Discrete Diffusion VLA (64) and Being-H0 (69) β the most-cited methods. Editorial judgment retains low-traction-but-important work: the Oral insight papers (From Pixels to Tokens, Pretrained-VLA-Forgetting, RoboMME) and the latent-reasoning / efficiency methods. 7 of 20 have industry involvement (NVIDIA Γ2, X-Humanoid, BeingBeyond, Figure AI, Huawei, Infinigence-AI).
Each paper links to a detail page (Problem Β· Method Β· Results Β· Significance Β· Links); page metrics use the same 2026-06-09 snapshot.
99 papers. Headline numbers in back-ticks are quoted from the abstract.
- From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models (Oral) β A systematic comparison of image-based versus action-based latent-action supervision strategies for VLAs under a unified baseline, finding that directly supervising the VLM with discrete latent action tokens is most effective.
-
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations (Oral) β A VLA model that learns Unified Vision-Motion Codes via a dual-branch VQ-VAE jointly encoding visual dynamics and robotic motion, trained in three stages to bridge cross-embodiment and human-demonstration domain gaps. β
validated through over 12,000 real-world rollouts across six robot embodiments and 120+ manipulation tasks - Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds β Introduces Any3D-VLA, which merges simulator-generated, sensor-derived, and model-estimated point clouds into a unified training framework to learn domain-agnostic 3D representations fused with 2D inputs for more robust VLA spatial understanding.
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting β Introduces Visual Attentive Prompting (VAP), a training-free perceptual adapter that grounds personalized objects via open-vocabulary detection and injects visual prompts to give frozen VLAs top-down selective attention for personalized manipulation commands.
-
Contrastive Representation Regularization for Vision-Language-Action Models β RS-CL is a robot-state-aware contrastive representation regularizer that aligns VLA representations with proprioceptive states using inter-state distances as soft supervision to improve manipulation performance. β
69.7% on RoboCasa-Kitchen; real-robot success 45.0% to 58.3% -
Demystifying Action Space Design for Robotic Manipulation Policies β A large-scale empirical study dissecting action space design along temporal and spatial axes (absolute vs. delta, joint-space vs. task-space) for imitation-based manipulation policies and its effect on learnability and control stability. β
Based on 13,000+ real-world rollouts on a bimanual robot and evaluation on 500+ trained models over four scenarios - EnsembleVLA: Ensemble Learning for Vision-Language Action Models β EnsembleVLA is an energy-based framework that formulates diffusion- and flow-based VLA models as energy-based models to compose multiple pretrained policies with learnable weights and confidence-aware gating.
-
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation β Proposes FOCA, a VLA adaptation method that predicts future interaction embeddings and aligns to future goal observations (with optional co-training on world-model synthetic videos) for data-efficient few-shot imitation. β
95.7% success with 20 demonstrations on LIBERO - Fourier Features Let Agents Learn High Precision Policies with Imitation Learning β The paper maps point clouds into high-dimensional Fourier space via a parametric projection to overcome neural networks' spectral bias, improving high-precision point-cloud imitation learning.
- GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation Policies β Introduces GeoMoLa, which learns discrete motion latent codes by predicting how point clouds evolve during manipulation (a 4D spatial-temporal objective) rather than reconstructing visual observations, using only single-view RGB-D input.
-
LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries β LangForce uses a dual-branch Bayesian decomposition with learnable Latent Action Queries to maximize conditional pointwise mutual information between actions and instructions, countering the vision shortcut in VLA training. β
11.3% improvement on the OOD SimplerEnv benchmark -
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs β Task-Agnostic Pretraining (TAP) pretrains a VLA on cheap off-task/play trajectories via an inverse-dynamics objective, then aligns the learned physical priors with language using minimal expert data. β
25% vs. 0% under camera perturbations -
Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation β A VLA framework that decouples manipulation into coarse 'move' and contact-critical 'operate' phases using a dual-expert policy routed by a learnable phase selector with MLLM-generated phase labels. β
average success rate of 68.9%, outperforming the monolithic Pi0 baseline by +24% - N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout β A mobile-manipulation base-positioning method (N2M) that learns policy-aware, viewpoint-invariant pose preferences from policy rollouts using a viewpoint augmentation strategy, avoiding pre-built scene reconstruction.
- NeurVLA: Unleashing Failure-Handling Capability of Vision-Language-Action Models via Neural-Symbolic Reasoning β NeurVLA is a neural-symbolic framework that jointly handles failure correction and prevention via reasoning and internalizes these failure-handling capabilities into VLA models for robotic manipulation.
-
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models β Reformulates VLA action prediction from discrete waypoints to continuous-time action function regression using an MLLM as a spectral modulator over a learnable motion prior, enabling differentiable velocity/acceleration/jerk supervision. β
achieves state-of-the-art results on CALVIN and LIBERO benchmarks - RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation β A retrieval-augmented VLA framework for training-free test-time adaptation that combines behavior-aligned context retrieval with a grounded execution pipeline to overcome in-context imitation learning's adaptation bottleneck.
-
RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization β RDT2 is a 7B-parameter robotic foundation model trained on over 10,000 hours of UMI-collected data with a three-stage RVQ/flow-matching/distillation recipe for zero-shot cross-embodiment manipulation. β
over 10,000 hours of demonstrations - Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action β A spatial memory framework (SOMA) that equips VLA models with persistent spatial-semantic memory from multi-view head-camera observations, enabling manipulation of objects outside the current camera view.
-
UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning β UniCoD learns a unified continuous-and-discrete representation by pretraining on over 1M internet manipulation videos and fine-tuning on robot data to map predictive representations to action tokens. β
9% and 12% over baselines in sim and real-world OOD tasks -
VLANeXt: Recipes for Building Strong VLA Models β A unified empirical study dissecting VLA design choices across foundational components, perception, and action modeling, distilling 12 findings into a recipe and a model (VLANeXt). β
outperforms prior state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks
-
Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment β Provides a systematic model-hardware co-characterization framework for low-cost VLA edge deployment, identifying a two-phase (compute-bound VLM, memory-bound action expert) inference pattern and proposing DP-Cache and V-AEFusion to reduce diffusion redundancy and enable asynchronous pipeline parallelism. β
up to 2.9x speedup on GPUs and 3.3x on edge NPUs with only marginal success degradation -
EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models β EcoVLA is a training-free adaptive channel-pruning framework with environment-aware sparsity updates and interleaved inference scheduling that accelerates VLA inference with minimal accuracy loss. β
up to 1.60x speedup with only 0.4% drop in success rate (2.18x combined with token pruning) -
Reflex: Real-Time Vision-Language-Action Control through Streaming Inference β Reflex enables real-time streaming inference for flow-matching VLAs by exploiting timestep-invariance to partition attention into static/sliding/dynamic regions for O(1) cache updates, plus AdaRMSNorm and an async pipeline. β
achieves a 2.58x inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% -
STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction β A lightweight spatiotemporal consistency prediction mechanism that constructs high-quality warm-start actions plus a velocity-aware perturbation injection scheme to accelerate diffusion-policy inference without sacrificing action quality. β
STEP with 2 steps can achieve an average 21.6% and 27.5% higher success rate than BRIDGER and DDIM on the RoboMimic benchmark and real-world tasks, respectively -
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model β Proposes the Differentiable Grid Sampler (GridS), a plug-and-play module that performs task-aware continuous resampling of visual tokens via differentiable interpolation to drastically compress VLA visual tokens while preserving spatial information. β
76% reduction in FLOPs with no degradation in the success rate (fewer than 10% original visual tokens) -
Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning β Accelerates diffusion-policy action generation with a rollout-adaptive, observation-conditioned prune-then-reuse mechanism that caches and substitutes redundant computations across timesteps and blocks. β
achieves up to 4x generation speedup without sacrificing performance -
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning β SpecPrune-VLA is a training-free two-level (action-level static and layer-level dynamic) token pruning method with an action-aware controller that accelerates VLA inference using global context plus local attention. β
1.57x speedup in LIBERO simulation and 1.70x on real-world tasks -
Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation β Proposes a policy-agnostic offline-RL scheduler that downsamples action chunks under a Constrained MDP, using a learned world model as a safety-constraint surrogate to accelerate embodied policies without retraining. β
approximately 1.8x execution speedup across diverse policies while preserving original success rates - Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models β Proposes AVA-VLA, a latent-reasoning VLA that models reasoning as unobservable latent variables refined by an RL-based denoising mechanism and adaptively terminated via a confidence-based early-exit strategy to cut inference latency.
- Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation β A skill-reasoning in-context-learning framework that decomposes seen demonstrations into atomic skill-action pairs and recomposes them for unseen tasks via compositional reasoning, using task-adaptive dynamic and coverage-aware static demonstration libraries.
-
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning β A unified VLA model using a Mixture-of-Transformers architecture to perform embodied multimodal chain-of-thought reasoning by sequentially combining textual task reasoning, visual subgoal prediction, and action prediction. β
surpassing baseline policy Pi0 by 34.1% on RoboTwin benchmark -
LaST_0: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model β Proposes a VLA framework that reasons before acting in a token-efficient latent spatio-temporal chain-of-thought space (future visual dynamics, 3D structure, proprioception), with a Mixture-of-Transformers dual-system design separating low-frequency reasoning and high-frequency action experts. β
improves mean success rates by 13%, 14% and 14% over prior SOTA VLA methods -
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models β Internalizes multi-modal Chain-of-Thought reasoning into continuous latent representations via a curriculum that transitions from explicit textual/visual CoT supervision to latent reasoning, removing explicit CoT generation at inference for efficient action control. β
up to a 90% reduction in inference latency compared to explicit CoT-based VLA approaches - SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models β Proposes a training-free, verifier-free single-forward-pass inference strategy that jointly modulates visual perception and action based on self-uncertainty, exploring more when uncertain and exploiting when confident.
-
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery β Introduces Sentinel-VLA, a metacognitive VLA with an active sentinel module for real-time status monitoring that triggers on-demand reasoning or error recovery, plus a self-evolving continual learning algorithm and orthogonal adapter to prevent forgetting. β
boosts the task success rate by over 30% compared to the SOTA model, PI0 - TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation β Proposes a policy-agnostic inference-time sampling framework that draws candidate actions from an Action-VAE latent space and selects among them with a verifier trained to predict task-progress outcomes from sequential robot data.
- Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models β Formulates sequential calibration for episodic VLA tasks via a sequential Brier score whose risk minimizer equals the policy's value function, enabling temporal-difference value estimation as a principled calibration mechanism over partial trajectories.
-
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model β Proposes a VLA framework with uncertainty-triggered adaptive test-time compute, using a Relative Action Critic that selects the best candidate action via pairwise comparisons instead of unstable absolute value estimation. β
On LIBERO-LONG, reduces the failure rate of the SOTA model PI0.5 by over 50%
-
Cross-Embodiment Robot Foundation World Models with Latent Actions β A Latent Action Conditioned Robot World Model (LAC-WM) that operates in a learned unified latent action space shared across embodiments, enabling better adaptation to unseen robots than explicit-action-conditioned baselines. β
LAC-WM achieves up to a 46.7% improvement in performance over EAC-WM -
DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos β DreamDojo is a foundation robot world model pretrained on 44,000 hours of egocentric human video using continuous latent actions, distilled to run in real time after robot fine-tuning. β
real-time performance at 10.93 FPS -
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model β A world-model-augmented VLA (DUST) using a multimodal diffusion transformer with separate modality streams, decoupled flow-matching loss, and asynchronous action/vision sampling to jointly predict states and actions. β
up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement; real-world Franka Research 3 outperforms baselines by 10% in success rate - From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation β MoLA converts imagined future videos into executable actions by using multiple modality-aware pretrained inverse dynamics models (semantic, depth, flow) to infer a mixture of latent actions bridging video imagination and policy execution.
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation β MVISTA-4D is an embodied 4D world model that generates geometrically consistent arbitrary-view RGBD from a single-view input and infers actions via test-time trajectory-latent optimization plus a residual inverse dynamics model.
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic Mask β The Mask World Model predicts the evolution of semantic masks instead of RGB pixels using video diffusion, imposing a geometric information bottleneck, and integrates this with a diffusion policy head for robust end-to-end control.
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation β Introduces a lightweight end-to-end flow world model that predicts multi-frame 3D flows from images and instructions to guide action generation via slow-fast collaboration for real-time, resource-efficient manipulation.
- Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World Model β Proposes dWorldEval, an action-centric discrete-diffusion world model that treats actions as first-class tokens in a unified token space with sparse keyframe memory and progress-as-text, enabling reliable automatic policy evaluation.
-
SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation β SoMA is a 3D Gaussian-Splat neural simulator that couples deformable dynamics, environmental forces, and robot joint actions in a unified latent space for end-to-end real-to-sim soft-body manipulation. β
improves resimulation accuracy and generalization by 20% - Structured 4D Latent World Model for Robot Planning β A world model that predicts the evolution of a scene's 3D structure in a structured latent space conditioned on observations and text, decodable into 3D formats and used as a planner whose futures are converted to actions by a goal-conditioned inverse dynamics module.
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models β Proposes VLA-MBPO, a model-based RL framework for VLA finetuning that adapts unified multimodal models for data-efficient world modeling, uses interleaved view decoding for multi-view consistency, and chunk-level branched rollout to mitigate compounding errors.
-
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model β VLAW iteratively improves a VLA policy and an action-conditioned video world model together, using real rollouts to refine the world model which then generates synthetic data to improve the policy. β
39.2% absolute success rate improvement over the base policy
- DADP: Domain Adaptive Diffusion Policy β Proposes a diffusion policy that disentangles static domain representations from transient dynamics via lagged-context dynamical prediction and injects them through domain-aware adjustment of the diffusion prior and target for zero-shot adaptation to unseen transition dynamics.
-
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies β Discrete Diffusion VLA is a unified-transformer policy that decodes discretized action chunks via discrete diffusion inside the VLM backbone, with adaptive decoding order and re-masking for error correction. β
96.5% avg. success on LIBERO - FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy β Proposes FocalPolicy, a visuomotor policy combining frequency-optimized chunking with locally anchored flow matching and a foresight composite objective to improve inter-chunk coherence and long-horizon action smoothness.
- From Noise to Control: Parameterized Diffusion Policies β Parameterized Diffusion Policy learns a diffusion policy over a smooth continuous latent manifold where distances reflect trajectory similarity, enabling controllable interpolation and generalization to novel constraints without weight updates.
- From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges β Proposes ResVLA, which reframes action generation as refinement-from-intent by spectrally decomposing motion into deterministic low-frequency anchoring and stochastic high-frequency residuals refined via residual diffusion.
- OMP: One-step Meanflow Policy with Directional Alignment β OMP is a one-step MeanFlow manipulation policy that adds directional velocity alignment and a differential approximation of the JVP operator to enable high-fidelity, real-time single-step generation.
-
Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation β Proposes BridgePolicy, which integrates observations directly into the diffusion process via a diffusion-bridge formulation (with multi-modal fusion and a semantic aligner) so sampling starts from an informative observation-conditioned prior rather than random noise. β
52 simulation tasks on three benchmarks and 5 real-world tasks -
The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space β Introduces Lie Diffuser Actor, a diffusion VLA policy operating intrinsically on the SE(3) manifold via left-invariant SDEs, tangent-space score prediction, and exponential-map retraction to avoid manifold drift and guarantee equivariance and geodesic optimality. β
On CALVIN ABC->D, LDA improves average task length from 3.06 to 3.30 (+7.8%)
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model (Oral) β Proposes BehaviorVLA, which learns generalized behavior representations for VLA models using a causal Mamba-based Visuomotor Behavior Encoder and a Phase-conditioned Behavior Decoder to improve robustness under distribution shift.
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning (Oral) β An empirical study finding that large-scale pretrained VLA models resist catastrophic forgetting far better than small policies trained from scratch, with simple Experience Replay sometimes achieving zero forgetting at small replay sizes.
- A Generalist Pair-wise Progress Critic Model for Vision-Language-Action Robots β Proposes VLAC, a unified autoregressive vision-language action-critic model that predicts pair-wise task-progress deltas to provide dense intrinsic rewards and robust actions for real-world reinforcement learning.
- DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization β A two-stage RL optimization framework for VLAs that captures cross-task latent representations via information-theoretic principles and refines policy optimization through a mixture-of-RL-residuals to improve generalization.
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement Learning β Proposes Focus Then Contact (FTC), a lightweight method combining residual RL base actions with an affordance-guided reward to accelerate convergence of human-in-the-loop real-world RL for contact-rich manipulation tasks.
- HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control β A hierarchical embodied memory framework decoupling control into a high-frequency Executor, a Sentry for working memory, and a Planner for long-term strategy, with a dynamic cross-modal knowledge system supporting add/update/delete operations for long-horizon tasks.
-
LAGEA: Language Guided Embodied Agents for Robotic Manipulation β Proposes LAGEA, which converts natural-language failure reflections from vision-language models into temporally grounded, decaying shaped rewards for reinforcement learning in robotic manipulation. β
improvements of 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks -
LARA: Latent Action Representation Alignment for Vision-Language-Action Models β Jointly optimizes a latent action model and a VLA model via representation alignment, so action trajectories prevent spurious visual changes while forward-dynamics regularization reduces VLA hallucinations. β
approximately 10% improvement in pre-training scenarios, 5% enhancement for post-training, and 15% gains in latent action model refinement - ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation β ReLAM learns an anticipation model that proposes keypoint-based subgoals from action-free videos to automatically generate dense rewards for hierarchical RL on long-horizon visual manipulation.
- Scaling by Diversified Experience for Vision-Language-Action Models β Introduces SyVLA, which uses an Intention Decoupling algorithm to isolate control features from reasoning context and a similar-sample-guided RL pipeline to stabilize policy updates and improve out-of-distribution generalization.
-
SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models β Models skill attributes hierarchically using motion code and the VerbNet framework and regulates a mixture-of-experts structure with transferable skill embeddings as soft constraints to enable compositional generalization across tasks. β
achieves an improvement of performance by 16.0% and 23.9% [zero-shot and few-shot transfer] -
Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations β QUEST, a model-based RL-from-demonstrations framework that adaptively switches between exploration and exploitation guided by uncertainty, using intrinsic rewards, ensemble-dynamics planning, and hybrid sampling for sparse-reward manipulation. β
outperforms state-of-the-art methods by 17% on average, with gains increasing to 60% on difficult tasks
-
Can VLMs Diagnose and Recover from VLA Manipulation Faults? β Introduces the VLA-FixBench fault dataset and FaultEval framework to benchmark 20 VLMs on diagnosing perception/planning/control failures, plus a VLM-VLA collaboration mechanism that localizes deviations and rolls back execution for targeted recovery. β
an idealized feedback loop can improve task success rates by 13% on LIBERO and 35% on real-world robots -
Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts β A diagnostic benchmark (LIBERO-Gen) that restructures VLA evaluation into in-distribution, compositional, and domain-generalization tiers to expose spurious invariance and brittleness masked by standard metrics. β
identifies Pi0.5 as the top performer (64.0% in Spatial-CG; 21.2% in Task-CG) - Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning β CAPS is a training-free inference-time framework that uses power-distribution sampling and SNR-triggered adaptive MCMC search to mitigate instruction drift in long-horizon VLA manipulation.
- Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models β An interpretability study introducing the Interventional Significance Score and Nuisance Mass Ratio metrics to quantify causal influence of visual regions on VLA actions and predict generalization under distribution shift.
-
PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation β A post-training framework (PACT) that aligns pretrained diffusion policies with constraint-feasible regions by distilling constraint gradients via reverse-KL optimization with a constraint-tightening curriculum, without demonstrations or task rewards. β
reduces safety violations by 31.0% on average while improving task success by 30.7% -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data β Proposes an information-bottleneck adapter that filters visual noise to make VLA models robust to unseen real-world visual disturbances without extra data or augmentation, adding fewer than 10M parameters. β
IB-Adapter consistently improves over the baseline by an average of 30% - TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches β Proposes TRAP, the first targeted adversarial attack framework for CoT-reasoning VLA models, using a physical adversarial patch to corrupt intermediate chain-of-thought reasoning and hijack the robot's manipulation output toward adversary-defined behaviors.
- Cross-Tactile Sensor Representation Learning β CTSRL learns sensor-agnostic visuo-tactile representations via a Cross-Sensor Modulator and a two-stage synthetic-then-real self-supervised paradigm for generalization to unseen tactile sensors.
-
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter β Proposes DECO, a decoupled multimodal diffusion transformer that disentangles vision, proprioception, and tactile signals through specialized conditioning pathways with a lightweight tactile adapter, and releases the 50-hour DECO-50 bimanual dexterous tactile dataset. β
72.25% average success rate and a 21% improvement over the baseline - DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation β A curriculum-based algorithm for functional retargeting that learns bimanual dexterous policies to track object states from human hand-object demonstrations using virtual object controllers with decaying strength, plus a simulation benchmark.
- EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video β EgoTactile is a benchmark and diffusion-based method (EgoPressureDiff) that estimates full-hand grasp pressure for everyday objects from egocentric video using a pretrained video diffusion backbone.
- Joint-Space Empowerment as a Theory of Dexterous Motor Coordination β Introduces joint-space empowerment, an information-theoretic principle quantifying an agent's control over its body, used to discover low-dimensional high-empowerment action manifolds for overactuated musculoskeletal systems that improve dexterity and sample efficiency.
-
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation β A scalable closed-loop simulation framework that uses MLLMs with simulation-in-the-loop verification and five-axis domain randomization to generate diverse synthetic data and a unified evaluation benchmark for dual-arm manipulation. β
3.6x improvement in few-shot real-world transfer (over a 10-demo baseline) and a 2.2x gain in zero-shot generalization -
Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language β Tabero is a vision-tactile-language benchmark and VTLA model with a decoupled force-position command interface executed by a hybrid controller for gentle, force-aware language-conditioned manipulation. β
reduces average grip force by over 70% under gentle instructions -
Vision-Language-Action Pretraining from Large-Scale Human Videos β The paper proposes physical instruction tuning that pretrains a VLA from large-scale human-hand videos with perspective spatial alignment and part-level motion tokenization to transfer to dexterous robot manipulation. β
millimeter-level reconstruction accuracy
-
RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies (Oral) β A standardized benchmark for evaluating memory in VLA models on long-horizon, history-dependent manipulation, with a taxonomy of temporal/spatial/object/procedural memory and memory-augmented variants on a pi0.5 backbone. β
16 manipulation tasks ... 14 memory-augmented VLA variants built on the pi0.5 backbone -
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation β Introduces AIR-VLA, the first VLA benchmark tailored for aerial manipulation systems, featuring physics-based simulation and a dataset of 3,000 teleoperated demonstrations covering manipulation, object understanding, semantic reasoning, and planning. β
dataset of 3,000 teleoperated demonstrations -
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation β A framework (CaP-Gym + CaP-Bench) for benchmarking code-as-policy coding agents on robot manipulation, plus training-free (CaP-Agent0) and RL-with-verifiable-rewards (CaP-RL) methods that improve performance via test-time computation. β
evaluation across 12 models on 7 simulation tasks - DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics β Introduces a differentiable simulator and benchmark suite for deformable linear object manipulation that models extensibility, elasticity, bending plasticity, and interactions, plus a DLO agent handling topological complexity via strategic grasping and task decomposition.
- Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation β Identifies a coverage-density trade-off in budget-constrained VLA adaptation and proposes Anchor-Centric Adaptation, a two-stage method that stabilizes a policy skeleton with repeated anchor demonstrations then selectively expands coverage to high-risk boundaries.
- FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects β FlatLab pairs a strategy-generator/action-execution framework for manipulating flat objects with a high-fidelity simulation benchmark for diverse rigid and deformable flat objects.
-
ManiSoft: Towards Vision-Language Manipulation for Soft Robotics β Introduces ManiSoft, a benchmark and simulator for vision-language manipulation with soft robotic arms, featuring soft-body dynamics, four deformable-control tasks, and 6,300 scenes with expert trajectories. β
6,300 diverse scenes with expert trajectories -
OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning β Introduces AugE-Toolkit and the OXE-AugE dataset, augmenting Open X-Embodiment with 9 robot embodiments and over 4.4 million trajectories to study how robot augmentation improves cross-embodiment generalist policy learning. β
improving success rates by 24-45% on previously unseen robot-gripper combinations across four real-world manipulation tasks -
SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics β A generative simulation benchmark grounded in a high-fidelity chemistry lab that integrates an LLM task-synthesis engine, an automated expert, and an interactive RL environment to evaluate and improve embodied safety in precision-critical manipulation. β
RL post-training pipeline improves success rates by 37% -
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation β Presents an efficient video augmentation framework that converts simulated VLA videos into realistic training videos via structured-condition extraction, caption rewriting, and conditional video transfer, with diffusion feature-reuse and coreset sampling for scalability. β
improves RDT-1B by 8% on RobotWin 2.0 and boosts pi0 by 5.1% on LIBERO-Plus -
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models β An open-source VLA benchmarking framework structuring tasks across task structure, language, and visual observation dimensions with Safety/Distractor/Extrapolation/Long-Horizon categories and perturbation-based diagnostics. β
11 task suites ... containing 170 total tasks at three difficulty levels (L0-L2)
- Joint Navigation and Manipulation Planning with 3D Interaction Chains β Introduces 3D Interaction Chains, a unified open-vocabulary mobile-manipulation framework that couples navigation and manipulation planning over a shared 3D feature map, scoring multi-stage waypoint chains by VLM-based feasibility and transition cost.
- Learning AttributeβAffordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding β An Attribute-Affordance Hierarchies framework localizes affordance regions on 3D objects from images or text by modeling hierarchical attribute-affordance relationships with hypergraphs and hyperbolic-space concept embeddings.
-
VLA is now the default manipulation paradigm at ICML too. 21 of 99 papers are core VLA architecture/backbone work, and VLA framing pervades the efficiency, reasoning, RL, world-model, and benchmark clusters. ICML β historically a methods venue β has fully absorbed the VLA agenda that CoRL/ICLR/CVPR drove over 2024β2025.
-
Efficiency & on-robot deployment is the second-largest cluster (9+ papers). SpecPrune-VLA, EcoVLA, See-What-Matters (GridS), Speedup-Patch, Sparse-ActionGen, Reflex (50 Hz streaming), STEP, and a modelβhardware XPU characterization paper all target real-time/edge inference. Reported speedups cluster around 1.5β4Γ. This is the strongest "make VLAs actually runnable" wave in any 2026 venue so far.
-
Reasoning is going latent and test-time, not textual-CoT. Latent-Reasoning-VLA (
β90% inference latencyvs explicit CoT), LaSTβ, AVA-VLA (early-exit), VLA-ATTC and SCALE (uncertainty-gated test-time compute), Sentinel-VLA (metacognitive monitoring). The explicit language-CoT VLA of 2025 is being replaced by latent reasoning + adaptive compute β same efficiency pressure as trend #2. -
World-model-augmented VLA is a mature cluster (12 papers). Latent-action world models (LAC-WM, MoLA), co-improvement loops (VLAW, model-based RL VLA), structured-4D / semantic-mask prediction, and real-to-sim neural simulators (SoMA, RoboFlow4D). Continues the ICLR-2026 "world model as policy/evaluator" thread, now with explicit VLA coupling.
-
A robustness/skepticism backlash, mirroring CVPR's LIBERO-Plus. Diagnostic benchmarks (LIBERO-Gen "Dismantling the Illusion", VLA-Arena perturbations, RoboMME memory), safety alignment (PACT), the first targeted adversarial attack on VLA chain-of-thought (TRAP), and causal interpretability metrics. The field is auditing the headline-number inflation it produced.
-
Ο0 / Ο0.5 is the universal baseline. HALO, Move-Then-Operate, Sentinel-VLA, VLA-ATTC, RoboMME, "Dismantling the Illusion" and others benchmark directly against Physical Intelligence's Ο0/Ο0.5 β it is now the de-facto reference policy the way OpenVLA was in 2024.
-
ICML's signature: empirical "demystifying" papers. VLANeXt (12 findings β a recipe), From Pixels to Tokens (latent-action supervision study), Demystifying Action Space Design (13,000+ rollouts, 500+ models), Pretrained VLAs Resist Forgetting. Where CoRL ships robots and CVPR ships perception, ICML ships controlled studies of VLA design choices β the most rigorous methodology cluster of the 2026 venues.
-
New benchmark glut (11 papers). RoboTwin 2.0 (bimanual), VLA-Arena, RoboMME (memory), ManiSoft (soft robots), DLO-Lab (deformable linear objects), FlatLab (flat objects), SafeLab (chemistry-lab safety), AIR-VLA (aerial manip), CaP-X (coding agents), OXE-AugE (4.4M-traj OXE augmentation). Benchmarks now specialize by object physics and embodiment rather than competing as general suites.
ICML 2026 (July) lands after CVPR 2026 (June) and the rebuilt ICLR 2026 (May). Threads that carry through:
- World-model-as-policy/evaluator (ICLR β CVPR GigaBrain/CoWVLA β ICML's 12-paper world-model cluster).
- Latent action representations (ICLR villa-X/UniVLA β ICML XR-1, LARA, LAC-WM, MoLA).
- Efficiency/pruning (ICLR FASTER/SP-VLA β ICML's deployment cluster, now hardware-aware).
- Robustness audits (CVPR LIBERO-Plus β ICML LIBERO-Gen / VLA-Arena / RoboMME).
- RDT2 appears in both the CVPR-2026 and ICML-2026 lists in this wiki β confirm the canonical venue.
β οΈ De-duplication note. Several titles (XR-1, RDT2, Discrete Diffusion VLA, Human-Video VLA pretraining) also have wiki pages tagged to ICLR 2026 or CVPR 2026. ICML, ICLR, and CVPR 2026 overlap in time and some are distinct papers with similar names; others may be the same work. These need a venue-of-record reconciliation pass before per-paper pages are created.
-
Source of truth:
https://icml.cc/static/virtual/data/icml-2026-orals-posters.json(6,804 events; 6,636 distinct papers; 168 Orals). -
Filter: title + keyword match on
vision-language-action / VLA / manipulation / dexterous / grasp / bimanual / humanoid / tactile / affordance / visuomotor / diffusion-policy / flow-policy / imitation / cross-embodiment / world-model+robot, with negative filters for off-topic "manipulation" (image forgery, market/LLM-judge manipulation, graph reasoning) and exclusion of autonomous-driving-only VLAs, pure navigation, and pure learning-theory. -
Per-paper verification: each candidate's abstract was fetched from its
icml.cc/virtual/2026/poster/<id>page; summaries and quoted numbers are abstract-derived. 220 candidates β 99 confirmed manipulation papers. - Known gaps: author institutions unavailable (OpenReview guest-blocked); abstracts not yet cross-checked against arXiv full text; venue-of-record overlaps with CVPR/ICLR 2026 not yet reconciled.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)