Skip to content

ICML 2026 DADP

hwoo.han edited this page Jun 11, 2026 · 1 revision

DADP: Domain Adaptive Diffusion Policy β€” Disentangling static domain identity from transient dynamics for zero-shot adaptation

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Pengcheng Wang, Qinghang Liu, Haotian Lin, Yiheng Li, Guojian Zhan, Masayoshi Tomizuka, Yixiao Wang Traction (2026-06): 1 citation (arXiv)

Averaged normalized performance of baselines across In-Distribution and Out-of-Distribution settings (Figure 1 from Wang et al., 2026)

Problem

Learning domain-adaptive policies that generalize to unseen transition dynamics (e.g., a robot or agent encountering a body mass, gravity, or friction it was never trained on) remains a fundamental challenge in learning-based control. A common recipe is domain representation learning: infer a latent vector that captures domain-specific information and condition the policy on it. The authors analyze how such representations are learned via dynamical prediction and identify a failure mode: when the context used for prediction is drawn from steps adjacent to the current step, the learned representation entangles static domain information (which should be constant across an episode) with transient dynamical properties (velocity, current state). This mixture confuses the conditioned policy and constrains zero-shot adaptation.

Method

DADP achieves robust adaptation through two ideas: unsupervised disentanglement of the domain code, and domain-aware injection into the diffusion process.

Design intuition: velocity inferred from another episode in the same domain cannot assist prediction, so only the static factor (e.g. gravity) is extracted (Figure 5 from Wang et al., 2026)

  1. Lagged Context Dynamical Prediction. Instead of conditioning future-state estimation on a temporally adjacent context, DADP conditions it on a historical offset context separated by a temporal gap Ξ”t. As Ξ”t grows, transient properties (which differ between the offset context and the target) stop being useful for prediction, so the encoder is forced to retain only the static domain factor. This disentangles static domain representations from transient dynamics without supervision and without contrastive learning or extra data generation.

  2. Domain-aware Diffusion Modulation. The learned domain representation is injected directly into the generative diffusion process by (i) biasing the prior distribution the diffusion samples from and (ii) reformulating the diffusion target, so domain information shapes both the start and the endpoint of denoising rather than being a passive conditioning token.

flowchart LR
  A[Trajectory] --> B[Lagged context offset Ξ”t]
  B --> C[Dynamical prediction]
  C --> D[Static domain representation]
  D --> E[Bias diffusion prior]
  D --> F[Reformulate diffusion target]
  E --> G[Domain-aware diffusion policy]
  F --> G
Loading

Results

Experiments span locomotion and manipulation benchmarks under both in-distribution (IID) and out-of-distribution (OOD) dynamics, reported as mean Β± std over 5 seeds against CORRO, Prompt-DT, and Meta-DT.

  • On Walker2d, DADP reaches 3991 (IID) / 3015 (OOD) vs. Meta-DT's 1304 / 889 and an expert reference of 7101.
  • On HalfCheetah, DADP scores 4100 (IID) / 3056 (OOD), exceeding the expert reference (4575 IID) on the OOD-robust trend and beating Meta-DT (3857 / 3174).
  • On manipulation tasks Door (1483 OOD) and Relocate (βˆ’5.63 OOD), DADP matches or surpasses the best baseline.
  • DADP shows the smallest standard deviation across seeds, indicating strong stability.
  • Disentanglement quality: increasing Ξ”t raises linear-probe accuracy of the embedding to 99.3–99.8% (Walker2d) and 99.9% (HalfCheetah) while driving reconstruction loss down by orders of magnitude β€” far above the supervised/short-gap baselines (e.g., 27.9% probe accuracy at the supervised setting on Walker2d).

Significance

DADP shows that how the prediction context is sampled is decisive: a single, simple temporal-offset trick yields cleanly disentangled static domain codes without contrastive objectives or synthetic data, and feeding those codes into both the prior and target of a diffusion policy delivers strong, stable zero-shot adaptation across locomotion and manipulation.

Links

← Back to ICML-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally