-
Notifications
You must be signed in to change notification settings - Fork 0
ICML 2026 Decompose and Recompose
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation β Atomic skillβaction pairs for zero-shot cross-task generalization
Venue: ICML 2026 (Poster) Category: Reasoning Affiliations: Xitie Zhang, Aming Wu, Yahong Han

Cross-task generalization is a core challenge in open-world robotic manipulation, and the key is extracting transferable manipulation knowledge from seen tasks. Recent in-context learning (ICL) approaches feed seen-task demonstrations to a large model to generate actions for unseen tasks without parameter updates. However, existing methods provide only low-level continuous action sequences as context, which fails to capture composable skill knowledge and causes the model to degenerate into superficial trajectory imitation.
Decompose and Recompose is a skill-reasoning framework built on atomic skillβaction pairs as intermediate representations that bridge high-level task semantics and low-level control.

- Decompose: seen demonstrations are broken into interpretable skillβaction alignments (atomic skillβaction pairs) via keyframe detection, producing composable intermediate representations.
- Dual-library demonstration retrieval: a task-adaptive dynamic library is built via visual-semantic retrieval combined with skill sequences from a planning agent, and a coverage-aware static library complements it to fill missing skill patterns. Together they yield skill-comprehensive demonstration sets.
- Recompose: the skill-augmented demonstrations explicitly elicit the LLM's compositional reasoning, so it composes existing skills and infers execution ordering for unseen tasks β zero-shot, with no parameter updates.
For execution, continuous control u = [p, q, g] (end-effector position, orientation quaternion, binary gripper) interfaces with a standard RLBench motion planner; for LLM interaction, translation and rotation are discretized into integer bins and represented as a 7-tuple action token, then decoded back to continuous control.
Evaluated on the AGNOSTOS benchmark (23 unseen tasks split into Level-1 and Level-2 tiers; standard protocol of 25 rollouts Γ 3 seeds = 75 episodes per task) plus real-world experiments:
- The method achieves the highest overall success rates across both difficulty levels versus diverse VLA baselines (foundation VLA models and other ICL approaches).
- It exceeds 60% success on four distinct tasks (Microwave, Seat, LampOff, USB), whereas individual baselines hit that threshold on at most three.
- Ablations: with no in-context demonstrations the model fails entirely (0% success) β demonstrations are essential. Overall success rises from 19.8% β 26.4% as demonstrations grow from 5 to 20 (Level-1 27.3%β32.5%, Level-2 13.1%β18.5%). For the coverage-aware static library, performance climbs from 24.9% (dynamic-only) to a best 26.4% at 3 static demos, then slightly drops at 4 β indicating an optimal balance between coverage and task relevance.
- Real-world tasks (e.g., stack cups, stack blocks) confirm the framework generates appropriate skill sequences and executes precise actions across diverse scenarios.
By making atomic skillβaction pairs the unit of in-context reasoning, Decompose and Recompose moves beyond trajectory imitation toward genuine compositional skill reuse, enabling training-free zero-shot transfer to novel objects and goals. The dual-library retrieval strategy β pairing task-adaptive dynamic retrieval with coverage-aware static complementation β is a practical recipe for supplying an LLM with the right skill vocabulary, and it sets a new bar on AGNOSTOS cross-task generalization.
- arXiv: 2605.01448
- ICML 2026: https://icml.cc/virtual/2026/poster/63250
β Back to ICML-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)