-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 DISC
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 2 Β· paper #147 Authors: Hanxiang Ren, Pei Zhou, Xunzhe Zhou, Yanchao Yang arXiv: 2605.20856 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1. Two rollout comparisons illustrating observation leakage. Left ("put the white bowl to the right of the plate"): entangled Octo instead approaches the microwave, executing a behavior tied to a similar scene, while DISC follows the instruction. Right ("turn on the stove and put the frying pan on it"): pretrained Οβ.β skips the stove-activation subgoal and just places the pan, whereas DISC completes both steps.
Language-conditioned manipulation policies typically process instruction and observation through shared parameters β "task-state entanglement." This lets networks learn scene-to-action shortcuts that bypass language grounding entirely (observation leakage), so policies succeed on familiar scenes but fail when instructions change or visual context is ambiguous.
DISC removes the failure structurally: rather than conditioning one universal policy on language, a hypernetwork generates the entire parameter set of a task-specific visuomotor policy from the instruction alone, and that generated policy only ever sees observations β never language β so there is no pathway for leakage. Generating coherent high-dimensional weights is hard, so DISC uses a two-stage hypernetwork: a Weight Initialization Network maps the language embedding to a semantically informed point in parameter space, and an Iterative Refinement Module embeds the structure of gradient-based optimization (forward evaluation, error estimation, correction) as a feed-forward inductive bias, producing globally consistent parameters without actual gradient computation. It is trained entirely from scratch on standard data budgets with no external pretraining.
DISC achieves 94.3% on LIBERO-90 and 92.2% on Meta-World, beating the strongest trained-from-scratch entangled baseline by 7.7% on LIBERO-90, with advantages widening on long-horizon tasks. It surpasses the large-scale pretrained Οβ (91.6%) and remains competitive with Οβ.β (95.7%) despite using no pretraining data. On a real-world combinatorial benchmark where visual context is shared across all 9 tasks, DISC reaches 86.4% versus 78.5% for the best entangled baseline β direct evidence that generated parameters, not visual shortcuts, drive behavior. The learned parameter manifold further supports few-shot adaptation and robustness to paraphrased instructions.
Offers an architectural (rather than data-scale) route to genuine language grounding, sharpening the instruction-following debates in Review-LBM-Cotraining and pretraining-centric VLA work.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)