-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 TMRL
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 3 Β· paper #208 Authors: Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta arXiv: 2605.12236 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Left: on a downstream task ("put the carrot in the pot") the BC conditional p(a|c) has collapsed support and fails. Center: injecting forward-diffusion noise into the context c aliases nearby contexts, moving the action distribution from sharp imitation p(a|c) toward the marginal p(a) β no overlap at Ο=0, partial overlap at medium Ο, full overlap at high Ο β so the policy "borrows" reach and place behaviors from other tasks. Right: the adapted policy completes the task.
BC-pretrained policies model narrow conditional action distributions: under distribution shift, support collapses, online rollouts earn no reward, and RL fine-tuning stalls. Prior fixes inject Gaussian action noise for coverage, which yields incoherent dithering and cannot be controlled during fine-tuning.
Context-Smoothed Pre-training (CSP) applies a forward-diffusion corruption kernel to the policy's inputs (contexts) β states, point clouds, or VLM embeddings β training a generative policy p(a | noisy c, Ο) across all noise levels, creating a continuum from sharp imitation to the marginal action distribution. Timestep-Modulated RL (TMRL) then trains a high-level policy that adaptively outputs the diffusion timestep (context-noise dial Ο) plus a latent, treating conditioning strength as an explicit exploration control. Theory shows context smoothing strictly increases action coverage between overlapping contexts. The recipe is policy-agnostic: state-input diffusion policies, point-cloud policies, and Ο0-style VLAs (noise on VLM embeddings before the action expert).
On OGBench (pointmaze-giant, cube-single), CSP beats BC and PostBC on success@K coverage, and TMRL delivers a 101% average improvement over the best baseline (DSRL, PostBC, RLPD, SPiRL), including near-100% on cube-single where DSRL is near zero. On LIBERO adaptation, only TMRL solves the libero-90 task. With a LEAP-hand dexterous grasping policy on point clouds, TMRL reaches 2.5Γ DSRL's final success. In the real world, fine-tuning Ο0 on WidowX (BridgeData-v2) and Franka DROID tasks (sausage-in-pot, shrimp-in-drawer, press-button), TMRL reaches high success within tens of episodes β under one hour β while DSRL fails to learn; converged policies visibly anneal Ο downward within a rollout.
Reframes pre-training-for-RL from action-noise injection to structured context smoothing, giving RL an interpretable knob that interpolates between imitation and exploration β and demonstrating practical real-robot VLA fine-tuning in under an hour. Related wiki threads: Review-Human-Video-Transfer Β· Review-Realtime-Execution.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)