Skip to content

RSS 2026 TMRL

hwoo.han edited this page Aug 9, 2026 · 2 revisions

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Imitation learning 3 Β· paper #208 Authors: Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta arXiv: 2605.12236 Β· program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Context aliasing enables action borrowing (Figure 1 of arXiv 2605.12236, Β© the authors)

Left: on a downstream task ("put the carrot in the pot") the BC conditional p(a|c) has collapsed support and fails. Center: injecting forward-diffusion noise into the context c aliases nearby contexts, moving the action distribution from sharp imitation p(a|c) toward the marginal p(a) β€” no overlap at Οƒ=0, partial overlap at medium Οƒ, full overlap at high Οƒ β€” so the policy "borrows" reach and place behaviors from other tasks. Right: the adapted policy completes the task.

Problem

BC-pretrained policies model narrow conditional action distributions: under distribution shift, support collapses, online rollouts earn no reward, and RL fine-tuning stalls. Prior fixes inject Gaussian action noise for coverage, which yields incoherent dithering and cannot be controlled during fine-tuning.

Method

Context-Smoothed Pre-training (CSP) applies a forward-diffusion corruption kernel to the policy's inputs (contexts) β€” states, point clouds, or VLM embeddings β€” training a generative policy p(a | noisy c, Οƒ) across all noise levels, creating a continuum from sharp imitation to the marginal action distribution. Timestep-Modulated RL (TMRL) then trains a high-level policy that adaptively outputs the diffusion timestep (context-noise dial Οƒ) plus a latent, treating conditioning strength as an explicit exploration control. Theory shows context smoothing strictly increases action coverage between overlapping contexts. The recipe is policy-agnostic: state-input diffusion policies, point-cloud policies, and Ο€0-style VLAs (noise on VLM embeddings before the action expert).

Results

On OGBench (pointmaze-giant, cube-single), CSP beats BC and PostBC on success@K coverage, and TMRL delivers a 101% average improvement over the best baseline (DSRL, PostBC, RLPD, SPiRL), including near-100% on cube-single where DSRL is near zero. On LIBERO adaptation, only TMRL solves the libero-90 task. With a LEAP-hand dexterous grasping policy on point clouds, TMRL reaches 2.5Γ— DSRL's final success. In the real world, fine-tuning Ο€0 on WidowX (BridgeData-v2) and Franka DROID tasks (sausage-in-pot, shrimp-in-drawer, press-button), TMRL reaches high success within tens of episodes β€” under one hour β€” while DSRL fails to learn; converged policies visibly anneal Οƒ downward within a rollout.

Significance

Reframes pre-training-for-RL from action-noise injection to structured context smoothing, giving RL an interpretable knob that interpolates between imitation and exploration β€” and demonstrating practical real-robot VLA fine-tuning in under an hour. Related wiki threads: Review-Human-Video-Transfer Β· Review-Realtime-Execution.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally