-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 Robust Param Merging
Venue: ICLR 2026 Category: VLA Training β Adaptation / robustness Trend tag: Continual learning / robust fine-tuning
flowchart LR
PT[Pretrained generalist VLA<br/>ΞΈ_pre = Ο0-FAST-DROID / Ο0-LIBERO] --> FT[Task-FT or Co-FT<br/>~50-100 demos]
FT --> THFT[Fine-tuned weights ΞΈ_ft]
PT --> Merge["ΞΈΜ = (1-Ξ±) Β· ΞΈ_pre + Ξ± Β· ΞΈ_ft<br/>Ξ± tuned on val OOD scene"]
THFT --> Merge
Merge --> Eval[Three evaluations]
Eval --> ID[Target task ID]
Eval --> OOD[Target task OOD]
Eval --> Gen[Generalist tasks]
Merge -->|continual| Merge2["ΞΈΜ_2 = (1-Ξ±) Β· ΞΈΜ_1 + Ξ± Β· ΞΈ_ft,2<br/>add next skill"]
Generalist VLAs (Ο0, OpenVLA, GR00T, Gemini Robotics) trained on large multi-task corpora generalise impressively out of the box, but practical deployments still fine-tune them on ~50-100 demonstrations for new tasks. In this low-data regime two failure modes appear simultaneously:
- Forgetting: the fine-tuned policy degrades on the generalist tasks it could solve before fine-tuning.
- Over-fitting: even on the target task, the policy fails on small variations not seen in the fine-tuning dataset (new object instances, lighting, distractors, viewpoints).
Figure 4 in the paper makes this concrete: standard task-FT gradually destroys generalist performance as gradient steps increase, and the gap between in-distribution (ID) and out-of-distribution (OOD) target-task success widens β i.e. the pretrained policy's generalisation ability is not transferring to the new task.
The authors show that even careful learning-rate / step-count tuning (Figure 16 in the appendix) does not resolve this: lower LR retains more generalist knowledge but underfits OOD on the target task; higher LR achieves ID success but kills OOD and generalist.
Given pretrained ΞΈ_pre and fine-tuned ΞΈ_ft, RETAIN produces a final policy by linear interpolation:
ΞΈΜ = (1 β Ξ±) Β· ΞΈ_pre + Ξ± Β· ΞΈ_ft (Eq. 2)
where Ξ± β [0, 1] is a tunable merging coefficient. No additional training, no inference-time overhead.
When the pretraining dataset (or a subset) is available, the authors fine-tune on a mix of D_pre and D_Ξ· first, then merge:
- RETAIN-task-FT: merge ΞΈ_pre with task-only fine-tuned weights.
- RETAIN-co-FT: merge ΞΈ_pre with co-fine-tuned weights. Consistently better (Sec. 6.2): co-FT prevents target-side overfitting; merging then explicitly re-injects pretrained knowledge.
VLAs are vision-language-action stacks (vision encoder ΞΈ_v, language model backbone ΞΈ_l, action expert ΞΈ_a). Allow independent coefficients (Eq. 3):
ΞΈΜ_v = (1βΞ±_v) ΞΈ_pre,v + Ξ±_v ΞΈ_ft,v ΞΈΜ_l = (1βΞ±_l) ΞΈ_pre,l + Ξ±_l ΞΈ_ft,l ΞΈΜ_a = (1βΞ±_a) ΞΈ_pre,a + Ξ±_a ΞΈ_ft,a
A 3D grid sweep on mugs-on-plates (Figure 11) reveals that Ξ±_l (language model) has the largest gradient on OOD performance β best at Ξ±_l = 0.8 β while Ξ±_v = Ξ±_a = 1 is optimal. Merging only the language-model parameters matches full-merging performance.
For a sequence of target tasks T_1, ..., T_N, accumulate merges:
ΞΈΜ_n = (1 β Ξ±) ΞΈΜ_{nβ1} + Ξ± ΞΈ_ft,n
The same merging operator is reused at each stage, building a single growing policy.
- Ξ± swept β {0.25, 0.5, 0.75} on DROID (with one OOD scene held out as validation; the chosen Ξ± is then applied unchanged to other test OOD scenes).
- LIBERO uses a similar val/test split.
- DROID experiments: Οβ-FAST-DROID (autoregressive next-token transformer, FAST tokenizer, trained on all of DROID + Physical Intelligence robot data).
- LIBERO experiments: Οβ (flow-based action expert) fine-tuned on LIBERO-{object, spatial, goal, 90}.
- whiteboard: 50 human-teleop demos, single fixed setup, 5 eraser positions Γ 2 orientations.
- plates: 100 demos, 5 plate colours Γ 2 dish racks Γ 2 orientations; 80 demos contain training-set distractors.
OOD tests vary backgrounds, object instances, camera angles. Each policy is evaluated 10 trials per scene; OOD has val + 2 test scenes.
Reported trends (paper plots, Figure 7):
- Baselines (Task-FT, Co-FT, LoRA, FreezeFT) reach 70-80 % ID success but only 30-50 % OOD success.
- RETAIN-task-FT and RETAIN-co-FT reach >60 % on plates OOD and ~80 % on whiteboard OOD β comparable to ID performance, indicating the policy generalised the new skill rather than memorising it.
- On generalist evaluations (44 distinct real-world DROID tasks distributed across 9 scenes and 17 different language instructions, per Tables 3-5 / App. A.6.7; LIBERO uses 20 random pretraining tasks, 5 from each of object/spatial/goal/90), RETAIN matches the pretrained Ο0-FAST-DROID baseline; Task-FT and LoRA degrade noticeably.
The paper's headline number: RETAIN finetuned policies achieve ~40 % higher OOD success on average than the best prior fine-tuning method on real DROID.
Tasks: pot-on-stove, mugs-on-plates, items-into-basket. ~45 demos each (filtered from 50 provided), tested on 3 OOD scenes per task with new initial positions, backgrounds, distractors.
- Most baselines saturate near-perfect on ID due to LIBERO's relative simplicity.
- OOD shows the same pattern as DROID β RETAIN improves over Co-FT, though by a smaller margin than DROID. The authors attribute this to the weaker generalist base: their LIBERO Ο0 was trained on only 5.3 k trajectories / 117 scenes, vs Ο0-FAST-DROID's 76 k trajectories / 564 scenes.
Three pretrained policies merged-then-evaluated:
- Ο0-FAST-DROID (full DROID + PI internal data) β biggest RETAIN gain on OOD.
- DROID-only (76k episodes).
- DROID-subset (20k episodes) β smallest gain.
ID is similar across all three. OOD gain scales monotonically with pretraining data: with the most general base, OOD performance is "nearly as good as ID" β i.e. the policy fully transfers its base generalisation to the new task.
- Ξ±_l (language model) most influential: best at Ξ±_l β 0.8 on mugs-on-plates.
- Ξ±_v, Ξ±_a = 1 best: vision encoder and action expert benefit from staying at the fine-tuned values.
- Merging only language-backbone parameters matches full-merge OOD performance across all three LIBERO tasks.
Sequential plates β whiteboard:
- RETAIN preserves performance on plates after a second round of fine-tuning on whiteboard, while co-FT in the same sequential setting regresses.
- On the second task (whiteboard), RETAIN also outperforms co-FT under both ID and OOD, despite no special replay.
Four LRs evaluated every 100 gradient steps. Larger LR β faster overfitting and generalist collapse to ~0 %. Smaller LR β slower forgetting but worse OOD performance (i.e. underfitting). No LR setting alone resolves the problem β only weight merging does.
- Cosine similarity of consecutive parameter-difference vectors during fine-tuning is far from 1 β fine-tuning path is highly non-linear.
- 2D PCA projection of the parameter trajectory shows oscillation (low LR) or curving (high LR); merged checkpoints land in a different region of parameter space, not on the fine-tuning trajectory.
- Singular-value decomposition of the difference matrix shows many non-zero singular values, confirming non-linearity in many directions.
- Ξ± sweep on LIBERO (Figure 10). OOD performance peaks at intermediate Ξ±; Ξ± β 0 (pretrained) gives ~0 % task success; Ξ± β 1 (fine-tuned) gives full overfitting; in between, merging benefits emerge.
- Modality-specific merging. Confirms language-backbone is the dominant contributor β informs efficient deployment.
- task-FT vs co-FT before merging. RETAIN-co-FT > RETAIN-task-FT on generalist evals nearly always; on OOD the gap is smaller. The authors attribute different roles: co-FT prevents target overfitting at training time; merging elicits pretrained knowledge in parameter space.
- Pretraining-data scaling. OOD gain from RETAIN scales with pretraining size (Figure 9 β Figure 15).
- Sequential vs single-task. Sequential RETAIN approaches single-task oracle; sequential co-FT does not.
- No-merge baselines. Comparison set: Task-FT, Co-FT, LoRA, FreezeFT (freeze language backbone, only update vision encoder + action expert head), Scratch.
The authors explicitly state:
- Mechanistic understanding incomplete. Why parameter merging produces this generalisation transfer is "an interesting area for future work"; the paper offers analysis (non-linear path) but no rigorous theoretical explanation.
- Hyperparameter dependence. The merging coefficient Ξ± requires tuning. The authors note RETAIN is robust to Ξ± in real-world experiments (one validation OOD scene generalises to test OOD scenes) but a heuristic for choosing Ξ± a priori is left to future work.
Implicit limitations from the paper:
- Smaller OOD gains on LIBERO than on DROID, attributed to the weaker LIBERO base β RETAIN's effectiveness depends on having a sufficiently generalist starting checkpoint.
- All real-world experiments are on a Franka with DROID-style setups; no humanoid or mobile-manipulator results.
- Continual learning shown for 2 sequential tasks; longer sequences are not evaluated.
- Action-expert merging (Ξ±_a < 1) is not helpful β the method's elegance partly comes from leaving the action head alone, but this also means RETAIN does not directly address VLA action-head specialisation.
RETAIN is, in the authors' words, the first work to investigate and analyze parameter merging for robot policies. The paper makes three contributions:
- Empirical: ~40 % absolute OOD improvement on real-robot DROID tasks at zero additional training or inference cost. This is the kind of "free lunch" the field has been looking for in the small-demo fine-tuning regime that dominates practical VLA deployments.
- Methodological: modality-specific merging analysis identifies the language-model backbone as the load-bearing parameter group β concrete guidance for memory-constrained deployments.
- Conceptual: weight merging scales with the generality of the base model (Figure 9). This connects directly to the Ο0.6 / Ο0.7 / GR00T Series story: as pretraining data and base-model generality grow, RETAIN gets more effective. It is not in tension with continued scale-up; it is a multiplier on top.
Within the 2026 robust-fine-tuning cluster:
- Align-Then-Steer adapts in distribution space (steering output statistics) while RETAIN adapts in parameter space. Both target the same problem with orthogonal mechanisms; combining them is unexplored.
- VLA Robustness is more about identifying failure modes; RETAIN is one prescriptive answer.
- Knowledge Insulation addresses forgetting via gradient-space techniques; RETAIN sidesteps the forgetting question entirely by never overwriting ΞΈ_pre.
- Compared to PEFT methods like LoRA, RETAIN does not freeze parameters β it allows full fine-tuning then merges, which empirically beats LoRA in both ID and OOD on DROID.
- For continual learning, RETAIN provides an alternative to EWC / Progress-and-Compress / experience replay, requiring no Fisher matrix or replay buffer β only the predecessor checkpoint.
The paper's clearest implication for the field: the next round of VLA system design should treat the pretrained checkpoint as a first-class deployable artifact, not a starting point to be overwritten. Cheap parameter merging then gives every new task the equivalent of a "soft reset" to the generalist.
- OpenReview: https://openreview.net/forum?id=uWJwQ5SZoM
- Project: https://retain.yajatyadav.com
- Align-Then-Steer β distribution-space adaptation
- VLA Robustness
- Knowledge Insulation
- Ο0.6 Β· Ο0.7
- Survey: VLA & Manipulation
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)