Skip to content

ICLR 2026 Robust Param Merging

Heungwoo edited this page Jun 1, 2026 · 4 revisions

RETAIN β€” Robust VLA Fine-tuning via Parameter Merging

Venue: ICLR 2026 Category: VLA Training β€” Adaptation / robustness Trend tag: Continual learning / robust fine-tuning

Approach diagram

flowchart LR
  PT[Pretrained generalist VLA<br/>ΞΈ_pre = Ο€0-FAST-DROID / Ο€0-LIBERO] --> FT[Task-FT or Co-FT<br/>~50-100 demos]
  FT --> THFT[Fine-tuned weights ΞΈ_ft]
  PT --> Merge["ΞΈΜƒ = (1-Ξ±) Β· ΞΈ_pre + Ξ± Β· ΞΈ_ft<br/>Ξ± tuned on val OOD scene"]
  THFT --> Merge
  Merge --> Eval[Three evaluations]
  Eval --> ID[Target task ID]
  Eval --> OOD[Target task OOD]
  Eval --> Gen[Generalist tasks]
  Merge -->|continual| Merge2["ΞΈΜƒ_2 = (1-Ξ±) Β· ΞΈΜƒ_1 + Ξ± Β· ΞΈ_ft,2<br/>add next skill"]
Loading

Problem

Generalist VLAs (Ο€0, OpenVLA, GR00T, Gemini Robotics) trained on large multi-task corpora generalise impressively out of the box, but practical deployments still fine-tune them on ~50-100 demonstrations for new tasks. In this low-data regime two failure modes appear simultaneously:

  1. Forgetting: the fine-tuned policy degrades on the generalist tasks it could solve before fine-tuning.
  2. Over-fitting: even on the target task, the policy fails on small variations not seen in the fine-tuning dataset (new object instances, lighting, distractors, viewpoints).

Figure 4 in the paper makes this concrete: standard task-FT gradually destroys generalist performance as gradient steps increase, and the gap between in-distribution (ID) and out-of-distribution (OOD) target-task success widens β€” i.e. the pretrained policy's generalisation ability is not transferring to the new task.

The authors show that even careful learning-rate / step-count tuning (Figure 16 in the appendix) does not resolve this: lower LR retains more generalist knowledge but underfits OOD on the target task; higher LR achieves ID success but kills OOD and generalist.

Detailed Method

Core proposal: linear weight interpolation

Given pretrained ΞΈ_pre and fine-tuned ΞΈ_ft, RETAIN produces a final policy by linear interpolation:

ΞΈΜƒ = (1 βˆ’ Ξ±) Β· ΞΈ_pre + Ξ± Β· ΞΈ_ft (Eq. 2)

where α ∈ [0, 1] is a tunable merging coefficient. No additional training, no inference-time overhead.

Co-fine-tuning (RETAIN-co-FT)

When the pretraining dataset (or a subset) is available, the authors fine-tune on a mix of D_pre and D_Ξ· first, then merge:

  • RETAIN-task-FT: merge ΞΈ_pre with task-only fine-tuned weights.
  • RETAIN-co-FT: merge ΞΈ_pre with co-fine-tuned weights. Consistently better (Sec. 6.2): co-FT prevents target-side overfitting; merging then explicitly re-injects pretrained knowledge.

Modality-specific merging

VLAs are vision-language-action stacks (vision encoder ΞΈ_v, language model backbone ΞΈ_l, action expert ΞΈ_a). Allow independent coefficients (Eq. 3):

ΞΈΜƒ_v = (1βˆ’Ξ±_v) ΞΈ_pre,v + Ξ±_v ΞΈ_ft,v ΞΈΜƒ_l = (1βˆ’Ξ±_l) ΞΈ_pre,l + Ξ±_l ΞΈ_ft,l ΞΈΜƒ_a = (1βˆ’Ξ±_a) ΞΈ_pre,a + Ξ±_a ΞΈ_ft,a

A 3D grid sweep on mugs-on-plates (Figure 11) reveals that Ξ±_l (language model) has the largest gradient on OOD performance β€” best at Ξ±_l = 0.8 β€” while Ξ±_v = Ξ±_a = 1 is optimal. Merging only the language-model parameters matches full-merging performance.

Continual sequential adaptation (Eq. 4)

For a sequence of target tasks T_1, ..., T_N, accumulate merges:

ΞΈΜƒ_n = (1 βˆ’ Ξ±) ΞΈΜƒ_{nβˆ’1} + Ξ± ΞΈ_ft,n

The same merging operator is reused at each stage, building a single growing policy.

Hyperparameter selection

  • Ξ± swept ∈ {0.25, 0.5, 0.75} on DROID (with one OOD scene held out as validation; the chosen Ξ± is then applied unchanged to other test OOD scenes).
  • LIBERO uses a similar val/test split.

Pretrained policies used

  • DROID experiments: Ο€β‚€-FAST-DROID (autoregressive next-token transformer, FAST tokenizer, trained on all of DROID + Physical Intelligence robot data).
  • LIBERO experiments: Ο€β‚€ (flow-based action expert) fine-tuned on LIBERO-{object, spatial, goal, 90}.

Comprehensive Results

Real-world DROID tasks (Figure 7)

  • whiteboard: 50 human-teleop demos, single fixed setup, 5 eraser positions Γ— 2 orientations.
  • plates: 100 demos, 5 plate colours Γ— 2 dish racks Γ— 2 orientations; 80 demos contain training-set distractors.

OOD tests vary backgrounds, object instances, camera angles. Each policy is evaluated 10 trials per scene; OOD has val + 2 test scenes.

Reported trends (paper plots, Figure 7):

  • Baselines (Task-FT, Co-FT, LoRA, FreezeFT) reach 70-80 % ID success but only 30-50 % OOD success.
  • RETAIN-task-FT and RETAIN-co-FT reach >60 % on plates OOD and ~80 % on whiteboard OOD β€” comparable to ID performance, indicating the policy generalised the new skill rather than memorising it.
  • On generalist evaluations (44 distinct real-world DROID tasks distributed across 9 scenes and 17 different language instructions, per Tables 3-5 / App. A.6.7; LIBERO uses 20 random pretraining tasks, 5 from each of object/spatial/goal/90), RETAIN matches the pretrained Ο€0-FAST-DROID baseline; Task-FT and LoRA degrade noticeably.

The paper's headline number: RETAIN finetuned policies achieve ~40 % higher OOD success on average than the best prior fine-tuning method on real DROID.

LIBERO simulation tasks (Figure 8, averaged over 3 tasks)

Tasks: pot-on-stove, mugs-on-plates, items-into-basket. ~45 demos each (filtered from 50 provided), tested on 3 OOD scenes per task with new initial positions, backgrounds, distractors.

  • Most baselines saturate near-perfect on ID due to LIBERO's relative simplicity.
  • OOD shows the same pattern as DROID β€” RETAIN improves over Co-FT, though by a smaller margin than DROID. The authors attribute this to the weaker generalist base: their LIBERO Ο€0 was trained on only 5.3 k trajectories / 117 scenes, vs Ο€0-FAST-DROID's 76 k trajectories / 564 scenes.

Pretraining-data scaling (Figure 9, plates task)

Three pretrained policies merged-then-evaluated:

  1. Ο€0-FAST-DROID (full DROID + PI internal data) β€” biggest RETAIN gain on OOD.
  2. DROID-only (76k episodes).
  3. DROID-subset (20k episodes) β€” smallest gain.

ID is similar across all three. OOD gain scales monotonically with pretraining data: with the most general base, OOD performance is "nearly as good as ID" β€” i.e. the policy fully transfers its base generalisation to the new task.

Modality-specific merging analysis (Figure 11)

  • Ξ±_l (language model) most influential: best at Ξ±_l β‰ˆ 0.8 on mugs-on-plates.
  • Ξ±_v, Ξ±_a = 1 best: vision encoder and action expert benefit from staying at the fine-tuned values.
  • Merging only language-backbone parameters matches full-merge OOD performance across all three LIBERO tasks.

Continual learning (Figure 12)

Sequential plates β†’ whiteboard:

  • RETAIN preserves performance on plates after a second round of fine-tuning on whiteboard, while co-FT in the same sequential setting regresses.
  • On the second task (whiteboard), RETAIN also outperforms co-FT under both ID and OOD, despite no special replay.

Learning-rate ablation (Figure 16)

Four LRs evaluated every 100 gradient steps. Larger LR β†’ faster overfitting and generalist collapse to ~0 %. Smaller LR β†’ slower forgetting but worse OOD performance (i.e. underfitting). No LR setting alone resolves the problem β€” only weight merging does.

Parameter-trajectory analysis (Figures 17-19)

  • Cosine similarity of consecutive parameter-difference vectors during fine-tuning is far from 1 β‡’ fine-tuning path is highly non-linear.
  • 2D PCA projection of the parameter trajectory shows oscillation (low LR) or curving (high LR); merged checkpoints land in a different region of parameter space, not on the fine-tuning trajectory.
  • Singular-value decomposition of the difference matrix shows many non-zero singular values, confirming non-linearity in many directions.

Ablation Studies

  • Ξ± sweep on LIBERO (Figure 10). OOD performance peaks at intermediate Ξ±; Ξ± β†’ 0 (pretrained) gives ~0 % task success; Ξ± β†’ 1 (fine-tuned) gives full overfitting; in between, merging benefits emerge.
  • Modality-specific merging. Confirms language-backbone is the dominant contributor β€” informs efficient deployment.
  • task-FT vs co-FT before merging. RETAIN-co-FT > RETAIN-task-FT on generalist evals nearly always; on OOD the gap is smaller. The authors attribute different roles: co-FT prevents target overfitting at training time; merging elicits pretrained knowledge in parameter space.
  • Pretraining-data scaling. OOD gain from RETAIN scales with pretraining size (Figure 9 β†’ Figure 15).
  • Sequential vs single-task. Sequential RETAIN approaches single-task oracle; sequential co-FT does not.
  • No-merge baselines. Comparison set: Task-FT, Co-FT, LoRA, FreezeFT (freeze language backbone, only update vision encoder + action expert head), Scratch.

Limitations

The authors explicitly state:

  1. Mechanistic understanding incomplete. Why parameter merging produces this generalisation transfer is "an interesting area for future work"; the paper offers analysis (non-linear path) but no rigorous theoretical explanation.
  2. Hyperparameter dependence. The merging coefficient Ξ± requires tuning. The authors note RETAIN is robust to Ξ± in real-world experiments (one validation OOD scene generalises to test OOD scenes) but a heuristic for choosing Ξ± a priori is left to future work.

Implicit limitations from the paper:

  • Smaller OOD gains on LIBERO than on DROID, attributed to the weaker LIBERO base β€” RETAIN's effectiveness depends on having a sufficiently generalist starting checkpoint.
  • All real-world experiments are on a Franka with DROID-style setups; no humanoid or mobile-manipulator results.
  • Continual learning shown for 2 sequential tasks; longer sequences are not evaluated.
  • Action-expert merging (Ξ±_a < 1) is not helpful β€” the method's elegance partly comes from leaving the action head alone, but this also means RETAIN does not directly address VLA action-head specialisation.

Significance & Positioning

RETAIN is, in the authors' words, the first work to investigate and analyze parameter merging for robot policies. The paper makes three contributions:

  1. Empirical: ~40 % absolute OOD improvement on real-robot DROID tasks at zero additional training or inference cost. This is the kind of "free lunch" the field has been looking for in the small-demo fine-tuning regime that dominates practical VLA deployments.
  2. Methodological: modality-specific merging analysis identifies the language-model backbone as the load-bearing parameter group β€” concrete guidance for memory-constrained deployments.
  3. Conceptual: weight merging scales with the generality of the base model (Figure 9). This connects directly to the Ο€0.6 / Ο€0.7 / GR00T Series story: as pretraining data and base-model generality grow, RETAIN gets more effective. It is not in tension with continued scale-up; it is a multiplier on top.

Within the 2026 robust-fine-tuning cluster:

  • Align-Then-Steer adapts in distribution space (steering output statistics) while RETAIN adapts in parameter space. Both target the same problem with orthogonal mechanisms; combining them is unexplored.
  • VLA Robustness is more about identifying failure modes; RETAIN is one prescriptive answer.
  • Knowledge Insulation addresses forgetting via gradient-space techniques; RETAIN sidesteps the forgetting question entirely by never overwriting ΞΈ_pre.
  • Compared to PEFT methods like LoRA, RETAIN does not freeze parameters β€” it allows full fine-tuning then merges, which empirically beats LoRA in both ID and OOD on DROID.
  • For continual learning, RETAIN provides an alternative to EWC / Progress-and-Compress / experience replay, requiring no Fisher matrix or replay buffer β€” only the predecessor checkpoint.

The paper's clearest implication for the field: the next round of VLA system design should treat the pretrained checkpoint as a first-class deployable artifact, not a starting point to be overwritten. Cheap parameter merging then gives every new task the equivalent of a "soft reset" to the generalist.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally