Code for the paper "How does the optimizer implicitly bias the model merging loss landscape?". ICLR 2026.
The paper shows that merging success depends on optimizer noise scale in a non-monotonic way, with a distinct optimum. This holds across learning rate, weight decay, batch size, and augmentation.