You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We evaluate a suite of optimizers on Transformer-style language models (130 M–1.2 B parameters) trained on up to 16× Chinchilla-optimal data. The goal is to quantify real speedups under rigorous per-optimizer tuning.
Models: 130 M, 300 M, 520 M, 1.2 B parameter Transformers (seq-len 4 096). Detailed hyperparameters:
Model
Params
Seq Len
Hidden Dim
Inter Dim
# Layers
# Heads
LLaMA-130M
130M
4096
512
2048
32
8
LLaMA-300M
300M
4096
768
3072
32
12
LLaMA-520M
520M
4096
1024
4096
32
16
LLaMA-1.2B
1.2B
4096
1536
6144
32
24
Data: Mixture of DCLM-baseline (3.8 T tokens), StarCoder V2 (0.25 T), ProofPile 2 (0.055 T); tokenized via LLaMA-3 tokenizer. Chinchilla is set to ~20× non-embedding param count.
Experiment Files: Please refer to the breakdown in the methodology section.
Results
For the in-depth analysis, please refer to the Wandb report. Here is a TL;DR.
Max Speedup ≤ 1.4× and Matrix Preconditioners Lead
No optimizer achieved the 2× step-wise speedup from prior claims; the best was ≈ 1.4× over AdamW. Muon, Soap, and Kron outperformed AdamW/NAdamW/Mars across regimes.
Regime-Dependent Winner
Muon wins in 1×–4× Chinchilla; Soap/Kron takes over at ≥ 8× and in over-trained (16×) settings.
We also note many subtleties about hyperparameters emerge from our evaluation.
Optimal Hyperparameters May Change Drastically Across Optimizers
For example in comparison of Lion and AdamW in 130M 1x Chinchilla setting.
Speedup may diminish even if one hyperparameter is slightly misset
For example in comparison of Soap and Mars in 520M 8x Chinchilla setting.
Full-Schedule Evaluation Required
Optimizers can flip rankings during learning rate decay—always use end-of-schedule metrics.
Description
We evaluate a suite of optimizers on Transformer-style language models (130 M–1.2 B parameters) trained on up to 16× Chinchilla-optimal data. The goal is to quantify real speedups under rigorous per-optimizer tuning.
Methodology
The experiment code is now at #1293
Data: Mixture of DCLM-baseline (3.8 T tokens), StarCoder V2 (0.25 T), ProofPile 2 (0.055 T); tokenized via LLaMA-3 tokenizer. Chinchilla is set to ~20× non-embedding param count.
Optimizers: we evaluated the following ten optimizers, our implementation code can be found at Add Modern Optimizers in Levanter levanter#955
AdamW: Decoupled Weight Decay Regularization (Loshchilov & Hutter, 2017)
NAdamW: Incorporating Nesterov Momentum into Adam (Dozat, 2016)
Mars: MARS: Unleashing the Power of Variance Reduction for Training Large Models (Yuan et al., 2024)
Cautious: Cautious Optimizers: Improving Training with One Line of Code (Liang et al., 2024)
Lion: Symbolic Discovery of Optimization Algorithms (Raffel et al., 2023)
Adam-mini: Adam-mini: Use Fewer Learning Rates To Gain More (Zhang et al., 2024)
Muon: Muon: An optimizer for the hidden layers of neural networks (GitHub)
Scion: LIONS-EPFL/scion (GitHub)
Kron (PSGD): kron_torch: PSGD with Kronecker-factored preconditioner (GitHub)
Soap: SOAP: Improving and Stabilizing Shampoo using Adam (Vyas et al., 2024)
Phase I: Full Hyperparameter Sweep
Phase II: Sensitive-Hyperparameters Identification
Phase III: Scaling-Law Fitting
over 12 (model size, data size, hyperparameter) triples.
Hypothesis or Goal
Links
Results
For the in-depth analysis, please refer to the Wandb report. Here is a TL;DR.
Max Speedup ≤ 1.4× and Matrix Preconditioners Lead
No optimizer achieved the 2× step-wise speedup from prior claims; the best was ≈ 1.4× over AdamW. Muon, Soap, and Kron outperformed AdamW/NAdamW/Mars across regimes.
Regime-Dependent Winner
Muon wins in 1×–4× Chinchilla; Soap/Kron takes over at ≥ 8× and in over-trained (16×) settings.
We also note many subtleties about hyperparameters emerge from our evaluation.
For example in comparison of Lion and AdamW in 130M 1x Chinchilla setting.
For example in comparison of Soap and Mars in 520M 8x Chinchilla setting.
Optimizers can flip rankings during learning rate decay—always use end-of-schedule metrics.