-
Notifications
You must be signed in to change notification settings - Fork 0
Cross LLM Comparison
How much does the LLM generation matter for autonomous hyperparameter optimization? This page tracks cross-generation comparison experiments: the same dataset, same hardware, same codebase, same training budget — but different Claude model versions driving the optimization.
Sonnet 4.0 vs 4.6: Complete (5/5 datasets). Sonnet 4.6 wins 3–2 on absolute val_bpb, discovers architectures Sonnet 4.0 misses, and crashes 2.1× less.
Haiku 4.5: In progress (ClimbMix — 22/101 experiments). Early results are striking: Haiku's inherited baseline already beats both Sonnets' best on ClimbMix, and after just one keep it reaches 1.2924 — the new absolute best.
| Metric | ClimbMix | FineWeb-Edu | Cosmopedia-v2 | SlimPajama | FineWeb-Edu-High |
|---|---|---|---|---|---|
| S4.0 baseline | 1.2969 | 1.4088 | 0.9640 | 1.5410 | 1.3730 |
| S4.6 baseline | 1.3213 | 1.3710 | 0.9806 | 1.5312 | 1.3478 |
| H4.5 baseline | 1.2966 | — | — | — | — |
| S4.0 best | 1.2959 | 1.3424 | 0.9606 | 1.5259 | 1.3463 |
| S4.6 best | 1.2997 | 1.3416 | 0.9549 | 1.5267 | 1.3345 |
| H4.5 best | 1.2924 | — | — | — | — |
| S4.0 improvement | −0.08% | −4.71% | −0.35% | −1.0% | −1.97% |
| S4.6 improvement | −1.63% | −2.14% | −2.63% | −0.30% | −0.99% |
| H4.5 improvement | −0.32% | — | — | — | — |
| S4.0 keeps | 1 | 17 | 4 | 3 | 20 |
| S4.6 keeps | 8 | 18 | 16 | 2 | 13 |
| H4.5 keeps | 1* | — | — | — | — |
| S4.0 crashes | 11 | 2 | 2 | 0 | 5 |
| S4.6 crashes | 5 | 1 | 0 | 2 | 2 |
| H4.5 crashes | 2* | — | — | — | — |
| Winner (absolute) | H4.5* | S4.6 | S4.6 | S4.0 | S4.6 |
* Haiku 4.5 in progress (22/101 experiments on ClimbMix). Remaining datasets pending.
Sonnet scorecard: Sonnet 4.6 wins 3–2 on absolute val_bpb (FineWeb-Edu, Cosmopedia-v2, FineWeb-Edu-High). Sonnet 4.0 wins on ClimbMix and SlimPajama.
Haiku early result: With inherited S4.6 config as its baseline, Haiku already holds the absolute best on ClimbMix (1.2924). However, this reflects baseline inheritance more than optimization skill — see fairness note below.
| Metric | Sonnet 4.0 (Mar 19) | Sonnet 4.6 (Mar 22) | Winner |
|---|---|---|---|
| Best val_bpb | 1.2959 | 1.2997 | Sonnet 4.0 |
| Improvement | −0.08% | −1.63% | Sonnet 4.6 (20×) |
| Keeps | 1 (1.0%) | 8 (6.8%) | Sonnet 4.6 |
| Crashes | 11 (11.0%) | 5 (4.2%) | Sonnet 4.6 |
Sonnet 4.0 wins on absolute val_bpb (pre-optimized baseline advantage), but Sonnet 4.6 found 20× more improvement and 7× more keeps.
| Metric | Sonnet 4.0 (Mar 17) | Sonnet 4.6 (Mar 22) | Winner |
|---|---|---|---|
| Best val_bpb | 1.3424 | 1.3416 | Sonnet 4.6 (by 0.0008) |
| Improvement | −4.71% | −2.14% | Sonnet 4.0 (more room) |
| Keeps | 17 (19.3%) | 18 (18.0%) | Comparable |
| Crashes | 2 (2.3%) | 1 (1.0%) | Sonnet 4.6 |
Near-identical results via completely different configs — 8 of 10 parameters differ. Proves FineWeb-Edu has multiple near-equivalent optima.
| Metric | Sonnet 4.0 (Mar 20) | Sonnet 4.6 (Mar 24) | Winner |
|---|---|---|---|
| Best val_bpb | 0.9606 | 0.9549 | Sonnet 4.6 (by 0.006) |
| Improvement | −0.35% | −2.63% | Sonnet 4.6 (7.5×) |
| Keeps | 4 (3.9%) | 16 (16.0%) | Sonnet 4.6 (4×) |
| Crashes | 2 (1.9%) | 0 (0.0%) | Sonnet 4.6 |
Sonnet 4.6's strongest showing. Discovered ASPECT_RATIO=21 (vs AR=32), breaking the AR=32 consensus from all Sonnet 4.0 runs. Zero crashes.
| Metric | Sonnet 4.0 (Mar 20) | Sonnet 4.6 (Mar 25) | Winner |
|---|---|---|---|
| Best val_bpb | 1.5259 | 1.5267 | Sonnet 4.0 (by 0.0008) |
| Improvement | −1.0% | −0.30% | Sonnet 4.0 |
| Keeps | 3 (3.0%) | 2 (2.0%) | Comparable |
| Crashes | 0 (0.0%) | 2 (2.0%) | Sonnet 4.0 |
The hardest dataset to optimize. Both models find SlimPajama nearly impervious — total improvements of just 1.0% and 0.30%. Sonnet 4.6 went 62 experiments without a single keep before finding a triple-synergy combination (β1, FINAL_LR_FRAC, WARMDOWN_RATIO together). Sonnet 4.0 edges out on absolute val_bpb.
| Metric | Sonnet 4.0 (Mar 21) | Sonnet 4.6 (Mar 25) | Winner |
|---|---|---|---|
| Best val_bpb | 1.3463 | 1.3345 | Sonnet 4.6 (by 0.012) |
| Improvement | −1.97% | −0.99% | Sonnet 4.0 (more room) |
| Keeps | 20 (20.0%) | 13 (12.9%) | Sonnet 4.0 |
| Crashes | 5 (5.0%) | 2 (2.0%) | Sonnet 4.6 |
Sonnet 4.6 wins decisively on absolute val_bpb (1.3345 vs 1.3463). The signature finding: a spectacular β2 walk — 5 consecutive keeps as β2 decreased from 0.98→0.975→0.970→0.968→0.966→0.964. Sonnet 4.0 had more keeps (20 vs 13) but started from a worse baseline with more room to improve.
FineWeb-Edu-High Configuration Comparison (click to expand)
| Parameter | Sonnet 4.0 | Sonnet 4.6 | Same? |
|---|---|---|---|
| ASPECT_RATIO | 32 | 21 | ❌ |
| MATRIX_LR | 0.047 | 0.068 | ❌ |
| EMBEDDING_LR | 0.375 | 0.60 | ❌ |
| SCALAR_LR | 0.39 | 0.18 | ❌ |
| WEIGHT_DECAY | 0.10 | 0.08 | ≈ |
| WARMDOWN_RATIO | 0.48 | 0.75 | ❌ |
| FINAL_LR_FRAC | 0.085 | 0.05 | ❌ |
| ADAM β1 | 0.66 | 0.45 | ❌ |
| ADAM β2 | 0.95 | 0.964 | ❌ |
| WINDOW_PATTERN | SSSL | SSSS | ❌ |
| MLP_RATIO | 4.25 | 4.0 | ❌ |
Every single parameter differs. Yet Sonnet 4.6 achieves 0.9% better val_bpb — the largest absolute gap in any comparison.
Haiku 4.5 started from Sonnet 4.6's accumulated best config (AR=21, β2=0.964, SCALAR_LR=0.18), not from stock defaults. This means:
| Starting Point | S4.0 | S4.6 | H4.5 |
|---|---|---|---|
| Config source | Prior characterization | Reset defaults | S4.6 accumulated |
| ASPECT_RATIO | 32 | 32 | 21 |
| ADAM β2 | 0.95 | 0.95 | 0.964 |
| Config quality | Pre-optimized | Stock | Fully optimized |
This is consistent with the sequential methodology (each model inherits from the prior run's state), but means Haiku's absolute val_bpb isn't directly comparable to the Sonnets' without accounting for baseline advantage.
| Metric | Sonnet 4.0 (Mar 19) | Sonnet 4.6 (Mar 22) | Haiku 4.5 (Mar 25) | Winner |
|---|---|---|---|---|
| Best val_bpb | 1.2959 | 1.2997 | 1.2924 | Haiku 4.5 |
| Baseline | 1.2969 | 1.3213 | 1.2966 | H4.5 (inherited) |
| Improvement | −0.08% | −1.63% | −0.32% | Sonnet 4.6 |
| Keeps | 1 (1.0%) | 8 (6.8%) | 1 (4.5%)* | Sonnet 4.6 |
| Crashes | 11 (11.0%) | 5 (4.2%) | 2 (9.1%)* | Sonnet 4.6 |
| Experiments | 101 | 119 | 22* | — |
| First keep | exp84 | exp25 | exp15 | Haiku 4.5 |
* In progress — 22/101 experiments completed.
Haiku found its first keep faster than both Sonnets (exp15 vs exp25 vs exp84) via SCALAR_LR tuning. Its absolute best (1.2924) is the new ClimbMix record, but this largely reflects the strong inherited baseline rather than optimization depth. After the single keep, all 7 subsequent experiments were discards — Haiku may be stuck near a local optimum.
| Dataset | S4.0 Best | S4.6 Best | Gap | Winner |
|---|---|---|---|---|
| ClimbMix | 1.2959 | 1.2997 | +0.0038 | S4.0 |
| FineWeb-Edu | 1.3424 | 1.3416 | −0.0008 | S4.6 |
| Cosmopedia-v2 | 0.9606 | 0.9549 | −0.0057 | S4.6 |
| SlimPajama | 1.5259 | 1.5267 | +0.0008 | S4.0 |
| FineWeb-Edu-High | 1.3463 | 1.3345 | −0.0118 | S4.6 |
Sonnet 4.0's wins are narrow (0.0038 and 0.0008) and benefited from pre-optimized baselines. Sonnet 4.6's wins are larger (up to 0.0118 on FineWeb-Edu-High) and came despite starting from worse baselines.
| Dataset | Sonnet 4.0 | Sonnet 4.6 | Reduction |
|---|---|---|---|
| ClimbMix | 11.0% | 4.2% | 2.6× fewer |
| FineWeb-Edu | 2.3% | 1.0% | 2.3× fewer |
| Cosmopedia-v2 | 1.9% | 0.0% | ∞ (zero crashes) |
| SlimPajama | 0.0% | 2.0% | S4.0 wins |
| FineWeb-Edu-High | 5.0% | 2.0% | 2.5× fewer |
| Total | 20 / 494 (4.0%) | 10 / 521 (1.9%) | 2.1× fewer |
Sonnet 4.6 crashes half as often overall. The sole exception is SlimPajama, where Sonnet 4.0 had zero crashes.
| Dataset | S4.0 params changed | S4.6 params changed |
|---|---|---|
| ClimbMix | 1 | 3 |
| FineWeb-Edu | 6 | 9 |
| Cosmopedia-v2 | 4 | 10 |
| SlimPajama | 3 | 4 |
| FineWeb-Edu-High | 8 | 10 |
| Average | 4.4 | 7.2 |
Sonnet 4.6 consistently explores 60% more parameter dimensions and builds compositional improvements across all datasets.
| Dataset | S4.0 AR | S4.6 AR | Agreement? |
|---|---|---|---|
| ClimbMix | 32 | 32 | ✅ |
| FineWeb-Edu | 32 | 32 | ✅ |
| Cosmopedia-v2 | 32 | 21 | ❌ |
| SlimPajama | 32 | 21 (inherited) | ❌ (different defaults) |
| FineWeb-Edu-High | 32 | 21 (inherited) | ❌ (different defaults) |
Sonnet 4.6's Cosmopedia-v2 AR=21 discovery propagated to subsequent runs via defaults. Neither SlimPajama nor FineWeb-Edu-High tried to change it — both explored narrower (AR=14, 16) and wider (AR=24, 26, 28) alternatives and found AR=21 near-optimal.
Sonnet 4.6 wins 3–2 on absolute val_bpb, but the margins are small on 4 of 5 datasets. The optimization landscape has natural performance floors that both models can approach. The exception is FineWeb-Edu-High, where Sonnet 4.6's β2 walk found a clearly better optimum.
By every meta-metric, Sonnet 4.6 outperforms:
- Crashes: 1.9% vs 4.0% (2.1× fewer)
- Parameter dimensions explored: 7.2 vs 4.4 per run (60% more)
- Compositional optimization: Consistently builds synergistic multi-parameter improvements
- β2 discovery: Systematic β2 walks (never attempted by Sonnet 4.0) produced keeps on 3 of 5 datasets
Sonnet 4.0 found AR=32 optimal on all 5 datasets. Sonnet 4.6 found AR=21 optimal on Cosmopedia-v2 and the result propagated — neither SlimPajama nor FineWeb-Edu-High reverted it. The "hardware-optimal architecture" depends on the LLM's exploration strategy, not just the hardware.
On FineWeb-Edu, both models find the same val_bpb via completely different configs (8/10 parameters differ). On FineWeb-Edu-High, every single parameter differs yet Sonnet 4.6 still beats Sonnet 4.0. The optimization landscape has multiple basins at similar or different depths.
Both models find SlimPajama nearly impervious to optimization (0.30–1.0% improvement, 2–3 keeps). Its 7-source diversity creates a flat optimization landscape where no parameter adjustment can meaningfully improve the model's ability to compress such varied data in 5 minutes.
-
Would AR=21 improve Sonnet 4.0's results? Sonnet 4.0 never tested AR<32 on any dataset. Running Sonnet 4.0 with AR=21 on Cosmopedia-v2 would reveal whether the AR=21 advantage is model-specific or universal.
-
Can we combine the best insights from both models? The models found complementary strategies (e.g., low WD vs low momentum on FineWeb-Edu). A hybrid config might beat both.
-
Is Sonnet 4.6's β2 walk transferable to Sonnet 4.0 datasets? The β2 walk was Sonnet 4.6's signature discovery. Testing β2<0.96 on Sonnet 4.0's best configs could unlock further improvements.
-
What would a third LLM generation find? ➜ In progress. Haiku 4.5 is now running across all 5 datasets. Early ClimbMix results show it can refine an already-optimized config (SCALAR_LR tuning), but hasn't yet discovered novel parameter regions. See Haiku 4.5 ClimbMix (Mar 25).
See individual run pages: H4.5 ClimbMix (Mar 25) | S4.6 FineWeb-Edu-High (Mar 25) | S4.6 SlimPajama (Mar 25) | S4.6 Cosmopedia-v2 (Mar 24) | S4.6 FineWeb-Edu (Mar 22) | S4.6 ClimbMix (Mar 22) | S4.0 FineWeb-Edu-High (Mar 21) | S4.0 SlimPajama (Mar 20) | S4.0 Cosmopedia-v2 (Mar 20) | S4.0 ClimbMix (Mar 19) | S4.0 FineWeb-Edu (Mar 17) | Cross-Dataset Comparison
