v2.5.0
What's Changed
- Rename
scatter_objecttobroadcast_object()by @epwalsh in #471 - Add exponential LR scheduler by @baileykuehl in #472
- Fix typos in code by @agupta2304 in #473
- Add Peri-LN transformer block, module for residual stream by @epwalsh in #457
- Fix QK norm parameter count in AttentionConfig.num_params() by @dirkgr in #477
- Update docker images and versions (most recent image - torch 2.9.1, cuda 13.0) by @tyler-romero in #459
- Internal olmo3 32b midtrain/long context scripts by @tyler-romero in #450
- Use correct peak LR in Olmo 3 32B midtrain by @soldni in #481
- Load multiple checkpoints at the same time and average them by @dirkgr in #396
- Fix type annotation issue in NumpyInterleavedFSLDataset (Issue #486) by @ada-ggf25 in #487
- Composable data loading by @epwalsh in #416
- [MoE] Upstream TrainModuleConfig by @tyler-romero in #490
- various improvements to set up for new model ladder by @epwalsh in #493
- set
num_execution_unitsdefault to 1 by @epwalsh in #495 - split off transformer and attention CPU tests by @epwalsh in #496
- Changes the prefix used by
GPUMemoryMonitorCallbackto avoid using the reservedsystemprefix by @finbarrtimbers in #497 - Avoid torch recompile when intra-doc masking enabled by @epwalsh in #504
- Gated Attention by @tyler-romero in #508
- Model ladder, revamped by @epwalsh in #494
- No-global-rope (GNoPE) by @tyler-romero in #509
- Fix metrics filter by changing
system/prefix to newgpu_memory/prefix by @finbarrtimbers in #499 - Add in-loop LM evals to ladder by @epwalsh in #512
- Bottom-up flops estimation & total flops tracking by @tyler-romero in #511
- MXFP8 Linear Support by @tyler-romero in #510
- ladder lmevaluator fix by @tyler-romero in #514
- Fix overflow when too many global flops are computed by @tyler-romero in #515
- Jacobm final merge sft by @jacob-morrison in #400
- Add instance filter to olmo3 ladder by @tyler-romero in #516
- Make data loading preprocessing more robust to race conditions by @epwalsh in #517
- Add scripts for slurm env by @epwalsh in #518
- update slurm scripts by @epwalsh in #519
- Allow cordoning nodes by @epwalsh in #520
- More slurm improvements by @epwalsh in #521
- Add temp & mem healthchecks and allow comments in cordoned nodelist by @tyler-romero in #522
- Consolidate bash helper functions by @epwalsh in #523
- Increase ulimit -n within slurm jobs by @tyler-romero in #524
- More Slurm util improvements by @epwalsh in #525
- Ladder improvements by @epwalsh in #529
- GAP Monitor Improvements by @tyler-romero in #527
- Ulysses CP Support by @tyler-romero in #532
- Support for Muon and Dion Optimizers by @tyler-romero in #528
- Add a model ladder for "more norms" and other internal ladder improvements by @epwalsh in #535
- Add 60M and 1M param model configs by @tyler-romero in #536
- Add Qwen3 model support (0.6B-32B) by @finbarrtimbers in #533
- update to Beaker-py v2, use gantry API internally by @epwalsh in #540
- Add 60M and 100M param ladder rungs by @tyler-romero in #543
- always include GOOGLE_CREDENTIALS by @epwalsh in #545
- Ensure final metrics always logged by @epwalsh in #546
- Log chinchilla multiple from
SpeedMonitorCallbackby @epwalsh in #548 - improve beaker client caching by @epwalsh in #549
- Add a stepped WSDS variant, integrated into ladder by @epwalsh in #538
- update base LR for stepped schedule by @epwalsh in #551
- Add Gemma 3 support with HF weight conversion by @finbarrtimbers in #534
- Add 14M param config by @tyler-romero in #552
- Fix flaky and broken transformer model tests by @tyler-romero in #537
- Add model ladders for recently run ablations by @tyler-romero in #553
- Fix 1M model by @tyler-romero in #555
- Now,
WandBCallbackandCometCallbacklog the Beaker URL to their config sections. by @finbarrtimbers in #547 - Skip init_process_group if already initialized by @finbarrtimbers in #539
- Fix failing test_build_world_mesh_cpu for pytorch 2.10 by @tyler-romero in #559
- Easier discovery and collection of ladder results by @tyler-romero in #556
- Reduce disk usage of convert_checkpoint_to_hf_test by @tyler-romero in #561
- More ladder work by @epwalsh in #560
- Add torch 2.10 image, fix FA3/TE incompatibility by @epwalsh in #562
- Add support for FA4 (CUTE implementation) by @epwalsh in #563
- Use uv in CI by @epwalsh in #564
- Add
Callback.pre_log_metrics()method by @epwalsh in #566 - Replace OmegaConf with dataclass-extensions by @epwalsh in #567
- Utilize Registrable for optim/scheduler configs by @epwalsh in #568
- Integrate QuACK kernels, minor improvements to Docker build by @epwalsh in #569
- SequenceMixer as base for both attention and recurrent layers by @tyler-romero in #570
- update deps by @epwalsh in #574
- Add GatedDeltaNet by @tyler-romero in #572
- Add option to load downstream eval tasks lazily by @epwalsh in #575
- [QoL] include throughput metrics when updating Beaker workload description by @epwalsh in #577
- Ensure Beaker workloads submitted from CI are canceled gracefully by @epwalsh in #579
- Evaluate upon completion of training by @tyler-romero in #580
- Fix documentation link for OLMo-core Beaker images by @natolambert in #581
- Update SFT documentation with tokenization and training tips by @natolambert in #571
- Join bookkeeping ops before checkpointing by @epwalsh in #583
- Propagate gantry launch options by @tyler-romero in #582
- Fix some logic around soft timeouts in
launch.beakerby @epwalsh in #584 - Make SLACK_WEBHOOK_URL not actually required by @tyler-romero in #585
- Optional disk-backed caching for remote IO ops, data loading config improvements by @epwalsh in #586
- Add chat template verification step to SFT README by @natolambert in #590
- Checkpointing improvements by @epwalsh in #588
- GatedDeltaNet improvements by @tyler-romero in #592
- Bump dataclass-extensions to fix pickle issue by @epwalsh in #595
- Fix config encoding/merging with non-init fields by @epwalsh in #596
- Internal cleanup of composable sampling classes by @epwalsh in #599
- Fix config decode when CLASS resolves a Registrable subclass by @baileykuehl in #600
- port over @dirkgr's fix for async callbacks by @epwalsh in #601
- Fix GDN initialization by @tyler-romero in #602
- Mark ephemeral checkpoints in metadata by @epwalsh in #605
- Add claude md for Olmo-core by @tyler-romero in #611
- Support flattening 3d parameters to 2d for use with Muon by @tyler-romero in #606
- docs: SFT conversion tips for tokenizer override and parallel jobs by @natolambert in #608
- updates to claude.md file by @epwalsh in #612
- Support block-pattern based initialization of hybrid transformers by @tyler-romero in #607
- Ladder5 v2 by @dirkgr in #610
- Add model merging callback with integration test by @baileykuehl in #558
- Fix A100 peak flops spec in speed monitor by @tyler-romero in #614
- Add token ID validation guard in DataCollator by @tyler-romero in #619
- fix typo in mixes init by @baileykuehl in #621
- Support perplexity evals with context and tensor parallelism by @tyler-romero in #616
- Add training smoketest skill for claude by @tyler-romero in #618
- Hybrid model pre/mid/lc configs by @tyler-romero in #613
- Add midtraining ingredient to to hybrid checkpoint list by @tyler-romero in #633
- Add new in-loop eval tasks and update eval dependency to PyPI release by @kyleclo in #587
- Full support for flash attention 4 (including generation kv cache support, installation from pypi) by @tyler-romero in #642
- Add documentation for native generation module and chat interface by @tyler-romero in #641
- Add Code Fresh per-language perplexity eval by @kyleclo in #639
- fix attn backend forcing for TransformerGenerationModule by @tyler-romero in #645
- ✨ Quality: Add documentation about checkpoint conversion limitations by @lukebaze in #648
New Contributors
- @agupta2304 made their first contribution in #473
- @ada-ggf25 made their first contribution in #487
- @finbarrtimbers made their first contribution in #497
- @jacob-morrison made their first contribution in #400
- @natolambert made their first contribution in #581
- @kyleclo made their first contribution in #587
- @lukebaze made their first contribution in #648
Full Changelog: v2.4.0...v2.5.0