Skip to content

Releases: allenai/Olmo-core

v3.0.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 15:17

What's new

Added πŸŽ‰

  • Added use_array_if_local to pack_documents_into_instances, segment_documents_into_instances, NumpyPackedFSLDataset and NumpyPackedFSLDatasetConfig, forwarded to iter_document_indices. Set it to False to take document boundaries from the source metadata file instead of inferring them by scanning the token array for the EOS token. Inferring is only correct when every document is EOS-terminated: a producer that truncates documents and drops the terminator with the tail causes the affected document to merge with the one after it, and LongDocStrategy.truncate then keeps only the head of the merged span, so the following document never reaches training. Measured on an SFT cache, 97.98% of tokens reached instances via the inferred path versus 100.00% via the metadata file, with an identical maximum document length. Whenever the metadata boundaries are the effective source -- set explicitly, or because the source is a URL, for which iter_document_indices always reads the metadata -- doc_lens is derived from them rather than by rescanning the packed tokens for EOS, so the block-diagonal attention mask cannot merge an unterminated document into the one after it. The hazard is now documented on iter_document_indices. Default behavior is unchanged.
  • Added opt-in paired SwiGLU backward and BF16-rounded weight-gradient accumulation for OLMoDDP experts, including Torch 2.13 support, checkpoint/recomputation coverage, and explicit backend and bucket-ownership guards.
  • Added independent per-head Q/K norm gains and scalable softmax, EMO document-pool routing/global load balancing, and opt-in FP32 gradient-accumulation/reduce-scatter fast paths with explicit hardware/version guards.
  • Extended hybrid MoE HF export for KDA, optional EMO and latent experts, per-head normalization gains, and scalable softmax, with exact tensor round-trip validation and legacy configuration migration.
  • Added an optional, dependency-free checkpoint-ready notification callback for independent upload services. It does not upload or delete checkpoints.

Removed πŸ‘‹

  • Removed the src/scripts/train/private-olmo.py training script.

  • Removed the olmo_core.model_ladder API, its internal CLI wrapper, all eleven ladder training scripts (including the standalone Gemma-like ladder), two Slurm launchers, and the API documentation. Existing ladder orchestration configs and commands require an earlier revision; this does not change the model or optimizer checkpoint formats. The internal experiment framework remains available.

Fixed βœ…

  • Metadata-backed packed datasets now include sidecar content hashes in packing-cache keys and dataset fingerprints, invalidating stale boundaries even after same-size corrections. Document lengths preserve EOS/BOS padding segmentation. Local array-backed defaults are unchanged (#843).
  • Apply opt-in Q/K gain expansion to eval-only and model-only DDP checkpoint loads, and reject forced expert assignments and biased KDA convolutions during HF export.
  • Validate normalization throughout MoE HF exports, preserve attention-only gates and resolved EOS/padding IDs, and reject unsupported shared-expert routing before conversion.
  • Reject MoE HF exports with incompatible Q/K normalization or inconsistent KDA output-norm epsilons instead of silently changing normalization behavior.
  • Preserve the resolved tokenizer BOS ID in exported model and generation configs, including when the training config leaves BOS unspecified.
  • Reject hybrid KDA HF exports with sliding-window attention and EMO exports with restricted evaluation pools in any routed layer. Disable scalable-softmax HF generation caching and reject explicit cache use.
  • Require the CUDA 13 CuTe compiler for experimental KDA, report incompatible installs before training, and exercise the kernels in a dedicated Blackwell CI job.
  • Keep legacy fused attention configs compatible when the new attention options are disabled. Reject unsupported scalable-softmax context parallelism and KV caching at setup.
  • Assign EOS tokens to their preceding document for EMO routing, matching attention document boundaries.
  • Preserve serialized tokenizer behavior during HF export, including source BOS settings when the training config leaves BOS unspecified.
  • Preserve FP32 router probabilities and accumulation when combining BF16 experts in the HF reference model.
  • Avoid graph breaks from no-op profiling decorators, including compiled router load-balancing collectives.
  • dispatch_flash_attn_4 passed cu_seqlens_q, cu_seqlens_k, max_seqlen_q and max_seqlen_k positionally. flash-attn 4 inserted a qv parameter at position 3 (present from ~4.0.0b19 onward), which shifts every following argument by one, so max_seqlen_q β€” an int β€” lands where cu_seqlens_k is expected and the call fails with AttributeError: 'int' object has no attribute 'shape'. These are now passed by keyword; the parameter names are unchanged across flash-attn 4 releases, so this is correct against both old and new versions. Only reachable on Blackwell, where has_flash_attn_4() returns True.

Changed ⚠️

  • OLMoDDPModel.apply_ddp() now raises NotImplementedError directing callers to apply_dp(), including under python -O. Removed its unreachable legacy DDP implementation.

  • OLMoDDPTrainModuleConfig.max_grad_norm now overrides the optimizer clipping threshold when set, matching the train-module configuration API used by TransformerTrainModuleConfig. When unset, the optimizer threshold is retained. Previously this field was ignored, so old configs with differing values now use the train-module value. The supplied optimizer config is not mutated. Removed the unused OLMoDDPTrainModule.max_grad_norm constructor argument and attribute; clipping remains inside the DDP optimizer.

  • Removed the unused, commented-out Beaker execution-unit helper and its commented call from the internal experiment setup.

  • Migrated repository attention callers to explicit backend selection. Transformer factories now resolve legacy use_flash arguments into backend configs, and the nGPT factory accepts attn_backend. Deprecated inputs remain supported for old configs and callers, including explicit-backend precedence and sliding-window automatic selection.

  • Replaced FusedAttention with FusedAttentionV2 and migrated its applicable tests to V2. Legacy AttentionConfig(name="fused") configs still load through V2 with the same packed QKV parameter names and shapes, preserving model and optimizer checkpoint compatibility. Legacy fused RoPE is mapped to regular RoPE: floating-point results (especially low-precision RoPE gradients) differ, so resumed training is not numerically identical. TransformerConfig.llama_like(fused_ops=True) now selects V2 with regular RoPE when it previously selected fused attention.

  • Raised the minimum supported PyTorch version to 2.10.0. Updated the stable Beaker images to PyTorch 2.10 with CUDA 12.8 and PyTorch 2.11 with CUDA 13.0, and expanded CI coverage to include PyTorch 2.10 and 2.11 with CUDA 12.8 and Docker builds through PyTorch 2.12 with CUDA 13.0.

Commits

a800c06 (chore) prepare for release v3.0.0
3c1a108 Add third party notice (#900)
56c04b8 Add tech-report reference scripts to src/examples/olmo_ddp (#899)
6ac500d Allow document boundaries to come from the metadata file (fixes silent token loss on non-EOS-terminated documents) (#843)
718d08f Wire DDP train-module gradient clipping and remove retired apply_ddp code (#881)
5e607f8 Remove model ladder framework and retired training scripts (#879)
306d4b3 Migrate use_flash callers to explicit attention backends (#878)
f6143b5 Add hybrid MoE attention, EMO routing and HF export (#872)
845bda9 Replace FusedAttention with FusedAttentionV2 (#873)
0589a03 Pass flash-attn 4 varlen arguments by keyword (#857)
bbe9473 Caleb/kda cute kernels (#863)
f726f96 Add the OLMoDDP fused MoE / expert-parallel training stack (#844)
92870a3 Raise minimum PyTorch version to 2.10 (#835)

v2.6.0

Choose a tag to compare

@github-actions github-actions released this 11 Aug 15:39

What's new

Added πŸŽ‰

  • Added max_checkpoints parameter to CheckpointerCallback (default: 3) to limit the number of permanent checkpoints retained. Oldest checkpoints are removed automatically when the limit is exceeded. Set to None to keep all (previous behavior).
  • Added OutputDiscardCheckpoint, an activation-recompute primitive for cases where the output of a checkpointed region dominates memory rather than its intermediates (e.g. precision casts, FFN up-projections). Forward runs under no_grad, the output's storage can be freed after downstream consumption, and a backward hook recomputes and rebinds the freed storage in place via a C++ share_storage extension (with a Python fallback for environments without a C++ toolchain).
  • Added Qwen3.5 dense model configs (0.8B, 4B, 9B, 27B) with hybrid Gated DeltaNet + full-attention architecture.
  • Added partial RoPE support via partial_rotary_factor on :class:~olmo_core.nn.rope.RoPEConfig.
  • Added HuggingFace weight conversion for qwen3_5_text hybrid models.
  • Added a configurable vision transformer encoder (VisionTransformer, configured via VisionEncoderConfig), vision-to-LM connector (VisionConnector), and MultimodalLM β€” a composite vision-language model that fuses image patch tokens into the LM token stream. Supports OpenAI CLIP, SigLIP, and SigLIP2 encoder variants with factory configs for all standard Molmo2 checkpoints.
  • Added HFConverterCallback, which can be used to convert models to huggingface format at the end of the training run.
  • Trainer now records checkpoint save and load durations as train/checkpoint_save_duration_s and train/checkpoint_load_duration_s metrics.
  • Added PowerLR, a power-law learning rate scheduler with linear warmup, power-decay phase (lr = initial_lr * (current / warmup) ** b for negative b, making the LR independent of the training horizon), and an optional linear decay tail. Registered as "power_lr".
  • Added ComposableScheduler, a piecewise LR scheduler built from ComposableSchedulerStage segments (linear/cosine interpolation between endpoint LRs) on an absolute time axis. Registered as "composable". Note: ComposableScheduler ignores the t_max passed to get_lr and emits a once-per-instance UserWarning to that effect.
  • Added OverrideDecay, a late-stage decay override usable on both ComposableScheduler and SequentialScheduler via an override_decay field. When current >= override_decay.start, the main schedule is interrupted mid-flight and the LR decays from the value the main schedule would have produced at start to a target LR over duration (linear or cosine). SequentialScheduler additionally warns that t_max is ignored once the override becomes active.
  • OLMO_RICH_LOGGING can now explicitly enable or disable rich console logging (0/false/no/off disables it); previously setting it to any value only force-enabled rich logging.
  • init_distributed() now bootstraps a minimal single-process environment (RANK=0, WORLD_SIZE=1, MASTER_ADDR/MASTER_PORT) when launch env vars are absent, so scripts can be run directly (without torchrun) for single-process debugging.
  • Added a configurable determinism_check option to activation checkpointing (default "default"); set it to "none" to skip torch's recompute metadata check for opaque linear-attention kernels under torch.compile.

Fixed βœ…

  • The CPU Test CI job now caches HF_HOME across runs so the HuggingFace roundtrip tests (Qwen3-0.6B, Gemma-3-270m) don't re-download their checkpoints every run.
  • Excluded mark_dynamic from torch.compile tracing (@torch.compiler.disable).
  • Clearer error messages (now include the offending values) when a rank batch size isn't divisible by the sequence length, or max_target_sequence_length isn't a multiple of sequence_length.
  • S3 uploads/downloads now also retry on transient SSL errors (ssl.SSLError, botocore/urllib3 SSLError).
  • Distributed checkpoint writes now clone each tensor before serialization to avoid accidentally writing the full backing storage of a view/shared tensor, with a guard that raises OLMoCheckpointError if a written tensor is unexpectedly larger than its nbytes.
  • Fixed LM in-loop evaluator data-order drift across repeated runs by resetting loader bookkeeping before each pass and making deterministic reshuffling the default.
  • Fixed Qwen3 implementation to match HuggingFace by applying RoPE in the input dtype (bf16) rather than upcasting to fp32.
  • Fixed HF model conversion for Llama, Qwen3, and Gemma so that converted checkpoints roundtrip correctly.
  • Fixed Beaker secret existence check to use the case-insensitive HTTP endpoint, avoiding spurious "secret not found" errors when secret names differ only in case.
  • Fixed Transformer.init_weights so that under interleaved pipeline parallelism (e.g. Interleaved1F1B, InterleavedZeroBubble) the multiple model chunks owned by a single rank no longer initialize to identical parameters. Adds a model_part_idx kwarg incorporated into the seed as model_part_idx * pp_size.
  • Disabled torch.compile tracing through TEAttentionBackend.forward, whose Python/pybind setup is not Dynamo-safe.
  • Fixed TransformerPipelineTrainModule.num_flops_per_token returning None under pipeline parallelism. Each PP rank only holds its stage's layers, so summing FLOPs from model_parts undercounts the model. Capture model.num_flops_per_token as a bound method before split_model deepcopies and drops layers, then call it at metric time. On meta device (the standard PP init path) this has no memory cost.

Changed ⚠️

  • Set transformers version to >= 5.4.0 for Qwen 3.5 and in sync with open-instruct
  • Added a documented deterministic option to LMEvaluator and LMEvaluatorCallbackConfig so callers can opt out of fixed eval ordering when desired.

Commits

b7e9671 (chore) prepare for release v2.6.0
77714b7 fix dion pypi (#816)
5b0d1cf (chore) prepare for release v2.6.0
66f768b Update anonymous paths in public scripts to point to data in hugging face bucket (#802)
064b172 bump fla to 0.5.2 (#798)
d3146cc Misc fixes (#800)
fa6c501 Add max_checkpoints to limit permanent checkpoint retention (#694)
c3802ed CI: cache HuggingFace models for the CPU Test job (#732)
2000b1b Add OutputDiscardCheckpoint (#682)
9aa3280 Add Qwen3.5 model support (0.8B, 4B, 9B, 27B) (#684)
36f99f0 Add conversion overrides for Llama, Qwen3, and Gemma 4 models so they roundtrip properly (#677)
885219b Add configurable determinism_check to activation checkpointing (#713)
8d22ca9 [1/n] Add vision transformer, connector, and MultimodalLM (#692)
59a339f Misc training-utility improvements (#691)
1713ea3 Improve checkpoint/S3 IO robustness (#690)
754d58d CI: run GPU tests at low priority (#696)
1af17a4 Fix PP FLOPs: capture full model before pipeline split (#680)
525cc25 Increase tolerance for flaky test (#681)
38704d1 Two small transformer core fixes: TE Dynamo + PP init seed (#679)
73637f7 Add PowerLR scheduler (#674)
2caaee9 Add ComposableScheduler (#671)
2e67bcb Update task timeout for GPU tests (#675)
e556a86 Revert "Add position_ids-based varlen RoPE support for packed inputs" (#672)
1eec696 Add position_ids-based varlen RoPE support for packed inputs (#654)
60d2487 Fixes the Qwen3 implementation to match HF (#663)
5e7ee43 Record checkpoint save/load durations as trainer metrics (#665)
3622318 Fix broken path in 32B LC (#670)
3e19fa2 Fix breaking paths (#669)
53c51c5 Use case-insensitive HTTP endpoint to check beaker secret existence (#666)
afe99b6 Pin flash-linear-attention version to 0.4.1 (#664)
2e57086 Add HFConverterCallback for end-of-training HuggingFace conversion (#660)
beca1f1 Correct dolmino mix urls (#630)
b376077 Fall back to anonymous GCS client when no credentials are available (#659)
60930ef Add gradient dumping to GAPMonitorCallback (#438)
befb60b make lm evaluator data deterministic (#652)

v2.5.0

Choose a tag to compare

@tyler-romero tyler-romero released this 03 Apr 17:32

What's Changed

Read more

v2.4.0

Choose a tag to compare

@github-actions github-actions released this 20 Nov 19:15

What's new

Added πŸŽ‰

  • Added option to skip ranges of steps in the trainer.
  • Send a Slack notification when a Beaker job appears to be stuck.
  • Added ignore_fingerprint_mismatch parameter to NumpyDataLoaderConfig to allow resuming training from a checkpoint with a different dataset mix.
  • Added helpful error messages when OLMo-mix-0625 files are not found, directing users to use OLMo-mix-0925 and the fingerprint override flag.
  • Added olmo_core.generate.chat module to allow interacting with OlmoCore models without conversion to other formats.
  • Added GAPMonitorCallback for monitoring gradients, activations, and parameters (GAP).
  • Added official Olmo 3 7B and 32B pretraining scripts and data mix.
  • Added official Olmo 3 7B and 32B midtraining scripts and data mix.
  • Added official Olmo 3 7B and 32B long-context scripts and data mix.
  • Added a NoOpOptimizer that does nothing, uses no memory, and can be used for debugging.
  • Added official config for Olmo 3 32B.
  • Olmo 3 model card and checkpoint manifests.

Fixed βœ…

  • Set missing NCCL_NVLSTREE_MAX_CHUNKSIZE env var that is now needed for running jobs on Augusta cluster.
  • Fixed bug with RemoteFileSystemReader that caused excess memory usage.
  • No longer overrides random's RNG seed when building SourceMixtureDatasetConfig.
  • Fix handling URLs in olmo_core.nn.hf.checkpoint.save_hf_model and in examples/huggingface.
  • Fix potential NaN loss that can occur when using instance masking.
  • Stability improvements developed while training Olmo3 32B.

Changed ⚠️

  • Removed unused field in YaRNRoPEScalingConfig.

Commits

1ed8900 (chore) prepare for release v2.4.0
2c179c2 (chore) prepare for release v2.4.0 (#467)
7e0431f Fix link to 7B midtrain script (#469)
843fe3d Olmo3 model cards, checkpoint manifest, and readme (#468)
cbdc2f1 Olmo3 32B cleanup and checkin (#460)
20548a0 Official Olmo3 32B long-context script (#465)
14b15cc Official Olmo3 32B midtrain script(s) and mix(es) (#466)
a25a514 Official Olmo3 32B pretrain config and data mix (#464)
6b73ba0 Official Olmo3-7B long context script (#458)
bdc61e4 Official Olmo3-7B midtraining scripts (#445)
55804bf 32B official config (#454)
a86131d Slight refactor of Yarn Scaling Config (#456)
68c7409 Handle target URLs properly in HF conversion (#453)
2504cc2 Instance mask correction to avoid nan loss (#452)
137274e Add callback to monitor grads, activations, params (#446)
0959a54 Improve mem usage of RemoteFileSystemReader (#451)
600d2fe Official Olmo3-7B pretraining scripts (#443)
accc310 make launch timeout configurable from CLI
aa0e629 Avoid overriding RNG seed when building SourceMixtureDatasetConfig (#449)
98ba2e4 NoOp optimizer (#444)
03e6836 OlmoCore native chat interface (#439)
7a0bbd7 unset 2 NCCL env vars per Google's recommendation
aacb6eb only send local Slack notifications when callback is enabled (#441)
bfc8d7a Min python version to 3.10 (#442)
96d43d4 Set missing NCCL_NVLSTREE_MAX_CHUNKSIZE env var (#440)
5ad6db5 hot fix for listing gcs dirs
2186957 Allow manual bypass of fingerprint mismatch when switching datasets (#435)
043505d hot fix to step regex
dd7e747 Send a Slack notification when a Beaker job appears to be stuck (#431)
e27a9b4 Add WSDS (Warmup-Stable-Decay-Simplified) Scheduler (#419)
c92320f Use a dataclass for 'Trainer.steps_to_skip' (#430)
9669268 clean up checkpointing code to minimize distributed communication (#428)
87d64b9 fix changelog
269bf02 Add option to skip ranges of steps in the trainer (#425)

v2.3.0

Choose a tag to compare

@github-actions github-actions released this 17 Oct 16:22

What's new

Fixed βœ…

  • Fixed parsing username+password git remote URLs in launch.beaker module.
  • Fixed bug with default setup steps in launch.beaker.BeakerLaunchConfig when a branch can't be resolved.
  • Cluster names in Beaker have changed.
  • Fixed mixture rounding error with SourceMixtureDataset, which was previously causing samples to be repeated at the end of training.
  • Don't DDOS Beaker from big jobs.
  • A configuration error is now raised if you pass in a URL for the trainer or dataset's working directory.
    Previously the URL would just get mangled into a local path, leading to unexpected behavior.
  • Fixed an issue where the ConsoleLoggerCallback would attempt to log before the first step.
  • Only call teardown_distributed_environment() when training ends cleanly to avoid a hang for the duration of the distributed backend's timeout when there's an error from one rank.
  • Fixed tensor parallelism issue with torch 2.8.
  • More fixes for Beaker cluster names.
  • Callback.post_train() will still be called even if the run is canceled before the dry-run batch.
  • GarbageCollectorCallback will restore gc settings even when Trainer.fit() exits on an error.
  • Make move_to_device blocking for MPS device to fix possible incorrect transfer of data from CPU to MPS.
  • Fixed bug where glob_directory() would fail to match certain glob patterns.
  • Added one more type of error to retry on when the Google Storage API throws it.
  • Perform a garbage collection after checkpointing to avoid running out of CPU memory.
  • Avoidable overflow error when using NumpyPackedFSLDataset.
  • Fixed issue with NumpyFSLDatasetMixture + SourceMixtureDataset where not all instances would have the same sequence length.
  • Attention backend will no longer default to flash in non-CUDA environments.

Changed ⚠️

  • The dir option to Trainer.maybe_load_checkpoint() is now optional and defaults to the save_folder.
  • Set fused_linear_cross_entropy_loss accum_dtype to fp32 in LMHead.
  • Increased NCCL_FASTRAK_PLUGIN_ACCEPT_TIMEOUT_MS from 10 minutes to 30 minutes.
  • SlackNotifierCallback will now notify on checkpoint saved and post epoch events.
  • BeakerLaunchConfig.launch() will now send Slack notifications by default when follow=True if the env var SLACK_WEBHOOK_URL is set.
  • src/examples/llama/ has been renamed to src/examples/llm/.
  • Refactored eval task groups into task_groups.py
  • The use_flash argument to the Attention classes is deprecated. Use backend="flash_2" instead.
  • Refactored NumpyDatasetConfig by splitting it into a separate config per underlying dataset class.
  • Refactored internal/experiment module to facilitate modifying datasets or supplying a fully custom ExperimentConfig.
  • Simplified SourceMixtureDatasetConfig by removing redundant sequence_length and dtype fields.
  • The model_id argument to convert_state_from_hf is deprecated. Conversion information is deduced from the model type.
  • Refactored the example conversion scripts to/from HF, including decreasing false failures in validation.
  • Small refactor to source_mixture.py to make it easier to define data mixes in yaml.
  • Reorganized/cleaned up internal training scripts.

Added πŸŽ‰

  • Added CLI script src/scripts/unshard.py for converting distributed checkpoints to regular PyTorch or safetensors format.
  • Added a custom block that does LayerNorm scaling.
  • Added OLMo-mix-0625-150Bsample data mix.
  • Added alias support to DataMix enum.
  • Added the HalfCos learning rate scheduler.
  • Added CONTRIBUTING.md guidelines.
  • Added a lightweight, gantry-like Beaker launch CLI: python -m olmo_core.launch.beaker.
  • Added Beaker images with torch 2.8. There is olmo-core-tch280cu128-2025-09-18 and olmo-core-tch280cu129-2025-09-18 for CUDA 12.8 and 12.9, respectively.
  • Added TransformerEngine to Docker images and a TransformerEngine attention backend.
  • Added Callback.close() method, which is always called when exiting Trainer.fit().
  • Added flash-attention 3 to Docker images, added flash_3 attention backend.
  • Added support for sliding window attention to the Torch attention backend. Performance is not optimized, so other backends should be preferred.
  • Added RoPEScalingConfig.to_hf_config() for each RoPE scaling method to support automatic conversion to HuggingFace format.
  • Guide to dataset mixing in docs/source/guides/data_mixing.rst.
  • Added support for converting FlexOlmo models (with both dropless and default MoEs) between OLMo Core and HF formats.
  • Added olmo3_7B model config.
  • Added additional internal configuration tools.
  • Added a new named data mix that we used for the 32B run
  • Added internal OLMo3 7B midtraining and long-context configs.
  • Added ability to convert OLMo3 models to/from HF format with support for rope scaling configs.
  • Added a script that can pull out a single training batch from a training job

Commits

5b32459 (chore) prepare for release v2.3.0
3f21b77 Reorganize internal scripts for Olmo3 (#423)
67dd1d6 Use exec to start torchrun (#422)
206d25e Cookbook migration part 2 - long context config (#404)
eabb869 Raise timeout error if training job doesn't start in 5 mins (#421)
cc69286 Script to dump training tokens (#418)
b5ba7be Script for unsharding (#420)
464d01e olmo3 conversion w/ support for rope scaling (#415)
3afa7dd Add Rope scaling configs to rope module's exports (#414)
3ef0c05 Cookbook migration part 1 - midtraining config (#403)
cdb7922 Add Dolma 3 sample, and a way to alias data mixes (#412)
77adc0b pull Slack webhook URL from secret if available (#413)
c8e32ab NumpyFSLDatasetMixture + SourceMixtureDataset fix (#411)
6ce62cc Fix OverflowError in pack_documents (#409)
e229c17 Cookbook migration part 0 - more setup (#397)
8eef6f3 Data mix for the 32B (#405)
9ba6154 Retry more on GS failures (#406)
73d11f8 Do GC after checkpointing (#407)
cdfd201 Typo
f477a8d Fix glob_directory bug (#402)
bc2a3f1 Support conversion of dropless MoE to FlexOlmo (#401)
7fc49de Support rope scaling configs for hf conversion (#394)
0f38dc6 HF Conversion Refactor (#390)
a53f825 source mixture dataset simplification and documentation (#399)
69bd9d2 MPS bug fixes (#395)
601d336 refactor internal experiment configuration assembler (#386)
874be5d Add flash_3 attention backend (#377)
3aac611 Remove old beaker refs (#393)
1aeb369 Add sliding window support for torch attention backend (#388)
7bebea3 Fully migrate to new cluster names, fix internal experiment launching on Augusta (#391)
1d4b67f Add flash-attn-3 install to dockerfile (#392)
04d5d5d Add Callback.close() method, other minor callback improvements (#389)
dcf3a0e Add attention backend abstraction and integrate transformer engine's attention (#384)
0356c1b Np Dataset Config Refactor (#381)
afb5827 Improve distributed error handling (#380)
45d2031 Consolidate CLI code in public scripts (#379)
cbcc735 Port task groups from cookbook to OlmoCore (#378)
71f0023 bump torch and other dependencies in our Docker build (#360)
b3c3b4e Add an all-in-one guide for researchers (#374)
65d411b CONTRIBUTING.md (#376)
e469561 Treat cordoned beaker hosts as occupied (#375)
75d4d3a Avoid DDOS-ing Beaker for real (#372)
258a7a8 fix
b068044 Set accum_grad to fp32 for fused_linear_cross_entropy_loss, bump liger-kernel version (#370)
b550f9b Add the ability to send Slack notifications from launch.beaker (#371)
801de4b Makes it possible to override the common config builder (#369)
59a2b2f Adds a new, highly specific LR scheduler (#368)
d4ad23f fix rounding error with mixing datasets (#316)
46f7211 [Feat] Add LNS training example (#320)
3d9e9cd redo base dir calculation for new cluster names (#364)
b852ec0 Up NCCL_FASTRAK_PLUGIN_ACCEPT_TIMEOUT_MS to 30 minutes (#365)
72b7c1b Moved google-cloud-compute dependency from dev to beaker group. (#363)
c00d715 catch issues with dolma metadata files earlier (#361)
fa5a5dc Refine hostname constraints for beaker experiments on Google clusters (#355)
386b0a8 Fix parsing username+password git remote URLs (#356)
6726ffc make release process more robust

v2.2.0

Choose a tag to compare

@github-actions github-actions released this 26 Aug 16:44

What's new

Added πŸŽ‰

  • Added option to set LR scheduler based on tokens instead of steps (e.g. --train_module.scheduler.units=tokens).
  • Added a "packed" numpy FSL variant that packs documents into sequences using the best-fit-decreasing bin packing algorithm following the work from Fewer Truncates Improve Language Modeling.
  • Added module olmo_core.testing.
  • Added a "interleaved" numpy FSL variant that interleaves several documents into sequences following the work from LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models.
  • Added sliding window attention as a feature
  • Added BatchSizeSchedulerCallback for setting a batch size schedule over the course of a training run.
  • Added optional TrainModule method, .pre_train(), which runs right after Callback.pre_train().
  • The BeakerCallback will save the config and Python requirements to the results dataset.
  • Added from_file method to Config class.
  • Added in-loop evals for OLMES basic skills eval
  • Added in-loop fast MCQA for in-loop evals and translated MBPP tasks
  • Added in-loop few-shot HumanEval BPB
  • Added fast and full in-loop recommendations, where fast is a roughly 2-3x faster subset of full
  • Added support for converting to HF models in lower precisions.
  • Added support for headwise QK norm.
  • Add BOS token in in-loop evals, when specified by the tokenizer (ai2-olmo-eval==0.8.4)
  • Add support for BOS token matching EOS token for intra-document masking in FSL numpy datasets.
  • Added option to allow profiler to record on multiple ranks.
  • Added support for accessing Google on non-Google clusters via auth with service account keys.
  • Added support for revisions in convert_checkpoint_from_hf.py and the load_hf_model method of olmo_core.nn.hf.checkpoint.
  • foreach support in SkipStepAdamW.
  • Added budget mode for activation checkpointing configuration.
  • Added io.remove_file() and io.glob_directory functions.
  • Added ABF, PI, and YaRN rope scaling strategies.
  • Added a script to compare two WandB runs
  • Added namespace option to nn.buffer_cache.BufferCache.
  • Added the option to configure head_stride for context parallelism with ring-flash-attn.
  • Added the option to group multiple npy source files together for packing with the packed FSL dataset by setting source_group_size to an integer greater than 1.
  • Added load_optim_state: Optional[bool] option to Trainer.load_checkpoint().
  • Added GenerationModule for OLMo-core native autoregressive generation with support for kv caching.

Changed ⚠️

  • Output of LMHead when labels is passed as input is now a 4-tuple instead of a 3-tuple, with (logits, loss, ce_loss, z_loss), where loss is the combined loss (ce_loss + z_loss).
  • The ConfigSaver callback will automatically set the config to save for other callbacks (WandBCallback, CometCallback, and BeakerCallback as of now).
  • Fixed bug causing slow evals in BPB/RC in-loop evals due to fast MC
  • Changed default precision of converted HF models in src/examples/huggingface/convert_checkpoint_to_hf.py to bfloat16.
  • Changed default cluster to saturn in src/examples/llama/train_launch.py.
  • Made some beaker secrets optional for internal experiments.
  • Changed SlidingWindowAttentionConfig to improve clarity.
  • Changed the default Beaker budget

Fixed βœ…

  • Modify TokenizerConfig.from_hf() to fallback to tokenizer_config.json if config.json is not found.
  • Fixed loading checkpoints with missing keys from transformer train modules using torch 2.7.
  • Made MoE load balancing loss more robust.
  • Fixed a bug with ReorderedNormTransformerBlock when using fine-grained FSDP wrapping and activation checkpointing together.
  • Fixed an issue preventing tensor parallelism from working with LMHead when using the "fused_linear" loss implementation.
  • Fixed a bug with LMHead when using "fused_linear" loss implementation where the ce_loss output included the z_loss added to it.
  • Fixed training on single GPU when using a SkipStepOptimizer.
  • Fixed the initialization of the CosWithWarmupAndLinearDecay learning rate scheduler
  • Ensured eval tasks are sorted to maintain the same order across ranks (the cookbook was configuring these in an unsorted way).
  • W&B callback uses working directory instead of save folder for local cache.
  • Reset speed monitor callback after changing batch size.
  • Fixed parallelism compatiblity between cp + tp and cp + pp and added test to catch regressions.
  • Ensure sharded parameters are initialized differently on separate ranks.
  • Fixed fingerprinting for FSL datasets
  • Fixed bug where step state in SkipStepAdamW was not incremented, biasing the optimizer steps. Added option to restore the bug for backwards compatibility.
  • Removed sklearn from upstream dependency ai2-olmo-eval.
  • Made removing ephemeral checkpoints more robust.
  • Made running bookkeeping operations more robust.
  • Ensure RoPE modules with different settings use a unique sub-cache for their buffers.
  • Fixed bug with context parallelism where every transformer block would use the same RoPE buffers even if their RoPE was configured differently.
  • Fixed MFU computation to work with FSDP, corrected some device specs.
  • Optimization: avoid redundant calls to model.train() in TransformerTrainModule.
  • NumpyDatasetConfig.expand_glob now works with remote directories.
  • Fixed Attention block sharding when TP and head-wise QK norm are both applied.

Commits

de89fbe (chore) prepare for release v2.2.0
54d3af0 run gpu tests with gantry (#357)
c82d13c GenerationModule with support for KV Caching (#324)
effdef3 Add option to Trainer.load_checkpoint() to ignore optim state (#351)
2cd5b82 Fix TP when headwise QK norm is applied (#353)
429054a Fix empty config.json output in convert_checkpoint_from_hf (#354)
e97f58d Fix skip step optimizer with TP (#352)
7f4b45f Option to group npy sources together for packing (#349)
9fb6366 Add io.glob_directory function (#348)
67854f9 Avoid redundant calls to model.train() (#345)
7e14a68 Fix CP bug with RoPE buffers (#341)
699971d Support configuring head_stride for ring-flash-attn (#344)
b0019cc Pull updates from olmo3 branches (#334)
026b882 Fix MFU Calculation (#343)
4162151 Add 'namespaces' to BufferCache to avoid collisions (#340)
085755c Port the WandB comparison tool from the old trainer (#338)
9e42a20 The Beaker default budget has changed (#337)
e395b82 Ensure LR scheduler's unit are tokens when using BZ scheduler (#335)
08340f9 fix off-by-one issue with SWA
abc12e5 Make async bookkeeping more robust (#333)
6d2f334 Add TrainModule.pre_train(), fix pain point with BZ scheduler (#332)
8fb85bc make Scheduler a subclass of Config (#331)
e1bac95 More Rope Scaling Implementations (PI, Yarn) (#330)
97a6f4e Make removing ephemeral checkpoints more robust (#329)
ebe0b19 loosen numpy requirement (#328)
5b924a6 Add missing return in init (#327)
34beda5 Bump ai2-olmo-eval==0.8.5 (#326)
f107e9c fix typo in release process
bddf65a fix initialization (again) (#319)
992a79e Memory budget strategy for activation checkpointing (#297)
0dda3ec DDP parameter dtype casting for 16-bit precision and flash attention support (#314)
26998de Improve clarity of SWA config (#301)
91630ea Add config option for to enable/disable the step-increment bugfix. (#317)
51e2049 "Better" sorting of Augusta ranks at runtime (#313)
fe50d9e Only check for beaker secrets in non-distributed settings (#302)
c41962a SkipStepAdamW foreach implementation; bug fix for step state in SkipStepAdamW (#309)
a04e93f ensure sharded parameters initialized different on different ranks (#307)
f7c394d Fix fingerprints for various FSLDatasets (#303)
5f93eed hot fix
20d833b Add support for revisions in conversion from HF (#304)
b0f1d9e use qualname instead of name
b744d9f Make async bookkeeping more robust (#305)
796c2dc Add support for accessing Google cloud on non-Google clusters (#299)
2bdb93c Allow profiler to record on multiple ranks (#298)
00f9b00 Fix for parallelism compatibilities in build_world_mesh (#293)
ef2845f Bump pytorch, ring-flash-attn, and liger-kernel versions (#295)
4cffd08 Add support for BOS token matching EOS token in documents (#291)
c5777a3 Make ruff line-length match black line-length (#290)
f2cc497 Reset speed monitor after changing batch size (#289)
9816e44 ai2-olmo-eval==0.8.4 (#288)
a42ce36 Add head QK norm support for Attention (#287)
e7d01b1 use work_dir instead of save_folder for W&B cache
2cd5a31 Update default cluster for src/examples/llama/train_launch.py (#286)
2b73817 Support lower precisions for conversion to HF (#277)
aa96c3a Bump ai2-olmo-eval==0.8.3 (RC/BPB speed fix) (#285)
9fac3c5 HF conversion improvements (#284)
db2b8a4 catch other types of error when importing liger-kernel (#283)
fbeaa97 Add "fast" in-loop task set (#282)
776778e Fast in-loop MCQA (#281)
c779ca5 Bump ai2-olmo-eval==0.7.2 (in-loop Basic Skills) (#279)
fd44f03 Update images to stable 2.7.0, use CUDA 12.8 by default (#280)
1acde9d Fix Numpy data loader indices (again) (#276)
b18e58a Ensure eval tasks are sorted for consistent order (#275)
e185944 Save metadata to Beaker results dir (#274)
490f03a make indices filename robust to change in batch size (#273)
80cf40e Add BatchSizeSchedulerCallback (#272)
cd3d995 Sliding Window Attention (#271)
53c28a4 use REST API to follow jobs again
ef28f2a update pins
1662d0d don't auto upgrade beaker-py for now
5cc16a2 Add Interleaved Numpy Dataset (#263)
1cb5add Scheduler init (#267)
2caadea Move test utilities to new submodule olmo_core.testing (#266)
5d4c7ec Fix singe-GPU training with SkipStepOptimizer (#265)
9a19a71 More fixes for LMHead with TP (#264)
5bd9006 minor improvements to log streaming
1d4bd1d Add a numpy FSL dataset variant ...

Read more

v2.1.0

Choose a tag to compare

@github-actions github-actions released this 14 Apr 18:08

What's new

Added πŸŽ‰

  • Added 50B Dolmino 11/24 mix.
  • Added support for auxiliary-loss-free MoE load-balancing, similar to DeepSeek-v3. You can activate this by setting bias_gamma to a non-zero float in your MoERouter config.
  • Added support for sequence-level MoE load balancing loss.
  • Compatibility with B200s.
  • Added support for warmup_fraction as an alternative to warmup_steps in all schedulers, allowing warmup to be specified as a fraction of total training steps.
  • A better config for the 1B model, ported from the old OLMo trainer.
  • Added auto_resume option to CometCallback for resume an existing run.
  • (BETA) Added methods load_hf_model and save_hf_model for saving supported OLMo Core models to HF transformers format.
    Also added lower-level methods for converting state between the formats.
  • Added the ability to run the evaluator callback on .pre_train() by setting eval_on_startup=True, and to cancel the run after the first time evals run by setting cancel_after_first_eval=True.
  • Added support for label mask files with numpy FSL datasets.
  • Added a git configuration to BeakerLaunchConfig.

Changed ⚠️

  • TransformerTrainModuleConfig can now be used to build a TransformerPipelineTrainModule by adding a pp_config spec. This makes the TransformerPipelineTrainModuleConfig redundant, but it will be kept around for backwards compatibility until the next major release.
  • Several state dict methods in TrainModule now take an optim option, which can disable the use of optimizer state.
  • Updated Float8Config for latest version of torchao.
  • Undo a fix applied to olmo_core.data.numpy_dataset.NumpyFSLDatasetMixture that was generating a mismatch between the shape of instances in the dataset and the shape of instances in the data loader.
  • Made the 1B and 7B scripts more similar to each other.
  • Changed underlying logic and top-level arguments of convert_checkpoint_from_hf.py and convert_checkpoint_to_hf.py.
  • Beaker experiments launched with the BeakerLaunchConfig will now log with ANSI colors enabled.

Fixed βœ…

  • Fixed calculation of total steps based on epochs at the end of a training job.
  • Fixed a bug where the trainer might try to save a duplicate final checkpoint if the run that already completed was restarted.
  • When submitting a Beaker job from a branch that's tracking a GitHub fork, OLMo-core now instructs Beaker to pull from the fork instead of from the main repo.
  • Made Beaker image resolution more robust.
  • Having t_max overrides in the default model configs is confusing and error prone, so we removed them.
  • Beaker launcher will only clone a single branch at runtime when possible, which can be much faster.

Commits

b8070fb (chore) prepare for release v2.1.0
7bc8aa2 remove erroneous license in test file
db91b7f Add a git config to BeakerLaunchConfig (#251)
36b791a [HF Converter] Expect model and optim state in model_and_optim subdirectory (#253)
1f2f6f9 Log with ANSI colors in Beaker (#252)
d0ab790 No more t_max (#247)
5653c92 rename * (unscaled) metrics to * unscaled
60a19c3 clone single branch when possible (#250)
e9a34e8 More MoE updates (#246)
c149b73 Update images for torch 2.7.0 (#249)
6d2bb0a Added 50B Dolmino-1124 mix (#248)
53e67ce Add option to cancel run after first evals (#244)
b493d50 fix in-loop normalization with v2 (#243)
a07ef78 Add a self-contained template train script (#242)
746408e Port the 1B from old OLMo (#234)
4ec0866 Add support for label masks with numpy datasets (#241)
ecb14e0 only resume if name matches (#240)
fc84edc Add option to auto resume Comet experiments (#239)
f5d85a9 OLMo Core to HF conversion refactor (#226)
23c6cb1 clean up logging output from source mixture tests
d502b7e Mapping new ladder to old ladder (#146)
a135883 fix calculation of max steps based on epoch at the end (#236)
2f66fd9 Added warmup_fraction to all schedulers (#235)
be06aa0 B200 compatibility (#232)
0973d4d make beaker image resolution more robust (#233)
78be552 Pick the correct remote (#230)
590138d Temp disables custom read_chunk_from_array in SourceMixture (#231)
082e0b1 Fix bug when restarting a completed run (#229)
6c626f2 Update float8 API for latest torchao (#228)
8919dff Some MoE changes/additions to support auxiliary-loss-free load-balancing (#227)
26e9476 Allow train modules to not load/save optimizer state (#225)
8c20a64 run cuda gc at the end of training
a907892 Merge transformer train module configs (#224)
b47e01c Added 32B stage2 checkpoints .csv (#220)

v2.0.1

Choose a tag to compare

@github-actions github-actions released this 18 Mar 20:56

What's new

Added πŸŽ‰

  • Added information about the official 32B training run.
  • Added automatic support for LL128 when running on Augusta.

Fixed βœ…

  • The official config for the 32B had unrealistic batch size settings.
  • Ignore group_overrides for frozen parameters instead of throwing an error.

Removed πŸ‘‹

Commits

27b1ae8 (chore) prepare for release v2.0.1
79ebc7f Add hybrid MoE transformer architecture (#223)
bce2b5b authenticate with Docker Hub to avoid rate limits
b1e0bbd Remove fused CE loss, reorganize MoE kernels/ops (#221)
56e06ee Ignore group_overrides for frozen params (#219)
9d80e8d Update logo for README header. (#218)
974e555 fix some typos, consistent naming
45fe007 Updated documentation (#217)
51aedcf More working config (#216)
47b2ad5 add release PR comments back in

v2.0.0

Choose a tag to compare

@github-actions github-actions released this 13 Mar 01:42

What's new

This major release introduces a few breaking changes. We've provided more information here: OLMo-core v2 design and upgrade guide.

Added πŸŽ‰

  • Added TrainModule abstraction with TransformerTrainModule implementation, which encapsulates both a model and optimizer.
  • Added namespace argument to Trainer.record_metric().
  • Added support for context parallelism.
  • Added support for expert parallelism with MoE models.
  • Added in-loop evals for Minerva, GSM, HumanEval, MBPP (ai2-olmo-eval==0.7.0)
  • Added CosWithWarmupAndLinearDecay learning rate scheduler
  • Added WSD learning rate scheduler

Changed ⚠️

  • The Trainer now takes a TrainModule instead of a model and optimizer, and several configuration options have been moved to TransformerTrainModule, including rank_microbatch_size, fused_loss, compile_loss, z_loss_multiplier, and autocast_precision.
  • Several TransformerModelConfig options have been to TransformerTrainModule / TransformerTrainModuleConfig, including dp_config, tp_config, float8_config, and compile.

Removed πŸ‘‹

  • Removed the following callbacks: MoEHandlerCallback, SchedulerCallback, MatrixNormalizerCallback, GradClipperCallback, and Float8HandlerCallback.
    The functionality from all of those callbacks has been moved to the TransformerTrainModule class.
  • Removed the callback methods .pre_eval_batch() and .post_eval_batch().

Fixed βœ…

  • Fixed the model ladder code when training on mps or cpu device

Commits

dfa8f2b (chore) prepare for release v2.0.0
95fb084 add work-around for pytorch/ao#1871 (#205)
3ce0c58 32B Documentation (#210)
41f8ddc Add a public "official" version of our 32B train script (#214)
7e58d12 Update data paths in example to public URLs (#213)
4327bb9 upload data to r2 and updated their paths (#208)
0e6ea23 Assorted improvements (#207)
9ceb1e4 Add CUDA 12.6 images (#209)
eda3afb guard against wrapping MoE modules for AC (#206)
6e5b16f Bump ai2-olmo-eval==0.7.0 (in-loop Minerva, GSM, HumanEval, MBPP) (#204)
eccdc00 Make it easier for external users to run train scripts (#203)
da33f5b fix entrypoint steps
947a293 clean up changelog
725adf3 V2 (#202)

v1.9.0

Choose a tag to compare

@github-actions github-actions released this 10 Mar 20:41

What's new

Fixed βœ…

  • Ensure certain optimizer param group fields are not overridden by the values in a checkpoint.

Added πŸŽ‰

  • Added instance_filter_config field to NumpyDatasetConfig.
  • Added conversion script for OLMo 2 checkpoints to Huggingface format.
  • Added BeakerCallback.
  • Added logging for in-loop eval throughput

Fixed βœ…

  • Ensure certain optimizer param group fields are not overridden by the values in a checkpoint.
  • Fixed issue where non-zero ranks would report partially-reduced values for training metrics.

Commits

41a7dbd (chore) prepare for release v1.9.0
d7301e6 32B scripts (#201)
d55562c Log in-loop eval throughput (#200)
260dafd Add support for BF16 optim state in SkipStepAdamW (#148)
e522437 fix inferring sequence length
0bef5aa allow dynamic batch sizes (#170)
fa11a40 Port over instance filtering from old codebase (#157)
8ef038a update formatting of bucket distribution
c9ca78a Add a BeakerCallback (#177)
e1cd8f6 use effective sequence length
32cb0fa Conversion script for OLMo 2 models trained with OLMo core to HuggingFace (#158)
feb57eb all-reduce train metrics (#166)
2b43d59 reset initial LR to configured value after loading (#163)
2902a9c Improve Config.from_dict (#156)
b4cee6d ignore class name field when config from dict
c1d1a53 update DTensor imports to use public module (#153)
4594231 activate virtual env before running script