Skip to content

v3.0.0

Latest

Choose a tag to compare

@github-actions github-actions released this 01 Oct 15:17
· 1 commit to main since this release

What's new

Added πŸŽ‰

  • Added use_array_if_local to pack_documents_into_instances, segment_documents_into_instances, NumpyPackedFSLDataset and NumpyPackedFSLDatasetConfig, forwarded to iter_document_indices. Set it to False to take document boundaries from the source metadata file instead of inferring them by scanning the token array for the EOS token. Inferring is only correct when every document is EOS-terminated: a producer that truncates documents and drops the terminator with the tail causes the affected document to merge with the one after it, and LongDocStrategy.truncate then keeps only the head of the merged span, so the following document never reaches training. Measured on an SFT cache, 97.98% of tokens reached instances via the inferred path versus 100.00% via the metadata file, with an identical maximum document length. Whenever the metadata boundaries are the effective source -- set explicitly, or because the source is a URL, for which iter_document_indices always reads the metadata -- doc_lens is derived from them rather than by rescanning the packed tokens for EOS, so the block-diagonal attention mask cannot merge an unterminated document into the one after it. The hazard is now documented on iter_document_indices. Default behavior is unchanged.
  • Added opt-in paired SwiGLU backward and BF16-rounded weight-gradient accumulation for OLMoDDP experts, including Torch 2.13 support, checkpoint/recomputation coverage, and explicit backend and bucket-ownership guards.
  • Added independent per-head Q/K norm gains and scalable softmax, EMO document-pool routing/global load balancing, and opt-in FP32 gradient-accumulation/reduce-scatter fast paths with explicit hardware/version guards.
  • Extended hybrid MoE HF export for KDA, optional EMO and latent experts, per-head normalization gains, and scalable softmax, with exact tensor round-trip validation and legacy configuration migration.
  • Added an optional, dependency-free checkpoint-ready notification callback for independent upload services. It does not upload or delete checkpoints.

Removed πŸ‘‹

  • Removed the src/scripts/train/private-olmo.py training script.

  • Removed the olmo_core.model_ladder API, its internal CLI wrapper, all eleven ladder training scripts (including the standalone Gemma-like ladder), two Slurm launchers, and the API documentation. Existing ladder orchestration configs and commands require an earlier revision; this does not change the model or optimizer checkpoint formats. The internal experiment framework remains available.

Fixed βœ…

  • Metadata-backed packed datasets now include sidecar content hashes in packing-cache keys and dataset fingerprints, invalidating stale boundaries even after same-size corrections. Document lengths preserve EOS/BOS padding segmentation. Local array-backed defaults are unchanged (#843).
  • Apply opt-in Q/K gain expansion to eval-only and model-only DDP checkpoint loads, and reject forced expert assignments and biased KDA convolutions during HF export.
  • Validate normalization throughout MoE HF exports, preserve attention-only gates and resolved EOS/padding IDs, and reject unsupported shared-expert routing before conversion.
  • Reject MoE HF exports with incompatible Q/K normalization or inconsistent KDA output-norm epsilons instead of silently changing normalization behavior.
  • Preserve the resolved tokenizer BOS ID in exported model and generation configs, including when the training config leaves BOS unspecified.
  • Reject hybrid KDA HF exports with sliding-window attention and EMO exports with restricted evaluation pools in any routed layer. Disable scalable-softmax HF generation caching and reject explicit cache use.
  • Require the CUDA 13 CuTe compiler for experimental KDA, report incompatible installs before training, and exercise the kernels in a dedicated Blackwell CI job.
  • Keep legacy fused attention configs compatible when the new attention options are disabled. Reject unsupported scalable-softmax context parallelism and KV caching at setup.
  • Assign EOS tokens to their preceding document for EMO routing, matching attention document boundaries.
  • Preserve serialized tokenizer behavior during HF export, including source BOS settings when the training config leaves BOS unspecified.
  • Preserve FP32 router probabilities and accumulation when combining BF16 experts in the HF reference model.
  • Avoid graph breaks from no-op profiling decorators, including compiled router load-balancing collectives.
  • dispatch_flash_attn_4 passed cu_seqlens_q, cu_seqlens_k, max_seqlen_q and max_seqlen_k positionally. flash-attn 4 inserted a qv parameter at position 3 (present from ~4.0.0b19 onward), which shifts every following argument by one, so max_seqlen_q β€” an int β€” lands where cu_seqlens_k is expected and the call fails with AttributeError: 'int' object has no attribute 'shape'. These are now passed by keyword; the parameter names are unchanged across flash-attn 4 releases, so this is correct against both old and new versions. Only reachable on Blackwell, where has_flash_attn_4() returns True.

Changed ⚠️

  • OLMoDDPModel.apply_ddp() now raises NotImplementedError directing callers to apply_dp(), including under python -O. Removed its unreachable legacy DDP implementation.

  • OLMoDDPTrainModuleConfig.max_grad_norm now overrides the optimizer clipping threshold when set, matching the train-module configuration API used by TransformerTrainModuleConfig. When unset, the optimizer threshold is retained. Previously this field was ignored, so old configs with differing values now use the train-module value. The supplied optimizer config is not mutated. Removed the unused OLMoDDPTrainModule.max_grad_norm constructor argument and attribute; clipping remains inside the DDP optimizer.

  • Removed the unused, commented-out Beaker execution-unit helper and its commented call from the internal experiment setup.

  • Migrated repository attention callers to explicit backend selection. Transformer factories now resolve legacy use_flash arguments into backend configs, and the nGPT factory accepts attn_backend. Deprecated inputs remain supported for old configs and callers, including explicit-backend precedence and sliding-window automatic selection.

  • Replaced FusedAttention with FusedAttentionV2 and migrated its applicable tests to V2. Legacy AttentionConfig(name="fused") configs still load through V2 with the same packed QKV parameter names and shapes, preserving model and optimizer checkpoint compatibility. Legacy fused RoPE is mapped to regular RoPE: floating-point results (especially low-precision RoPE gradients) differ, so resumed training is not numerically identical. TransformerConfig.llama_like(fused_ops=True) now selects V2 with regular RoPE when it previously selected fused attention.

  • Raised the minimum supported PyTorch version to 2.10.0. Updated the stable Beaker images to PyTorch 2.10 with CUDA 12.8 and PyTorch 2.11 with CUDA 13.0, and expanded CI coverage to include PyTorch 2.10 and 2.11 with CUDA 12.8 and Docker builds through PyTorch 2.12 with CUDA 13.0.

Commits

a800c06 (chore) prepare for release v3.0.0
3c1a108 Add third party notice (#900)
56c04b8 Add tech-report reference scripts to src/examples/olmo_ddp (#899)
6ac500d Allow document boundaries to come from the metadata file (fixes silent token loss on non-EOS-terminated documents) (#843)
718d08f Wire DDP train-module gradient clipping and remove retired apply_ddp code (#881)
5e607f8 Remove model ladder framework and retired training scripts (#879)
306d4b3 Migrate use_flash callers to explicit attention backends (#878)
f6143b5 Add hybrid MoE attention, EMO routing and HF export (#872)
845bda9 Replace FusedAttention with FusedAttentionV2 (#873)
0589a03 Pass flash-attn 4 varlen arguments by keyword (#857)
bbe9473 Caleb/kda cute kernels (#863)
f726f96 Add the OLMoDDP fused MoE / expert-parallel training stack (#844)
92870a3 Raise minimum PyTorch version to 2.10 (#835)