Skip to content

v2.6.0

Latest

Choose a tag to compare

@github-actions github-actions released this 11 Aug 15:39
· 9 commits to main since this release

What's new

Added 🎉

  • Added max_checkpoints parameter to CheckpointerCallback (default: 3) to limit the number of permanent checkpoints retained. Oldest checkpoints are removed automatically when the limit is exceeded. Set to None to keep all (previous behavior).
  • Added OutputDiscardCheckpoint, an activation-recompute primitive for cases where the output of a checkpointed region dominates memory rather than its intermediates (e.g. precision casts, FFN up-projections). Forward runs under no_grad, the output's storage can be freed after downstream consumption, and a backward hook recomputes and rebinds the freed storage in place via a C++ share_storage extension (with a Python fallback for environments without a C++ toolchain).
  • Added Qwen3.5 dense model configs (0.8B, 4B, 9B, 27B) with hybrid Gated DeltaNet + full-attention architecture.
  • Added partial RoPE support via partial_rotary_factor on :class:~olmo_core.nn.rope.RoPEConfig.
  • Added HuggingFace weight conversion for qwen3_5_text hybrid models.
  • Added a configurable vision transformer encoder (VisionTransformer, configured via VisionEncoderConfig), vision-to-LM connector (VisionConnector), and MultimodalLM — a composite vision-language model that fuses image patch tokens into the LM token stream. Supports OpenAI CLIP, SigLIP, and SigLIP2 encoder variants with factory configs for all standard Molmo2 checkpoints.
  • Added HFConverterCallback, which can be used to convert models to huggingface format at the end of the training run.
  • Trainer now records checkpoint save and load durations as train/checkpoint_save_duration_s and train/checkpoint_load_duration_s metrics.
  • Added PowerLR, a power-law learning rate scheduler with linear warmup, power-decay phase (lr = initial_lr * (current / warmup) ** b for negative b, making the LR independent of the training horizon), and an optional linear decay tail. Registered as "power_lr".
  • Added ComposableScheduler, a piecewise LR scheduler built from ComposableSchedulerStage segments (linear/cosine interpolation between endpoint LRs) on an absolute time axis. Registered as "composable". Note: ComposableScheduler ignores the t_max passed to get_lr and emits a once-per-instance UserWarning to that effect.
  • Added OverrideDecay, a late-stage decay override usable on both ComposableScheduler and SequentialScheduler via an override_decay field. When current >= override_decay.start, the main schedule is interrupted mid-flight and the LR decays from the value the main schedule would have produced at start to a target LR over duration (linear or cosine). SequentialScheduler additionally warns that t_max is ignored once the override becomes active.
  • OLMO_RICH_LOGGING can now explicitly enable or disable rich console logging (0/false/no/off disables it); previously setting it to any value only force-enabled rich logging.
  • init_distributed() now bootstraps a minimal single-process environment (RANK=0, WORLD_SIZE=1, MASTER_ADDR/MASTER_PORT) when launch env vars are absent, so scripts can be run directly (without torchrun) for single-process debugging.
  • Added a configurable determinism_check option to activation checkpointing (default "default"); set it to "none" to skip torch's recompute metadata check for opaque linear-attention kernels under torch.compile.

Fixed ✅

  • The CPU Test CI job now caches HF_HOME across runs so the HuggingFace roundtrip tests (Qwen3-0.6B, Gemma-3-270m) don't re-download their checkpoints every run.
  • Excluded mark_dynamic from torch.compile tracing (@torch.compiler.disable).
  • Clearer error messages (now include the offending values) when a rank batch size isn't divisible by the sequence length, or max_target_sequence_length isn't a multiple of sequence_length.
  • S3 uploads/downloads now also retry on transient SSL errors (ssl.SSLError, botocore/urllib3 SSLError).
  • Distributed checkpoint writes now clone each tensor before serialization to avoid accidentally writing the full backing storage of a view/shared tensor, with a guard that raises OLMoCheckpointError if a written tensor is unexpectedly larger than its nbytes.
  • Fixed LM in-loop evaluator data-order drift across repeated runs by resetting loader bookkeeping before each pass and making deterministic reshuffling the default.
  • Fixed Qwen3 implementation to match HuggingFace by applying RoPE in the input dtype (bf16) rather than upcasting to fp32.
  • Fixed HF model conversion for Llama, Qwen3, and Gemma so that converted checkpoints roundtrip correctly.
  • Fixed Beaker secret existence check to use the case-insensitive HTTP endpoint, avoiding spurious "secret not found" errors when secret names differ only in case.
  • Fixed Transformer.init_weights so that under interleaved pipeline parallelism (e.g. Interleaved1F1B, InterleavedZeroBubble) the multiple model chunks owned by a single rank no longer initialize to identical parameters. Adds a model_part_idx kwarg incorporated into the seed as model_part_idx * pp_size.
  • Disabled torch.compile tracing through TEAttentionBackend.forward, whose Python/pybind setup is not Dynamo-safe.
  • Fixed TransformerPipelineTrainModule.num_flops_per_token returning None under pipeline parallelism. Each PP rank only holds its stage's layers, so summing FLOPs from model_parts undercounts the model. Capture model.num_flops_per_token as a bound method before split_model deepcopies and drops layers, then call it at metric time. On meta device (the standard PP init path) this has no memory cost.

Changed ⚠️

  • Set transformers version to >= 5.4.0 for Qwen 3.5 and in sync with open-instruct
  • Added a documented deterministic option to LMEvaluator and LMEvaluatorCallbackConfig so callers can opt out of fixed eval ordering when desired.

Commits

b7e9671 (chore) prepare for release v2.6.0
77714b7 fix dion pypi (#816)
5b0d1cf (chore) prepare for release v2.6.0
66f768b Update anonymous paths in public scripts to point to data in hugging face bucket (#802)
064b172 bump fla to 0.5.2 (#798)
d3146cc Misc fixes (#800)
fa6c501 Add max_checkpoints to limit permanent checkpoint retention (#694)
c3802ed CI: cache HuggingFace models for the CPU Test job (#732)
2000b1b Add OutputDiscardCheckpoint (#682)
9aa3280 Add Qwen3.5 model support (0.8B, 4B, 9B, 27B) (#684)
36f99f0 Add conversion overrides for Llama, Qwen3, and Gemma 4 models so they roundtrip properly (#677)
885219b Add configurable determinism_check to activation checkpointing (#713)
8d22ca9 [1/n] Add vision transformer, connector, and MultimodalLM (#692)
59a339f Misc training-utility improvements (#691)
1713ea3 Improve checkpoint/S3 IO robustness (#690)
754d58d CI: run GPU tests at low priority (#696)
1af17a4 Fix PP FLOPs: capture full model before pipeline split (#680)
525cc25 Increase tolerance for flaky test (#681)
38704d1 Two small transformer core fixes: TE Dynamo + PP init seed (#679)
73637f7 Add PowerLR scheduler (#674)
2caaee9 Add ComposableScheduler (#671)
2e67bcb Update task timeout for GPU tests (#675)
e556a86 Revert "Add position_ids-based varlen RoPE support for packed inputs" (#672)
1eec696 Add position_ids-based varlen RoPE support for packed inputs (#654)
60d2487 Fixes the Qwen3 implementation to match HF (#663)
5e7ee43 Record checkpoint save/load durations as trainer metrics (#665)
3622318 Fix broken path in 32B LC (#670)
3e19fa2 Fix breaking paths (#669)
53c51c5 Use case-insensitive HTTP endpoint to check beaker secret existence (#666)
afe99b6 Pin flash-linear-attention version to 0.4.1 (#664)
2e57086 Add HFConverterCallback for end-of-training HuggingFace conversion (#660)
beca1f1 Correct dolmino mix urls (#630)
b376077 Fall back to anonymous GCS client when no credentials are available (#659)
60930ef Add gradient dumping to GAPMonitorCallback (#438)
befb60b make lm evaluator data deterministic (#652)