What's new
Added 🎉
- Added
max_checkpointsparameter toCheckpointerCallback(default: 3) to limit the number of permanent checkpoints retained. Oldest checkpoints are removed automatically when the limit is exceeded. Set toNoneto keep all (previous behavior). - Added
OutputDiscardCheckpoint, an activation-recompute primitive for cases where the output of a checkpointed region dominates memory rather than its intermediates (e.g. precision casts, FFN up-projections). Forward runs underno_grad, the output's storage can be freed after downstream consumption, and a backward hook recomputes and rebinds the freed storage in place via a C++share_storageextension (with a Python fallback for environments without a C++ toolchain). - Added Qwen3.5 dense model configs (0.8B, 4B, 9B, 27B) with hybrid Gated DeltaNet + full-attention architecture.
- Added partial RoPE support via
partial_rotary_factoron :class:~olmo_core.nn.rope.RoPEConfig. - Added HuggingFace weight conversion for
qwen3_5_texthybrid models. - Added a configurable vision transformer encoder (
VisionTransformer, configured viaVisionEncoderConfig), vision-to-LM connector (VisionConnector), andMultimodalLM— a composite vision-language model that fuses image patch tokens into the LM token stream. Supports OpenAI CLIP, SigLIP, and SigLIP2 encoder variants with factory configs for all standard Molmo2 checkpoints. - Added
HFConverterCallback, which can be used to convert models to huggingface format at the end of the training run. - Trainer now records checkpoint save and load durations as
train/checkpoint_save_duration_sandtrain/checkpoint_load_duration_smetrics. - Added
PowerLR, a power-law learning rate scheduler with linear warmup, power-decay phase (lr = initial_lr * (current / warmup) ** bfor negativeb, making the LR independent of the training horizon), and an optional linear decay tail. Registered as"power_lr". - Added
ComposableScheduler, a piecewise LR scheduler built fromComposableSchedulerStagesegments (linear/cosine interpolation between endpoint LRs) on an absolute time axis. Registered as"composable". Note:ComposableSchedulerignores thet_maxpassed toget_lrand emits a once-per-instanceUserWarningto that effect. - Added
OverrideDecay, a late-stage decay override usable on bothComposableSchedulerandSequentialSchedulervia anoverride_decayfield. Whencurrent >= override_decay.start, the main schedule is interrupted mid-flight and the LR decays from the value the main schedule would have produced atstartto a target LR overduration(linear or cosine).SequentialScheduleradditionally warns thatt_maxis ignored once the override becomes active. OLMO_RICH_LOGGINGcan now explicitly enable or disable rich console logging (0/false/no/offdisables it); previously setting it to any value only force-enabled rich logging.init_distributed()now bootstraps a minimal single-process environment (RANK=0,WORLD_SIZE=1,MASTER_ADDR/MASTER_PORT) when launch env vars are absent, so scripts can be run directly (withouttorchrun) for single-process debugging.- Added a configurable
determinism_checkoption to activation checkpointing (default"default"); set it to"none"to skip torch's recompute metadata check for opaque linear-attention kernels undertorch.compile.
Fixed ✅
- The CPU
TestCI job now cachesHF_HOMEacross runs so the HuggingFace roundtrip tests (Qwen3-0.6B, Gemma-3-270m) don't re-download their checkpoints every run. - Excluded
mark_dynamicfromtorch.compiletracing (@torch.compiler.disable). - Clearer error messages (now include the offending values) when a rank batch size isn't divisible by the sequence length, or
max_target_sequence_lengthisn't a multiple ofsequence_length. - S3 uploads/downloads now also retry on transient SSL errors (
ssl.SSLError, botocore/urllib3SSLError). - Distributed checkpoint writes now clone each tensor before serialization to avoid accidentally writing the full backing storage of a view/shared tensor, with a guard that raises
OLMoCheckpointErrorif a written tensor is unexpectedly larger than itsnbytes. - Fixed LM in-loop evaluator data-order drift across repeated runs by resetting loader bookkeeping before each pass and making deterministic reshuffling the default.
- Fixed Qwen3 implementation to match HuggingFace by applying RoPE in the input dtype (bf16) rather than upcasting to fp32.
- Fixed HF model conversion for Llama, Qwen3, and Gemma so that converted checkpoints roundtrip correctly.
- Fixed Beaker secret existence check to use the case-insensitive HTTP endpoint, avoiding spurious "secret not found" errors when secret names differ only in case.
- Fixed
Transformer.init_weightsso that under interleaved pipeline parallelism (e.g.Interleaved1F1B,InterleavedZeroBubble) the multiple model chunks owned by a single rank no longer initialize to identical parameters. Adds amodel_part_idxkwarg incorporated into the seed asmodel_part_idx * pp_size. - Disabled
torch.compiletracing throughTEAttentionBackend.forward, whose Python/pybind setup is not Dynamo-safe. - Fixed
TransformerPipelineTrainModule.num_flops_per_tokenreturningNoneunder pipeline parallelism. Each PP rank only holds its stage's layers, so summing FLOPs frommodel_partsundercounts the model. Capturemodel.num_flops_per_tokenas a bound method beforesplit_modeldeepcopies and drops layers, then call it at metric time. On meta device (the standard PP init path) this has no memory cost.
Changed ⚠️
- Set transformers version to >= 5.4.0 for Qwen 3.5 and in sync with open-instruct
- Added a documented
deterministicoption toLMEvaluatorandLMEvaluatorCallbackConfigso callers can opt out of fixed eval ordering when desired.
Commits
b7e9671 (chore) prepare for release v2.6.0
77714b7 fix dion pypi (#816)
5b0d1cf (chore) prepare for release v2.6.0
66f768b Update anonymous paths in public scripts to point to data in hugging face bucket (#802)
064b172 bump fla to 0.5.2 (#798)
d3146cc Misc fixes (#800)
fa6c501 Add max_checkpoints to limit permanent checkpoint retention (#694)
c3802ed CI: cache HuggingFace models for the CPU Test job (#732)
2000b1b Add OutputDiscardCheckpoint (#682)
9aa3280 Add Qwen3.5 model support (0.8B, 4B, 9B, 27B) (#684)
36f99f0 Add conversion overrides for Llama, Qwen3, and Gemma 4 models so they roundtrip properly (#677)
885219b Add configurable determinism_check to activation checkpointing (#713)
8d22ca9 [1/n] Add vision transformer, connector, and MultimodalLM (#692)
59a339f Misc training-utility improvements (#691)
1713ea3 Improve checkpoint/S3 IO robustness (#690)
754d58d CI: run GPU tests at low priority (#696)
1af17a4 Fix PP FLOPs: capture full model before pipeline split (#680)
525cc25 Increase tolerance for flaky test (#681)
38704d1 Two small transformer core fixes: TE Dynamo + PP init seed (#679)
73637f7 Add PowerLR scheduler (#674)
2caaee9 Add ComposableScheduler (#671)
2e67bcb Update task timeout for GPU tests (#675)
e556a86 Revert "Add position_ids-based varlen RoPE support for packed inputs" (#672)
1eec696 Add position_ids-based varlen RoPE support for packed inputs (#654)
60d2487 Fixes the Qwen3 implementation to match HF (#663)
5e7ee43 Record checkpoint save/load durations as trainer metrics (#665)
3622318 Fix broken path in 32B LC (#670)
3e19fa2 Fix breaking paths (#669)
53c51c5 Use case-insensitive HTTP endpoint to check beaker secret existence (#666)
afe99b6 Pin flash-linear-attention version to 0.4.1 (#664)
2e57086 Add HFConverterCallback for end-of-training HuggingFace conversion (#660)
beca1f1 Correct dolmino mix urls (#630)
b376077 Fall back to anonymous GCS client when no credentials are available (#659)
60930ef Add gradient dumping to GAPMonitorCallback (#438)
befb60b make lm evaluator data deterministic (#652)