A ~500M-parameter decoder-only language model, pretrained entirely from scratch in Python/PyTorch. The transformer, the byte-level BPE tokenizer, the data pipeline, the Muon+AdamW optimizer, and the training loop are all implemented in this repo, with no transformers and no pre-built model components. Development and correctness verification happen locally on Apple Silicon (CPU/MPS); the real pretraining run targets a rented cloud GPU.
tinylm/model.py implements every layer from scratch: RMSNorm pre-norm, RoPE, grouped-query attention with QK-norm, SwiGLU, tied embeddings, no biases anywhere. Two presets ship in tinylm/config.py:
| Hyperparameter | d26 (target) |
smoke (local dev) |
|---|---|---|
| Parameters | 489,297,408 total (≈447M non-embedding) | 13,111,296 |
| Layers | 26 | 6 |
| d_model | 1280 | 256 |
| Attention heads | 20 (head_dim 64) | 4 |
| KV heads (GQA) | 4 | 2 |
| FFN | SwiGLU, hidden 3456 | SwiGLU, hidden 768 |
| Norm | RMSNorm, pre-norm + QK-norm | same |
| Positional encoding | RoPE, θ = 10,000 | same |
| Context length | 2048 | 512 |
| Vocab | 32,768 (byte-level BPE, trained by us) | same |
| Embeddings | Tied input/output | same |
| Biases | None | None |
Optimization uses a Muon + AdamW hybrid: Muon (momentum + Newton-Schulz orthogonalization) for all 2-D hidden weight matrices, AdamW for embeddings and RMSNorm gains, under a warmup-stable-decay (WSD) LR schedule. Training runs in bf16 autocast, with torch.compile on CUDA and plain DDP via torchrun for multi-GPU; single-GPU is the degenerate case of the same code path.
uv sync
uv run pytest -q # 100 tests, a few secondsThe test suite covers every component (tokenizer, model, Muon, data sharding, training loop, eval) against hand-derived reference implementations, and needs no downloaded data.
To exercise the full training loop end-to-end on a laptop, use the smoke config (13.1M params, 6 layers) against a small locally-prepared shard set; this trains on Apple Silicon MPS (or CPU) in roughly 15-30 minutes:
uv run python scripts/prepare_data.py --data-dir data/smoke \
--max-tokens 60000000 --val-tokens 2000000 --shard-tokens 25000000
uv run python -m tinylm.train --config smoke --tokenizer tokenizer/tokenizer.json(prepare_data.py needs a trained tokenizer first; see below.)
scripts/train_tokenizer.py streams HuggingFaceFW/fineweb-edu (sample-10BT) text via datasets, trains our from-scratch byte-level BPE (tinylm.tokenizer.BPETokenizer) to a 32,768 vocab, saves it, and reports bytes/token compression on 100 held-out documents:
uv run python scripts/train_tokenizer.py \
--out tokenizer/tokenizer.json \
--max-bytes 250_000_000 \
--vocab-size 32768 \
--num-proc 8 \
--min-word-freq 2Every flag has a sane default (see --help); this is the one that produces the canonical tokenizer checked into cloud runs.
scripts/prepare_data.py streams the same FineWeb-Edu subset, encodes it with the trained tokenizer, and writes uint16 token shards (tinylm.data.ShardWriter) plus a JSON index: validation shard(s) filled first, then train shards, with tok.eot_id appended after every document:
uv run python scripts/prepare_data.py \
--tokenizer tokenizer/tokenizer.json \
--data-dir data/fineweb-edu \
--val-tokens 10_000_000 \
--shard-tokens 100_000_000 \
--max-tokens 0--max-tokens 0 (the default) consumes the entire sample-10BT subset (~10B tokens, roughly Chinchilla-optimal for the d26 model size); pass a smaller value for a quick local subset. By default the script takes a fast path: it exports the trained BPE merges to a HuggingFace tokenizers object (export_fast()), verifies it is token-identical to our own encoder on the first 200 documents (verify_fast), and then bulk-encodes in batches of 512 docs; pass --slow to force the pure-Python encoder instead (much slower, used only as a fallback/cross-check).
Both train_tokenizer.py and prepare_data.py also accept --dataset, --dataset-config, --split, and --text-key, so any HuggingFace text dataset can stand in for FineWeb-Edu; the docs site has a full runbook (Runbooks > Train on a custom dataset) including a laptop-scale TinyStories example.
Training a model for the first time? The docs site has a from-zero walkthrough (Runbooks > Your first training run) that explains every step below in detail, including installing the tools. The condensed version:
uv syncuv run pytest -q: confirm the suite is green.- Train a tokenizer on a small byte budget (a few minutes):
uv run python scripts/train_tokenizer.py --max-bytes 20_000_000. - Prepare a small shard set:
uv run python scripts/prepare_data.py --data-dir data/smoke --max-tokens 60000000 --val-tokens 2000000 --shard-tokens 25000000. - Train:
uv run python -m tinylm.train --config smoke --tokenizer tokenizer/tokenizer.json; logs toout/smoke/log.csv, checkpoints toout/smoke/ckpt_last.pt, periodic sample generations printed to stdout. - Sample from the checkpoint:
uv run python -m tinylm.sample --ckpt out/smoke/ckpt_last.pt --tokenizer tokenizer/tokenizer.json --prompt "Once upon a time".
This whole loop runs on CPU or Apple Silicon MPS and needs no GPU: torch.compile is automatically disabled off-CUDA (no --no-compile flag needed), and precision falls back to fp32 when --dtype is left at its auto default.
The smoke run above is deliberately undertrained (~0.75 tokens per parameter). The same config trained to its full ~20 tokens-per-parameter budget (16,000 steps, ~262M tokens) is an overnight run on the same laptop, roughly 5.5-6.5 hours at the recorded MPS throughput:
uv run python scripts/prepare_data.py --data-dir data/fineweb-275m --max-tokens 275000000
caffeinate -is uv run python -m tinylm.train --config smoke \
--tokenizer tokenizer/tokenizer.json \
--data-dir data/fineweb-275m --out-dir out/smoke-full --steps 16000The docs site runbook (Runbooks > Fully local training) has the full arithmetic, what to expect from the curve, and how to size a ~50M-parameter custom preset that still trains locally.
The real pretraining run, d26, ~10B FineWeb-Edu tokens (approximately Chinchilla-optimal for 489M params), is sized for a rented GPU box:
| Setup | Wall time | Approx. cost |
|---|---|---|
| 1× H100 | ≈20 h | ≈$50-60 (at $2.5-3/h) |
| 8× H100 (DDP) | ≈2.5 h | similar total spend, higher hourly rate |
Steps on a fresh Ubuntu GPU box (Lambda Labs, RunPod, etc.):
- Copy the repo (and, if you already have one,
tokenizer/tokenizer.json) onto the box. - Run the bootstrap script: idempotent, installs
uvif missing and runsuv sync, then prints the exact next commands (it skips the tokenizer step iftokenizer/tokenizer.jsonis already present):bash scripts/cloud_setup.sh
- If no tokenizer was shipped with the repo, train one (~30-60 min CPU):
uv run python scripts/train_tokenizer.py
- Prepare the full 10B-token dataset (~1-3 h CPU, ~20 GB disk):
uv run python scripts/prepare_data.py
- Launch training inside
tmuxso it survives a dropped SSH connection:Detach withtmux new -s train uv run python -m tinylm.train --config d26 --tokenizer tokenizer/tokenizer.json # or, on an 8-GPU box: uv run torchrun --standalone --nproc_per_node=8 -m tinylm.train --config d26 --tokenizer tokenizer/tokenizer.jsonCtrl-b d; reattach any time withtmux attach -t train. - If the box dies or the session is interrupted, resume from the last atomic checkpoint (bit-exact):
uv run python -m tinylm.train --config d26 --tokenizer tokenizer/tokenizer.json --resume
- Monitor progress via
out/d26/log.csv(or Weights & Biases, if--wandb <project>was passed): step, train/val loss, LR scale, tokens/sec, and MFU (populated on CUDA against a recognized GPU). - Once trained, evaluate and sample from the final checkpoint:
uv run python -m tinylm.eval_hellaswag --ckpt out/d26/ckpt_last.pt --tokenizer tokenizer/tokenizer.json --limit 1000 uv run python -m tinylm.sample --ckpt out/d26/ckpt_last.pt --tokenizer tokenizer/tokenizer.json --prompt "Once upon a time"