Kimi K3 (2.8T) runs on a single card in 3.72GB
Kimi K3 is the largest open-source model released to date. AirLLM runs it on one consumer-class GPU.
Measured end to end on a single RTX 6000 Ada (48GB) against the full 1.56TB checkpoint, generating real tokens:
| Peak VRAM during generation | 3.72 GB |
| Peak VRAM after init | 0.83 GB |
| Init (one-time per process) | 900 s |
| Generation | 292 s/token, disk-bound |
The reason a 2.8T model needs less VRAM than a 671B one is that sparse MoE checkpoints stream one expert at a time rather than a whole layer. K3 holds 896 experts per layer and routes each token to 16 of them — expanded, a layer's experts are ~55GB, but a token only needs ~1GB. AirLLM loads just those. MXFP4 weights also cross PCIe packed and expand on the GPU, moving 4x less data.
Fitting the checkpoint on disk needed the same kind of trick: a naive split would want 3.12TB for a 1.56TB model. K3's shards turn out to be pure, one module each, so split layers are hardlinked to the originals instead of copied.
Before you run K3
K3 brings three requirements of its own, none of them optional:
pip install airllm compressed-tensors flash-attn- flash-attn is mandatory. K3's model code overwrites whatever attention implementation you request.
- A CUDA 12 build of torch, because no prebuilt flash-attn wheel exists for CUDA 13 yet.
- transformers 4.56.x, as K3's remote code does not load on 5.x.
Also in this release
Four fixes that came out of the K3 work but are not K3-specific, and should help other models:
- Adopt the checkpoint's shape when a model class builds a different one for a parameter it never materialises.
- Expand packed weights for modules whose forward reads a plain
weightrather than the packed payload. - Read
scheme.formatacross compressed-tensors versions, where it is an enum in some and a plain string in others. - Propagate the attention implementation into nested sub-configs, which multimodal wrappers keep their decoder under.
Full details in #316.