Skip to content

v0.2.2 — Blackwell OOM fix + MoE VRAM optimization + --host

Choose a tag to compare

@github-actions github-actions released this 27 Apr 00:58
· 17 commits to main since this release

Fixes

Blackwell (RTX 50-series) OOM — root cause fixed

--kv-unified pre-allocates KV cache as one contiguous VRAM block. On SM120 with CUDA 12.4 binary + CUDA 13.x driver, this over-allocates massively — even ctx=8K OOMs on 24GB cards with 15GB models. Now skipped on Blackwell; llama.cpp uses paged KV allocation instead. Warmup start also changed from ideal×2 to ideal on Blackwell.

MoE VRAM reserve — dynamic instead of hardcoded

Old: hardcoded 1536MB reserve for MoE attention weights. A 20GB MoE model has ~6GB attention on GPU — 1536MB was wildly wrong, causing KV cache selection to overestimate free VRAM.

New: model_size × 0.30 (floor 1024MB). After warmup, measured VRAM is written back to profile for even more accurate subsequent calculations. Fixes users seeing 4GB+ unused VRAM while stuck on small ctx.

iso3 detection — no more marker file dependency

EnsureBinary now returns isTurboQuant directly. Bundled binary = turboquant, downloaded = not. No external .kaiwu file needed.

New feature

--host flag

kaiwu run model --host 0.0.0.0   # listen on all interfaces (LAN)
kaiwu run model                    # default 127.0.0.1 (local only)

Full changelog

https://github.com/val1813/kaiwu/blob/main/README.md#changelog