v0.2.2 — Blackwell OOM fix + MoE VRAM optimization + --host
Fixes
Blackwell (RTX 50-series) OOM — root cause fixed
--kv-unified pre-allocates KV cache as one contiguous VRAM block. On SM120 with CUDA 12.4 binary + CUDA 13.x driver, this over-allocates massively — even ctx=8K OOMs on 24GB cards with 15GB models. Now skipped on Blackwell; llama.cpp uses paged KV allocation instead. Warmup start also changed from ideal×2 to ideal on Blackwell.
MoE VRAM reserve — dynamic instead of hardcoded
Old: hardcoded 1536MB reserve for MoE attention weights. A 20GB MoE model has ~6GB attention on GPU — 1536MB was wildly wrong, causing KV cache selection to overestimate free VRAM.
New: model_size × 0.30 (floor 1024MB). After warmup, measured VRAM is written back to profile for even more accurate subsequent calculations. Fixes users seeing 4GB+ unused VRAM while stuck on small ctx.
iso3 detection — no more marker file dependency
EnsureBinary now returns isTurboQuant directly. Bundled binary = turboquant, downloaded = not. No external .kaiwu file needed.
New feature
--host flag
kaiwu run model --host 0.0.0.0 # listen on all interfaces (LAN)
kaiwu run model # default 127.0.0.1 (local only)Full changelog
https://github.com/val1813/kaiwu/blob/main/README.md#changelog