v0.2.7 — Speed boost + Multi-GPU fix + --host fix
Speed optimizations (automatic, no user action needed)
MTP speculative decoding (+40-80% for Qwen3.6)
Models with native MTP heads (Qwen3.6 series) now get --num-speculative-tokens 3 automatically. The model predicts 3 tokens per forward pass instead of 1. The NativeMTP field was already in the profile but never passed to llama-server — now it is.
N-gram lookup (+20-50% for code/structured output)
Models without MTP get --lookup 8 — zero-cost speculative decoding that finds repeated patterns in context history. Especially effective for code generation and structured output.
KV cache defragmentation
--defrag-thold 0.1 auto-compacts KV cache when fragmentation exceeds 10%. Prevents effective context from shrinking during long conversations.
Bug fixes
Multi-GPU: strip CUDA_VISIBLE_DEVICES
Some environments restrict llama-server to a single GPU. Now removed from child process env so all detected GPUs are used.
--host flag applies to proxy
--host 0.0.0.0 now affects both llama-server AND the proxy (port 11435). Previously only llama-server listened on 0.0.0.0, proxy was hardcoded to 127.0.0.1.
Blackwell JIT warmup
First run on RTX 50-series triggers llama-server --version to populate CUDA JIT cache. All subsequent launches start in ~2s.
Multi-GPU: skip --kv-unified
Prevents KV cache from being allocated entirely on GPU 0.
Full changelog
https://github.com/val1813/kaiwu/blob/main/README.md#changelog