Skip to content

v0.2.7 — Speed boost + Multi-GPU fix + --host fix

Choose a tag to compare

@github-actions github-actions released this 28 Apr 04:51
· 9 commits to main since this release

Speed optimizations (automatic, no user action needed)

MTP speculative decoding (+40-80% for Qwen3.6)

Models with native MTP heads (Qwen3.6 series) now get --num-speculative-tokens 3 automatically. The model predicts 3 tokens per forward pass instead of 1. The NativeMTP field was already in the profile but never passed to llama-server — now it is.

N-gram lookup (+20-50% for code/structured output)

Models without MTP get --lookup 8 — zero-cost speculative decoding that finds repeated patterns in context history. Especially effective for code generation and structured output.

KV cache defragmentation

--defrag-thold 0.1 auto-compacts KV cache when fragmentation exceeds 10%. Prevents effective context from shrinking during long conversations.

Bug fixes

Multi-GPU: strip CUDA_VISIBLE_DEVICES

Some environments restrict llama-server to a single GPU. Now removed from child process env so all detected GPUs are used.

--host flag applies to proxy

--host 0.0.0.0 now affects both llama-server AND the proxy (port 11435). Previously only llama-server listened on 0.0.0.0, proxy was hardcoded to 127.0.0.1.

Blackwell JIT warmup

First run on RTX 50-series triggers llama-server --version to populate CUDA JIT cache. All subsequent launches start in ~2s.

Multi-GPU: skip --kv-unified

Prevents KV cache from being allocated entirely on GPU 0.

Full changelog

https://github.com/val1813/kaiwu/blob/main/README.md#changelog