v0.8.18 — the CPU KV pool is sized from the checkpoint and from free RAM
Patch release. One defect, reachable since 0.8.17.
Defect
plan_cpu_kvcache_gib sized the CPU KV pool from ram_bytes alone — a quarter of it,
clamped to [2, 8] — and never saw the checkpoint. On a 30.8 GiB box serving the 13.9 GiB
gemma-4-26B-A4B it asked for 7 GiB of KV, giving 27.7 GiB of anonymous memory against
30.8 GiB installed.
Observed: kswapd0 at 100%, 85% iowait, buff/cache 100 MiB, ssh unreachable. No OOM kill
— with no swap (the cloud default) the only reclaimable pages were the mmap'd weights being
read back in.
0.8.17 made it reachable: before it the installer recommended only dense trellis
checkpoints on CPU, which are small.
The GPU path was unaffected. plan_gpu_memory_utilization sizes from
weights + overhead + minimum KV; on a GPU the weights and KV are in VRAM while the loader
and page cache are in RAM. On CPU all four share one pool.
Fix
pool = min(_CPU_ANON_FRACTION * MemTotal, MemAvailable - _CPU_HEADROOM_BYTES)
- weights - runtime
MemAvailable, not MemFree or free(1)'s used column, which exclude reclaimable page
cache. Read once at supervisor start.
--gpu-memory-utilizationno longer suppresses the checkpoint-size lookup on CPU, where
the flag has no effect.glq-chatandglq-codeshare one helper.- Unknown checkpoint size caps the pool at 4 GiB instead of taking a quarter of RAM.
- The supervisor prints the plan and, when it does not fit, the shortfall in GiB:
RAM plan: 13.9 GiB weights + 4 GiB KV pool + ~7 GiB runtime = 24.9 GiB, against 29.8 GiB free
Constants
Per-process PSS from smaps_rollup, summed against /proc/meminfo AnonPages (agreeing to
0.1 GiB), steady state, glq-chat serving the 26B-A4B on a 30.8 GiB box:
| GiB | |
|---|---|
VLLM::Worker — 13.9 weights + 5.0 pool + 3.7 activations |
22.56 |
| chat/gradio UI | 2.09 |
| EngineCore + API server | 1.06 |
| AnonPages | 25.64 of 30.81 = 83% |
| MemAvailable | 4.42 |
Cost beyond weights and pool: 6.8 GiB. _CPU_RUNTIME_OVERHEAD_BYTES 7 GiB,
_CPU_ANON_FRACTION 0.85, _CPU_HEADROOM_BYTES 4 GiB.
Separation of the two observed configurations: 7 GiB pool = 90% anonymous (thrashed),
5 GiB pool = 83% (served, 4.42 GiB available). The KV pool is not in shared memory
(Shmem 0.00, /dev/shm 1 MiB).
Validation
- 159 tests: both measured configurations pinned, busy machine plans smaller than idle,
no-room machine clamps to the floor, MemAvailable fallback,available_ram_bytesparsing. - Installed from source on a 30.8 GiB CPU box, served end to end: ready in 182 s, clean
shutdown, nothing left resident. - Busy-machine case with RAM actually occupied: idle 4 GiB pool, 10 GiB occupied 2 GiB
(floor) with the shortfall reported, released 4 GiB.
Distro matrix not run: this release touches none of install.sh, pre-flight or the kernel
build. Existing checkpoints unaffected — no CUDA path, checkpoint format or quantize path
changed.