Skip to content

v0.8.18 — the CPU KV pool is sized from the checkpoint and from free RAM

Choose a tag to compare

@cnygaard cnygaard released this 06 Sep 09:35
· 30 commits to main since this release
208915d

Patch release. One defect, reachable since 0.8.17.

Defect

plan_cpu_kvcache_gib sized the CPU KV pool from ram_bytes alone — a quarter of it,
clamped to [2, 8] — and never saw the checkpoint. On a 30.8 GiB box serving the 13.9 GiB
gemma-4-26B-A4B it asked for 7 GiB of KV, giving 27.7 GiB of anonymous memory against
30.8 GiB installed.

Observed: kswapd0 at 100%, 85% iowait, buff/cache 100 MiB, ssh unreachable. No OOM kill
— with no swap (the cloud default) the only reclaimable pages were the mmap'd weights being
read back in.

0.8.17 made it reachable: before it the installer recommended only dense trellis
checkpoints on CPU, which are small.

The GPU path was unaffected. plan_gpu_memory_utilization sizes from
weights + overhead + minimum KV; on a GPU the weights and KV are in VRAM while the loader
and page cache are in RAM. On CPU all four share one pool.

Fix

pool = min(_CPU_ANON_FRACTION * MemTotal, MemAvailable - _CPU_HEADROOM_BYTES)
       - weights - runtime

MemAvailable, not MemFree or free(1)'s used column, which exclude reclaimable page
cache. Read once at supervisor start.

  • --gpu-memory-utilization no longer suppresses the checkpoint-size lookup on CPU, where
    the flag has no effect. glq-chat and glq-code share one helper.
  • Unknown checkpoint size caps the pool at 4 GiB instead of taking a quarter of RAM.
  • The supervisor prints the plan and, when it does not fit, the shortfall in GiB:
    RAM plan: 13.9 GiB weights + 4 GiB KV pool + ~7 GiB runtime = 24.9 GiB, against 29.8 GiB free

Constants

Per-process PSS from smaps_rollup, summed against /proc/meminfo AnonPages (agreeing to
0.1 GiB), steady state, glq-chat serving the 26B-A4B on a 30.8 GiB box:

GiB
VLLM::Worker — 13.9 weights + 5.0 pool + 3.7 activations 22.56
chat/gradio UI 2.09
EngineCore + API server 1.06
AnonPages 25.64 of 30.81 = 83%
MemAvailable 4.42

Cost beyond weights and pool: 6.8 GiB. _CPU_RUNTIME_OVERHEAD_BYTES 7 GiB,
_CPU_ANON_FRACTION 0.85, _CPU_HEADROOM_BYTES 4 GiB.

Separation of the two observed configurations: 7 GiB pool = 90% anonymous (thrashed),
5 GiB pool = 83% (served, 4.42 GiB available). The KV pool is not in shared memory
(Shmem 0.00, /dev/shm 1 MiB).

Validation

  • 159 tests: both measured configurations pinned, busy machine plans smaller than idle,
    no-room machine clamps to the floor, MemAvailable fallback, available_ram_bytes parsing.
  • Installed from source on a 30.8 GiB CPU box, served end to end: ready in 182 s, clean
    shutdown, nothing left resident.
  • Busy-machine case with RAM actually occupied: idle 4 GiB pool, 10 GiB occupied 2 GiB
    (floor) with the shortfall reported, released 4 GiB.

Distro matrix not run: this release touches none of install.sh, pre-flight or the kernel
build. Existing checkpoints unaffected — no CUDA path, checkpoint format or quantize path
changed.