NVMAI 3.9
NVMAI 3.9 — streaming and memory release
Four changes to how weights and context reach the GPU. Each is verified
byte-identical against golden reference output on both 4-bit and 8-bit.
The default context no longer costs throughput
KV storage was sized for --max-context. The shipped default of 262144 reserved
512 MiB per layer — 20.0 GiB across 40 layers. Lazily touched, so it barely
appeared in RSS, but on a 24 GB machine the mappings alone cost real time, even for a
25-token prompt that uses none of it.
Storage now starts at 8192 tokens and doubles on demand, clamped to maxContext.
| ctx 8192 | ctx 262144 | gap | |
|---|---|---|---|
| before | 13.80 s | 22.44 s | 1.63× |
| after | — | — | 1.02× — gone |
A conversation that does reach 262144 pays five buffer copies in total rather than
one reservation up front. Written tokens are carried across every growth, which is
the property the tests guard: losing them would leave a long conversation generating
from corrupted history rather than failing outright.
Expert reads bypass the page cache
The slot budget is now the machine's real footprint, not a number the OS quietly
supplements. Streaming 16.88 GiB with this on left the machine at 78% free memory.
It costs throughput — every cache miss becomes a real device read — and
NVMAI_BOUNDED_IO=0 restores the previous behaviour. It is the default because a
footprint you can account for is the point of streaming a 35B model on 24 GB.
--ram-budget is the knob
--ram-budget 8G # slots derived from this and the model's expert stride
Defaults derive per quantization: 128 slots at 4-bit, 64 at 8-bit. The slot cache
has to hold the routing working set — a trace over 383 real tokens measured 131
distinct experts per layer across a 128-token window — and with page-cache reads
disabled there is no fallback for whatever it does not hold.
--expert-cache-slots still overrides.
6-bit is withdrawn
Across load, kernels, installer, launchers and docs. Its non-power-of-two packing
measured 46.8 GB/s against 60 for both 4-bit and 8-bit, and a 26 GB model does not
fit 24 GB. A 6-bit model now refuses with an explanation and a route forward, not a
generic error — and specifically not a corruption error, which would send you
re-downloading an intact file.
Also included
NVMAIKernelsC, a C99/NEON target for loops where Swift's vector types do not
lower well. Its int4 GEMV is 3.4× the Swift version it replaced.- A parallel
F_NOCACHEexpert reader: 2.43 GB/s on one thread, 3.92 on four,
measured against a working set larger than the page cache. - Instrumentation: true GPU occupancy by merged intervals, per-transition idle
attribution, prefill kernel roles, andNVMAI_ROUTE_TRACEfor real routing dumps.
A note on the numbers
Absolute throughput figures are deliberately absent. The development machine is not
currently quiet enough to measure them honestly — absolute rates there swing about
2× with thermal and memory state, which is documented alongside the method. Every
figure quoted above is either a same-conditions before/after or a device-level
measurement, neither of which depends on machine state.
Full method, every measurement, and the approaches tried and abandoned:
docs/v4-core-design.md
711 tests / 128 suites.
Binaries
macOS 26+, Apple Silicon only. Not code-signed or notarized — verify the checksum,
then clear quarantine:
xattr -dr com.apple.quarantine /path/to/nvmai-3.9-macos-arm64
nvmai-3.9-macos-arm64.tar.gz sha256:
2391d09ee93ed00052ed63abe84a1456bda035ae6cdaf853291364c2d2b9d261