NVMAI 3.6
Native Swift and Metal inference for Qwen 3.6 35B-A3B on Apple Silicon.
Model residency
The server can now release the model's memory instead of holding it for the
process lifetime.
--lazy-loadbinds the port immediately and defers the load to the first
inference request.--idle-unload-seconds <n>releases the weights afternidle seconds,
measured from when the last request finished. The next request reloads
transparently. Implies--lazy-load.POST /v1/models/unloadreleases them on demand, draining in-flight
requests first so a manual unload never fails a generation.
Measured on M3 24 GB, 4-bit, --max-context 8192:
| phase | physical footprint |
|---|---|
| idle, before the first request | 2.9 MB |
after GET /health |
3.1 MB (health never triggers a load) |
| after the first inference request | 5.7 GB |
| after the idle unload | 245 MB |
Both flags default to off. The eager path is byte-identical — same startup
banner, same generated output.
Pair --idle-unload-seconds with --prompt-cache-disk: the in-memory prefix
cache is discarded with the session, and the server warns at startup if you
don't.
Correctness
- Sampler could emit an out-of-range token id.
logit_softcap_softmax
published its cross-SIMD merge through threadgroup memory, but every thread
wrote the defaults with no barrier before SIMD-group 0 stored the real
values. A late default store leftfinal_m = -infandfinal_inv_d = 0, and
the normalize loop then produced NaN for the entire probability row —
after which the top-k reduction found no finite mass and the draw returned a
sentinel as a token id. In production that is a garbage token, not just a
flaky test. - The
-fastalias keeps its strip on cached follow-up turns. The prompt
cache keyed on the raw request while the KV range it described was prefilled
from the stripped messages, so a cached continuation could splice an
unstripped tail onto a stripped prefix and replay the CLI's
<system-reminder>scaffolding.
Operability
A stale install receipt — what you get after moving or renaming a model
directory — now names the NVMAIRepack --verify-install command that re-issues
it in place, instead of failing opaquely. No re-download.
Production gates
tools/lint.sh (no unaudited as!/try!; a function-length ratchet that also
expires exemptions once a function is fixed) and a ThreadSanitizer job now run
in CI. tools/golden-baseline.sh captures and checks deterministic greedy
generation per quantization — the only check that exercises real inference,
since the unit tests never load a model.
Verification
680 tests / 121 suites pass. Clean release build with 0 compiler warnings.
Golden baselines identical on 4-, 6- and 8-bit.
Downloads
nvmai-3.6-macos-arm64.tar.gz — prebuilt binaries for macOS 26+ on Apple
Silicon. Not code-signed or notarized, so Gatekeeper will block them on
first run; build from source or clear the quarantine attribute yourself after
verifying the checksum. Keep the .bundle directories next to the executables
— they carry the Metal shader library. No model weights are included; install
one with NVMAIRepack.
SHA-256: 4c9e81c8e0ea25a22a65eed5a9506c57f8de564b8028e30c088b8927f6d3b044
Full changelog: https://github.com/Pummelchen/NVMAI/wiki/Changelog