Skip to content

NVMAI 3.6

Choose a tag to compare

@Pummelchen Pummelchen released this 16 Aug 12:46
· 994 commits to main since this release

Native Swift and Metal inference for Qwen 3.6 35B-A3B on Apple Silicon.

Model residency

The server can now release the model's memory instead of holding it for the
process lifetime.

  • --lazy-load binds the port immediately and defers the load to the first
    inference request.
  • --idle-unload-seconds <n> releases the weights after n idle seconds,
    measured from when the last request finished. The next request reloads
    transparently. Implies --lazy-load.
  • POST /v1/models/unload releases them on demand, draining in-flight
    requests first so a manual unload never fails a generation.

Measured on M3 24 GB, 4-bit, --max-context 8192:

phase physical footprint
idle, before the first request 2.9 MB
after GET /health 3.1 MB (health never triggers a load)
after the first inference request 5.7 GB
after the idle unload 245 MB

Both flags default to off. The eager path is byte-identical — same startup
banner, same generated output.

Pair --idle-unload-seconds with --prompt-cache-disk: the in-memory prefix
cache is discarded with the session, and the server warns at startup if you
don't.

Correctness

  • Sampler could emit an out-of-range token id. logit_softcap_softmax
    published its cross-SIMD merge through threadgroup memory, but every thread
    wrote the defaults with no barrier before SIMD-group 0 stored the real
    values. A late default store left final_m = -inf and final_inv_d = 0, and
    the normalize loop then produced NaN for the entire probability row —
    after which the top-k reduction found no finite mass and the draw returned a
    sentinel as a token id. In production that is a garbage token, not just a
    flaky test.
  • The -fast alias keeps its strip on cached follow-up turns. The prompt
    cache keyed on the raw request while the KV range it described was prefilled
    from the stripped messages, so a cached continuation could splice an
    unstripped tail onto a stripped prefix and replay the CLI's
    <system-reminder> scaffolding.

Operability

A stale install receipt — what you get after moving or renaming a model
directory — now names the NVMAIRepack --verify-install command that re-issues
it in place, instead of failing opaquely. No re-download.

Production gates

tools/lint.sh (no unaudited as!/try!; a function-length ratchet that also
expires exemptions once a function is fixed) and a ThreadSanitizer job now run
in CI. tools/golden-baseline.sh captures and checks deterministic greedy
generation per quantization — the only check that exercises real inference,
since the unit tests never load a model.

Verification

680 tests / 121 suites pass. Clean release build with 0 compiler warnings.
Golden baselines identical on 4-, 6- and 8-bit.

Downloads

nvmai-3.6-macos-arm64.tar.gz — prebuilt binaries for macOS 26+ on Apple
Silicon. Not code-signed or notarized, so Gatekeeper will block them on
first run; build from source or clear the quarantine attribute yourself after
verifying the checksum. Keep the .bundle directories next to the executables
— they carry the Metal shader library. No model weights are included; install
one with NVMAIRepack.

SHA-256: 4c9e81c8e0ea25a22a65eed5a9506c57f8de564b8028e30c088b8927f6d3b044

Full changelog: https://github.com/Pummelchen/NVMAI/wiki/Changelog