Skip to content
This repository was archived by the owner on Oct 3, 2026. It is now read-only.

0.1.96 — Faster Mac GPU queue submission

Choose a tag to compare

@Geramy Geramy released this 24 Sep 14:15

Mac HRX now uses qualified same-queue GPU barriers automatically, preserving software deferral for cross-queue work and pending host callbacks. No experimental switch is required.

Measured on Apple Silicon, Thunderbolt 5 and Radeon AI PRO R9700 with Qwen3.8-27B MLX Q6:

  • Final default: 17.51 decode tokens/s and 88.93 prompt tokens/s, median of three requests after two warmups, 512 input / 129 output, MTP disabled.
  • Matched off/on/on/off comparison: 4.43% faster decoding, twenty identical responses, zero compilation or disk-cache misses in measured requests.
  • Validated ring wrap, cross-queue ordering, host callbacks, error propagation and actual queue-resource backpressure; host regression tests pass ASan/UBSan.

Host/driver bundle version is 0.1.96 (196). The app builds and its development signature verifies. GPU firmware and initialization sequences are unchanged; the runtime performance qualification used installed driver 195. Rebuild HRX with scripts/build-hrx-macos.sh to obtain the userspace improvement.

The source archives below contain the release. Build/sign the host using the repository instructions. Prefill and INT8 experiments are saved on separate LSE testing branches and remain outside the default. See the optimization checkpoint for evidence and remaining work.