Skip to content

OrbitQuant 0.9.3

Choose a tag to compare

@iamwavecut iamwavecut released this 11 Sep 21:02

OrbitQuant 0.9.3 adds a vectorized DP4A CUDA GEMV for packed W4A4 decode. It avoids mostly empty Tensor Core tiles for 1–8 rows while retaining packed weights and the existing norm/scale/bias arithmetic. Larger batches and K-major weights retain the previous dispatch.

On RTX PRO 6000 Blackwell with PyTorch 2.10/CUDA 12.8, the tested small-row operators were 3.4–55.9x faster; these are operator timings, not whole-model speedups. All eight benchmark output hashes matched. The release CUDA wheel passed 128 GPU tests (30 tests for other backends skipped), and a strict Python ABI3 audit.

Install orbitquant==0.9.3 and update the matching native kernel binary to 1.0.1 from the kernels-v1 release, or rebuild from the bundled sources. An already installed older native binary remains compatible but does not gain the optimization. Set ORBITQUANT_W4A4_DISABLE_GEMV=1 before process startup to restore the previous kernel path.

Benchmarks and runtime contract · PyPI · Native binaries

Compatibility note: the optional latest-HF overlay check finds a GPT-2 restore failure with Transformers 5.17.0 and current development main. The same three failures were reproduced with 0.9.2 source; this CUDA change does not address that existing compatibility issue. The locked dependency suite and current-HF checks passed.