Skip to content

OrbitQuant 0.9.6

Choose a tag to compare

@iamwavecut iamwavecut released this 12 Sep 02:03

Packed W4A4 CUDA decode now reuses each decoded weight vector across two activation rows. Native 1.0.3 selects this path for row-major weights, 2–8 rows, N>=2048 and K=1024–16384, while retaining the existing single-row and smaller-output kernels.

On RTX PRO 4500 Blackwell, YuE2 CFG2 fixed-token decode improved from 3.9533 to 3.6541 ms per step. An interleaved full-song comparison at CFG1.5 generated identical 183.96 s audio in 33.40 versus 31.78 s on average. These measurements do not imply a gain for CFG1 or other GPUs. Bias uses explicit FP32 fused multiply-add, so last-bit differences from older compiler-dependent epilogues are possible.

The Transformers adapter also fixes GPT-2 and GPTBigCode loading when upstream initialization accesses an already packed c_proj.weight. Missing unquantized child parameters still initialize normally.

Validation: 579 local tests passed; current, released and development Hugging Face compatibility gates passed; CUDA and Metal backend gates passed; the published CUDA wheel passed 208 tests. All 19 native variants built successfully. The published Python wheel passed an independent 13-case architecture/load/save/restore check.

Install orbitquant==0.9.6. The native kernel is a separate package: use the matching 1.0.3 wheel from kernels-v1. Existing installed or cached native packages retain precedence and are not replaced by upgrading the Python package alone.