TurboQuant v1.5.3 — Double WHT Per-Head for D=64
TurboQuant v1.5.3
Key Changes
- Double WHT per-head (D=64): Cross-head WHT abandoned. S1→WHT64→S2→WHT64 per-head.
- QJL re-enabled for D=64: 1-bit correction critical for multi-turn (9+ turns verified).
Recommended Settings
Tested:
| Model | head_dim | K cache | V cache | Status |
|---|---|---|---|---|
| Gemma 4 26B (MoE) | 256/512 | tbqp3 |
tbq3 |
✅ Tested |
| GLM-4.7-Flash (MLA) | 576(K)/512(V) | tbqp3 |
tbq3 |
✅ Tested |
| GPT-OSS 120B | 64 | tbq4 |
tbq3 |
✅ Tested (35/35) |
Should work (not tested by us):
| Model | head_dim | K cache | V cache |
|---|---|---|---|
| Llama 3.x | 128 | tbqp3 |
tbq3 |
| Qwen 2.5 / Qwen 3 | 128 | tbqp3 |
tbq3 |
| Mistral | 128 | tbqp3 |
tbq3 |
head_dim=64 모델: 3-bit K(tbqp3)는 한국어 대화/멀티턴은 지원하나 행렬 연산 정밀도 부족. 4-bit K(tbq4) 사용 권장.
Math Accuracy (GPT-OSS 120B, head_dim=64, 35 problems, temp=0)
| Config | K | V | Math | Korean | Multi-turn |
|---|---|---|---|---|---|
| f16/f16 | f16 | f16 | 35/35 | ✅ | ✅ |
| tbq4/tbq3 | tbq4_2 | tbq3_2 | 35/35 | ✅ | ✅ |
| tbqp3/tbq3 | tbqp3_3 | tbq3_2 | ❌ (matrix) | ✅ | ✅ (9+ turns) |
Run Options
# Required for ALL TurboQuant configurations:
--flash-attn on --n-gpu-layers 999
# Gemma 4 26B
--cache-type-k tbqp3 --cache-type-v tbq3
# GLM-4.7-Flash
--cache-type-k tbqp3 --cache-type-v tbq3
# GPT-OSS 120B (head_dim=64)
--cache-type-k tbq4 --cache-type-v tbq3
# Llama 3.x / Qwen / Mistral (head_dim=128, not tested by us)
--cache-type-k tbqp3 --cache-type-v tbq3
⚠️ Requirements: CUDA only (no ROCm/Metal/CPU). Flash attention ON. All layers on GPU. No MoE CPU offloading.
Target GPUs
sm_70 (V100), sm_80 (A100), sm_86 (3090 Ti), sm_89 (4090), sm_90 (H100), sm_100/120 (5090/DGX Spark)
Build from source
git clone https://github.com/AmesianX/TurboQuant.git
cd TurboQuant
mkdir build && cd buildFull build (all TBQ instances — slower compile, supports all K/V combinations):
cmake -DCMAKE_BUILD_TYPE=Release \
-DBUILD_SHARED_LIBS=ON \
-DGGML_CUDA=ON \
-DGGML_BLAS=ON \
-DGGML_F16C=ON \
-DGGML_FMA=ON \
-DLLAMA_BUILD_TESTS=OFF \
-DLLAMA_BUILD_EXAMPLES=OFF \
-DLLAMA_BUILD_SERVER=ON \
-DGGML_CUDA_GRAPHS=ON \
..
make -j$(nproc)TBQ tuning build (essential instances only — faster compile, recommended for development):
cmake -DCMAKE_BUILD_TYPE=Release \
-DBUILD_SHARED_LIBS=ON \
-DGGML_CUDA=ON \
-DGGML_BLAS=ON \
-DGGML_F16C=ON \
-DGGML_FMA=ON \
-DLLAMA_BUILD_TESTS=OFF \
-DLLAMA_BUILD_EXAMPLES=OFF \
-DLLAMA_BUILD_SERVER=ON \
-DGGML_CUDA_GRAPHS=ON \
-DGGML_CUDA_FA_TBQ_TUNING=ON \
..
make -j$(nproc)Full changelog: v1.5.2...v1.5.3