Skip to content

TurboQuant v1.5.3 — Double WHT Per-Head for D=64

Choose a tag to compare

@AmesianX AmesianX released this 11 Apr 17:06
· 939 commits to main since this release

TurboQuant v1.5.3

Key Changes

  • Double WHT per-head (D=64): Cross-head WHT abandoned. S1→WHT64→S2→WHT64 per-head.
  • QJL re-enabled for D=64: 1-bit correction critical for multi-turn (9+ turns verified).

Recommended Settings

Tested:

Model head_dim K cache V cache Status
Gemma 4 26B (MoE) 256/512 tbqp3 tbq3 ✅ Tested
GLM-4.7-Flash (MLA) 576(K)/512(V) tbqp3 tbq3 ✅ Tested
GPT-OSS 120B 64 tbq4 tbq3 ✅ Tested (35/35)

Should work (not tested by us):

Model head_dim K cache V cache
Llama 3.x 128 tbqp3 tbq3
Qwen 2.5 / Qwen 3 128 tbqp3 tbq3
Mistral 128 tbqp3 tbq3

head_dim=64 모델: 3-bit K(tbqp3)는 한국어 대화/멀티턴은 지원하나 행렬 연산 정밀도 부족. 4-bit K(tbq4) 사용 권장.

Math Accuracy (GPT-OSS 120B, head_dim=64, 35 problems, temp=0)

Config K V Math Korean Multi-turn
f16/f16 f16 f16 35/35 ✅ ✅
tbq4/tbq3 tbq4_2 tbq3_2 35/35 ✅ ✅
tbqp3/tbq3 tbqp3_3 tbq3_2 ❌ (matrix) ✅ ✅ (9+ turns)

Run Options

# Required for ALL TurboQuant configurations:
--flash-attn on --n-gpu-layers 999

# Gemma 4 26B
--cache-type-k tbqp3 --cache-type-v tbq3

# GLM-4.7-Flash
--cache-type-k tbqp3 --cache-type-v tbq3

# GPT-OSS 120B (head_dim=64)
--cache-type-k tbq4 --cache-type-v tbq3

# Llama 3.x / Qwen / Mistral (head_dim=128, not tested by us)
--cache-type-k tbqp3 --cache-type-v tbq3

⚠️ Requirements: CUDA only (no ROCm/Metal/CPU). Flash attention ON. All layers on GPU. No MoE CPU offloading.

Target GPUs

sm_70 (V100), sm_80 (A100), sm_86 (3090 Ti), sm_89 (4090), sm_90 (H100), sm_100/120 (5090/DGX Spark)

Build from source

git clone https://github.com/AmesianX/TurboQuant.git
cd TurboQuant
mkdir build && cd build

Full build (all TBQ instances — slower compile, supports all K/V combinations):

cmake -DCMAKE_BUILD_TYPE=Release \
      -DBUILD_SHARED_LIBS=ON \
      -DGGML_CUDA=ON \
      -DGGML_BLAS=ON \
      -DGGML_F16C=ON \
      -DGGML_FMA=ON \
      -DLLAMA_BUILD_TESTS=OFF \
      -DLLAMA_BUILD_EXAMPLES=OFF \
      -DLLAMA_BUILD_SERVER=ON \
      -DGGML_CUDA_GRAPHS=ON \
      ..
make -j$(nproc)

TBQ tuning build (essential instances only — faster compile, recommended for development):

cmake -DCMAKE_BUILD_TYPE=Release \
      -DBUILD_SHARED_LIBS=ON \
      -DGGML_CUDA=ON \
      -DGGML_BLAS=ON \
      -DGGML_F16C=ON \
      -DGGML_FMA=ON \
      -DLLAMA_BUILD_TESTS=OFF \
      -DLLAMA_BUILD_EXAMPLES=OFF \
      -DLLAMA_BUILD_SERVER=ON \
      -DGGML_CUDA_GRAPHS=ON \
      -DGGML_CUDA_FA_TBQ_TUNING=ON \
      ..
make -j$(nproc)

Full changelog: v1.5.2...v1.5.3