Skip to content

1.5.0

Latest

Choose a tag to compare

@github-actions github-actions released this 13 Sep 14:00
· 1 commit to master since this release
  • Fix NemotronHForCausalLM (support Nemotron-3-Super)
  • Faster quantization
  • Experimental zero-copy pinned arena mode for CPU offload (Linux only for now)
  • Faster MoE prefill and decode (see below)
  • Improved speculative decoding for MoE models
  • Faster dense model prefill on some architectures
  • Remove some sources of nondeterminism
  • Bugfixes and cleanup
Test case 4k prefill, t/s decode, t/s, S=1
Qwen3.8-Flash-Next, 3.05bpw
Pro 6000, ngram_ram
5788 → 7327 (+27%) 101.04 → 115.27 (+14%)
Qwen3.8-Flash-Next, 3.05bpw
Pro 6000 (gen5 x16) + TR 7960X, 80% CPU offload, pinned arena, ngram_ram
2801 → 4323 (+54%) 55.66 → 66.65 (+20%)
Qwen3.8-Flash-Next, 3.05bpw
5090 (gen5 x8) + TR 7960X, 80% CPU offload, pinned arena, ngram_ram
2615 → 3256 (+25%) 53.42 → 63.27 (+18%)
Qwen3.8-Flash-Next, 3.05bpw
4090 (gen4 x4) + TR 7960X, 80% CPU offload, pinned arena, ngram_ram
1165 → 1204 (+3%) 56.07 → 59.52 (+6%)
Qwen3.8-Flash-Next, 3.05bpw
3090Ti+3090 (gen4 x4) + AVX2, 8 cores, 20% CPU offload, pinned arena, ngram_ram
1572 → 1841 (+17%) 58.40 → 62.77 (+7%)
Gemma4-26B-A3B, 3.10bpw
Pro 6000
15626 → 18679 (+20%) 185.45 → 202.32 (+9%)
Gemma4-26B-A3B, 3.10bpw
5090
12066 → 14695 (+22%) 180.77 → 199.55 (+10%)
Gemma4-26B-A3B, 3.10bpw
4090
8779 → 10605 (+21%) 172.92 → 181.72 (+5%)
Gemma4-26B-A3B, 3.10bpw
3090
4742 → 5980 (+26%) 129.77 → 134.82 (+4%)
GLM5.3-Flash, 2.05bpw
Pro 6000
3446 → 3938 (+14%) 90.18 → 93.65 (+4%)
GLM5.3-Flash, 3.05bpw
5090 (gen5 x8) + TR 7960X, 80% CPU offload, pinned arena
809 → 1056 (+30%) 30.53 → 30.61 (+0%)
GLM5.3-Flash, 3.05bpw
5090 (gen5 x8) + AVX2, 8 cores, 80% CPU offload, pinned arena
841 → 1068 (+27%) 9.78 → 10.30 (+5%)
Qwen3.8-27B, 4.00bpw
Pro 6000
5549 → 5547 (-0%) 76.88 → 76.53 (-0%)
Qwen3.8-27B, 4.00bpw
5090
3766 → 4651 (+24%) 78.54 → 78.58 (+0%)
Qwen3.8-27B, 4.00bpw
4090
2636 → 2819 (+7%) 54.60 → 54.61 (+0%)
Qwen3.8-27B, 4.00bpw
3090Ti
1442 → 1552 (+8%) 48.55 → 48.66 (+0%)
gpt-oss-20b, 3.00bpw
Pro 6000
22229 → 22249 (+0%) 263.77 → 303.27 (+15%)
gpt-oss-20b, 3.00bpw
5090
16579 → 17936 (+8%) 255.68 → 287.21 (+12%)
gpt-oss-20b, 3.00bpw
4090
12520 → 12703 (+1%) 222.59 → 242.71 (+9%)
gpt-oss-20b, 3.00bpw
3090Ti
6455 → 7087 (+10%) 158.09 → 177.93 (+13%)

Full Changelog: v1.4.9...v1.5.0