Skip to content

v0.13.1

@wtdcode wtdcode tagged this 16 Sep 08:42
Fixes on top of v0.13.0:
- mamba align state seed used an unresolved block size, corrupting output or
  crashing on prefix-cache hits (#77, issues #80/#84).
- compressed-tensors PLE table no longer requires an `ignore` entry (#72).
- fp8 PLE rows move as bytes in the pinned-host lookup, so Qwen3.8-Flash-Next
  FP8 with engram cpu_offload starts on sm80/sm86.
- shared experts overlap on the aux stream on CUDA; DSv4-Flash decode is ~40%
  faster at TP8 (issue #85).
- split-K decode kernel uses 8 warps on CUDA; VLLM_DSV4_FIXED_DECODE_SPLITS
  is wired up and defaults to 0.
- VLLM_DETERMINISTIC_MOE_ALIGN defaults to 0.
- VLLM_UNREPLICATE_ATTN_GEMMS=1 no longer aborts at startup.
Assets 2
Loading