Skip to content

BatchGen v1.0.10.post1

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 07 May 03:13
· 221 commits to main since this release
7f41304

BatchGen v1.0.10.post1

What's New

  • Makes GLM-5-FP8 and GLM-5.1-FP8 segmented CUDA graph serving configurable through server CLI flags, so --enable-cuda-graph enables the DSA and MoE graph paths without requiring graph-specific environment variables. See #149.
  • Tightens the GLM-5 graph path contract so DSA and MoE graph requirements are propagated end-to-end from server startup through worker setup and model wrappers. See #149.
  • Adds fixes for decode metadata stability, DSA GPU KV loading, DSA graph page-table refresh, FlashMLA replay metadata preparation, and captured-sequence-length handling in the segmented graph path. See #149.

Compatibility and Installation

  • Python: 3.11+.
  • CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
  • Wheels: this release publishes a new batchgen wheel and reuses the compatible Hopper dependency wheels from v1.0.10.
  • Install: use the wheel assets from the GitHub release page for v1.0.10.post1.

Notes

For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8, the recommended server-side CUDA graph flags are:

--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93

These flags select the segmented DSA and MoE graph path through the public server interface. Whole-model CUDA graph remains separate from this release path.