BatchGen v1.0.10.post1
·
221 commits
to main
since this release
BatchGen v1.0.10.post1
What's New
- Makes GLM-5-FP8 and GLM-5.1-FP8 segmented CUDA graph serving configurable through server CLI flags, so
--enable-cuda-graphenables the DSA and MoE graph paths without requiring graph-specific environment variables. See #149. - Tightens the GLM-5 graph path contract so DSA and MoE graph requirements are propagated end-to-end from server startup through worker setup and model wrappers. See #149.
- Adds fixes for decode metadata stability, DSA GPU KV loading, DSA graph page-table refresh, FlashMLA replay metadata preparation, and captured-sequence-length handling in the segmented graph path. See #149.
Compatibility and Installation
- Python: 3.11+.
- CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
- Wheels: this release publishes a new
batchgenwheel and reuses the compatible Hopper dependency wheels fromv1.0.10. - Install: use the wheel assets from the GitHub release page for
v1.0.10.post1.
Notes
For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8, the recommended server-side CUDA graph flags are:
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93These flags select the segmented DSA and MoE graph path through the public server interface. Whole-model CUDA graph remains separate from this release path.