BatchGen v1.0.10.post2
·
214 commits
to main
since this release
BatchGen v1.0.10.post2
What's New
- Adds the opt-in GLM-5-FP8 full-DSA CUDA graph path, which captures the full decode attention segment through projections, DSA selection/FlashMLA, KV writes, absorb, and final attention projection while preserving the existing public attention input/output API.
- Reduces full-DSA CUDA graph retained HBM by sharing large scratch buffers across GLM-5 attention layers and releasing static buffers when graph buckets are dropped.
- Keeps the previously released segmented DSA/MoE CUDA graph path available; full-DSA graph is explicitly gated so deployments can choose the safer segmented path or the new full-DSA path.
Compatibility and Installation
- Python: 3.11+.
- CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
- Wheels: this release publishes a new
batchgenwheel and reuses the unchanged Hopper dependency wheels fromv1.0.10.post1. - Install: use the wheel assets from the GitHub release page for
v1.0.10.post2.
Notes
For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8 with the full-DSA CUDA graph path, the recommended server-side flags are:
BATCHGEN_GLM5_DSA_FULL_CUDA_GRAPH=1
BATCHGEN_GLM5_WHOLE_MODEL_CUDA_GRAPH=0
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93These flags select the full-DSA attention CUDA graph path plus the public CUDA graph serving interface. Whole-model CUDA graph remains separate from this release path.