Skip to content

BatchGen v1.0.10.post2

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 07 May 11:54
· 214 commits to main since this release

BatchGen v1.0.10.post2

What's New

  • Adds the opt-in GLM-5-FP8 full-DSA CUDA graph path, which captures the full decode attention segment through projections, DSA selection/FlashMLA, KV writes, absorb, and final attention projection while preserving the existing public attention input/output API.
  • Reduces full-DSA CUDA graph retained HBM by sharing large scratch buffers across GLM-5 attention layers and releasing static buffers when graph buckets are dropped.
  • Keeps the previously released segmented DSA/MoE CUDA graph path available; full-DSA graph is explicitly gated so deployments can choose the safer segmented path or the new full-DSA path.

Compatibility and Installation

  • Python: 3.11+.
  • CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
  • Wheels: this release publishes a new batchgen wheel and reuses the unchanged Hopper dependency wheels from v1.0.10.post1.
  • Install: use the wheel assets from the GitHub release page for v1.0.10.post2.

Notes

For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8 with the full-DSA CUDA graph path, the recommended server-side flags are:

BATCHGEN_GLM5_DSA_FULL_CUDA_GRAPH=1
BATCHGEN_GLM5_WHOLE_MODEL_CUDA_GRAPH=0
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93

These flags select the full-DSA attention CUDA graph path plus the public CUDA graph serving interface. Whole-model CUDA graph remains separate from this release path.