BatchGen v1.0.10
BatchGen v1.0.10
Highlights
- Adds GLM-5-FP8 segmented CUDA graph decode support for the DSA and MoE paths.
- Adds graph-safe DSA selected-KV handling with FlashMLA metadata passed as replay inputs.
- Adds fixed-capacity CUDA graph page-table storage so graph capture does not mutate active decode page tables.
- Adds GLM-5 MoE CUDA graph segment support and stricter graph state handling after model memory cleanup.
- Includes focused CUDA graph tests for DSA projection/selection, FlashMLA metadata, page-table stability, DSA graph replay, and MoE graph replay.
Packaging
batchgenis released as1.0.10.batchgen_kernelsis released as0.3.2+sm90a.- FlashAttention, FlashMLA, and DeepGEMM wheels are reused from the previous compatible release because their pinned versions did not change.
- This is a wheel release; no Docker image is published for this release.
Runtime notes
For the segmented GLM-5-FP8 graph path, enable segmented graph mode and the GLM-5 DSA/MoE graph flags. Whole-model CUDA graph remains separate from this release path.
Assets
The release includes BatchGen, BatchGen kernels, reused dependency wheels, and SHA256SUMS for verification.