Skip to content

BatchGen v1.0.10

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 06 May 05:19
· 238 commits to main since this release
8f592e9

BatchGen v1.0.10

Highlights

  • Adds GLM-5-FP8 segmented CUDA graph decode support for the DSA and MoE paths.
  • Adds graph-safe DSA selected-KV handling with FlashMLA metadata passed as replay inputs.
  • Adds fixed-capacity CUDA graph page-table storage so graph capture does not mutate active decode page tables.
  • Adds GLM-5 MoE CUDA graph segment support and stricter graph state handling after model memory cleanup.
  • Includes focused CUDA graph tests for DSA projection/selection, FlashMLA metadata, page-table stability, DSA graph replay, and MoE graph replay.

Packaging

  • batchgen is released as 1.0.10.
  • batchgen_kernels is released as 0.3.2+sm90a.
  • FlashAttention, FlashMLA, and DeepGEMM wheels are reused from the previous compatible release because their pinned versions did not change.
  • This is a wheel release; no Docker image is published for this release.

Runtime notes

For the segmented GLM-5-FP8 graph path, enable segmented graph mode and the GLM-5 DSA/MoE graph flags. Whole-model CUDA graph remains separate from this release path.

Assets

The release includes BatchGen, BatchGen kernels, reused dependency wheels, and SHA256SUMS for verification.