Skip to content

BatchGen v1.0.10.post4

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 18 May 06:55
· 162 commits to main since this release

BatchGen v1.0.10.post4

What's New

  • Adds GLM whole-model CUDA graph support for the GLM-5 FP8 release path. Users can enable the production graph path with --enable-cuda-graph and configure bucket coverage with --cuda-graph-max-bucket-size and --cuda-graph-num-buckets; GLM-specific graph-selection environment variables are no longer required for the default path. See #155.
  • Captures configured GLM decode buckets during decode configuration instead of recapturing in the decode loop. Over-bucket or stale graph state uses eager decode fallback rather than hot-path graph capture. See #155.
  • Expands and hardens the GLM graph implementation across DSA, MoE, decoder-layer, and whole-model graph boundaries, including safer empty-rank handling, replay comparison coverage, graph-buffer lifecycle handling, and synchronized bucket capture. See #155 and #158.
  • Keeps GLM graph-path diagnostic logging quiet by default so normal decode runs avoid the previous verbose graph-path log stream. See #155.
  • Updates the GLM graph release compatibility floor to batchgen_kernels 0.3.3 and aligns the packaged kernel source version. See #156 and #157.
  • Reverts the MiniMax MoE decode buffer growth change from this release line. See #151 and #152.

Compatibility and Installation

  • Python: unchanged from the previous BatchGen post release.
  • CUDA/PyTorch: unchanged from the previous BatchGen post release.
  • Kernel package: GLM graph support in this release requires the matching batchgen_kernels 0.3.3 wheel.
  • Wheels: release assets attached to GitHub release v1.0.10.post4.
  • Install: use the wheel assets from the GitHub release page.

Notes

  • Docker image publication is intentionally outside the scope of this formal wheel release.