BatchGen v1.0.10.post4
·
162 commits
to main
since this release
BatchGen v1.0.10.post4
What's New
- Adds GLM whole-model CUDA graph support for the GLM-5 FP8 release path. Users can enable the production graph path with
--enable-cuda-graphand configure bucket coverage with--cuda-graph-max-bucket-sizeand--cuda-graph-num-buckets; GLM-specific graph-selection environment variables are no longer required for the default path. See #155. - Captures configured GLM decode buckets during decode configuration instead of recapturing in the decode loop. Over-bucket or stale graph state uses eager decode fallback rather than hot-path graph capture. See #155.
- Expands and hardens the GLM graph implementation across DSA, MoE, decoder-layer, and whole-model graph boundaries, including safer empty-rank handling, replay comparison coverage, graph-buffer lifecycle handling, and synchronized bucket capture. See #155 and #158.
- Keeps GLM graph-path diagnostic logging quiet by default so normal decode runs avoid the previous verbose graph-path log stream. See #155.
- Updates the GLM graph release compatibility floor to
batchgen_kernels0.3.3 and aligns the packaged kernel source version. See #156 and #157. - Reverts the MiniMax MoE decode buffer growth change from this release line. See #151 and #152.
Compatibility and Installation
- Python: unchanged from the previous BatchGen post release.
- CUDA/PyTorch: unchanged from the previous BatchGen post release.
- Kernel package: GLM graph support in this release requires the matching
batchgen_kernels0.3.3 wheel. - Wheels: release assets attached to GitHub release
v1.0.10.post4. - Install: use the wheel assets from the GitHub release page.
Notes
- Docker image publication is intentionally outside the scope of this formal wheel release.