BatchGen v1.0.9.post3
Highlights
GLM-5-FP8 MoE perf release on top of v1.0.9.post2 (#142).
- Wire
act_quant_3dv2 into the WGMMA pipeline. The post1 release notes claimed anact_quant_3drewrite from(E,) + serial token loop→(E, mtp), one CTA per (expert, token), but the rewrite landed only intests/kernels/test_act_quant_3d_v2.pyand the standalonebatchgen_kernelsC extension — production (fp8_wgmma_pipeline.py's inline JIT module) still ran v1, leaving GLM-5-FP8 on a 16-CTA / 132-SM (~12% occupancy) path. post3 swaps the inline kernel for the v2 layout (grid=(E, mtp)) and updates the launcher accordingly. - CUDA-graph MoE capture is on by default for GLM-5.
BATCHGEN_GLM5_CUDA_GRAPHdefault flipped from"0"→"1".BATCHGEN_GLM5_CUDA_GRAPH=0is the new explicit opt-out (eager MoE). - Default
BATCHGEN_GLM5_MTP_BUCKETS = (128, 256, 512, 768)— the bucket set validated end-to-end this evening on H20 with 2K MMLU + 2K LongBench through the full 128K-decode regime. Empty string opts out to legacy singleton.
Validation: install + 2-node server launch + 2K+2K batch hot-patched against the post2 wheels reached Application startup complete, ran the prefill + decode regime, and submitted both batches without element-size-mismatch / SIGSEGV. post3 is the same source landed in the wheel.
Wheels
batchgen-1.0.9.post3— rebuilt on H20 node 0 from tagv1.0.9.post3.batchgen_kernels-0.3.1.post1— reused unchanged from v1.0.9.post1.flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post3"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post3-py3-none-any.whl"Behavior change for operators
If you previously launched GLM-5 servers without BATCHGEN_GLM5_CUDA_GRAPH=1, the post3 server will now capture MoE graphs by default and pre-allocate per-bucket 3D buffers for (128, 256, 512, 768). To restore the old eager-MoE behavior, set BATCHGEN_GLM5_CUDA_GRAPH=0 (or set BATCHGEN_GLM5_MTP_BUCKETS= to an empty string for the legacy singleton path).