Skip to content

BatchGen v1.0.9.post3

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 28 Apr 08:36
· 246 commits to main since this release
e12eca3

Highlights

GLM-5-FP8 MoE perf release on top of v1.0.9.post2 (#142).

  • Wire act_quant_3d v2 into the WGMMA pipeline. The post1 release notes claimed an act_quant_3d rewrite from (E,) + serial token loop → (E, mtp), one CTA per (expert, token), but the rewrite landed only in tests/kernels/test_act_quant_3d_v2.py and the standalone batchgen_kernels C extension — production (fp8_wgmma_pipeline.py's inline JIT module) still ran v1, leaving GLM-5-FP8 on a 16-CTA / 132-SM (~12% occupancy) path. post3 swaps the inline kernel for the v2 layout (grid=(E, mtp)) and updates the launcher accordingly.
  • CUDA-graph MoE capture is on by default for GLM-5. BATCHGEN_GLM5_CUDA_GRAPH default flipped from "0" → "1". BATCHGEN_GLM5_CUDA_GRAPH=0 is the new explicit opt-out (eager MoE).
  • Default BATCHGEN_GLM5_MTP_BUCKETS = (128, 256, 512, 768) — the bucket set validated end-to-end this evening on H20 with 2K MMLU + 2K LongBench through the full 128K-decode regime. Empty string opts out to legacy singleton.

Validation: install + 2-node server launch + 2K+2K batch hot-patched against the post2 wheels reached Application startup complete, ran the prefill + decode regime, and submitted both batches without element-size-mismatch / SIGSEGV. post3 is the same source landed in the wheel.

Wheels

  • batchgen-1.0.9.post3 — rebuilt on H20 node 0 from tag v1.0.9.post3.
  • batchgen_kernels-0.3.1.post1 — reused unchanged from v1.0.9.post1.
  • flash_attn_3, flash_mla, deep_gemm — reused unchanged from v1.0.9.

Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).

Install:

RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post3"
pip install \
  "${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
  "${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen-1.0.9.post3-py3-none-any.whl"

Behavior change for operators

If you previously launched GLM-5 servers without BATCHGEN_GLM5_CUDA_GRAPH=1, the post3 server will now capture MoE graphs by default and pre-allocate per-bucket 3D buffers for (128, 256, 512, 768). To restore the old eager-MoE behavior, set BATCHGEN_GLM5_CUDA_GRAPH=0 (or set BATCHGEN_GLM5_MTP_BUCKETS= to an empty string for the legacy singleton path).