Releases: batchgen-project/batchgen
Release list
BatchGen v1.0.11
Prebuilt Hopper (sm90a) wheel set for the reference env (py3.11, torch 2.9.0+cu128, CUDA 12.8): FA3, FlashMLA, DeepGEMM,
batchgen_kernels, batchgen. Enables install_deps.sh wheel fast-path.
BatchGen v1.0.10.post4
BatchGen v1.0.10.post4
What's New
- Adds GLM whole-model CUDA graph support for the GLM-5 FP8 release path. Users can enable the production graph path with
--enable-cuda-graphand configure bucket coverage with--cuda-graph-max-bucket-sizeand--cuda-graph-num-buckets; GLM-specific graph-selection environment variables are no longer required for the default path. See #155. - Captures configured GLM decode buckets during decode configuration instead of recapturing in the decode loop. Over-bucket or stale graph state uses eager decode fallback rather than hot-path graph capture. See #155.
- Expands and hardens the GLM graph implementation across DSA, MoE, decoder-layer, and whole-model graph boundaries, including safer empty-rank handling, replay comparison coverage, graph-buffer lifecycle handling, and synchronized bucket capture. See #155 and #158.
- Keeps GLM graph-path diagnostic logging quiet by default so normal decode runs avoid the previous verbose graph-path log stream. See #155.
- Updates the GLM graph release compatibility floor to
batchgen_kernels0.3.3 and aligns the packaged kernel source version. See #156 and #157. - Reverts the MiniMax MoE decode buffer growth change from this release line. See #151 and #152.
Compatibility and Installation
- Python: unchanged from the previous BatchGen post release.
- CUDA/PyTorch: unchanged from the previous BatchGen post release.
- Kernel package: GLM graph support in this release requires the matching
batchgen_kernels0.3.3 wheel. - Wheels: release assets attached to GitHub release
v1.0.10.post4. - Install: use the wheel assets from the GitHub release page.
Notes
- Docker image publication is intentionally outside the scope of this formal wheel release.
BatchGen v1.0.10.post2
BatchGen v1.0.10.post2
What's New
- Adds the opt-in GLM-5-FP8 full-DSA CUDA graph path, which captures the full decode attention segment through projections, DSA selection/FlashMLA, KV writes, absorb, and final attention projection while preserving the existing public attention input/output API.
- Reduces full-DSA CUDA graph retained HBM by sharing large scratch buffers across GLM-5 attention layers and releasing static buffers when graph buckets are dropped.
- Keeps the previously released segmented DSA/MoE CUDA graph path available; full-DSA graph is explicitly gated so deployments can choose the safer segmented path or the new full-DSA path.
Compatibility and Installation
- Python: 3.11+.
- CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
- Wheels: this release publishes a new
batchgenwheel and reuses the unchanged Hopper dependency wheels fromv1.0.10.post1. - Install: use the wheel assets from the GitHub release page for
v1.0.10.post2.
Notes
For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8 with the full-DSA CUDA graph path, the recommended server-side flags are:
BATCHGEN_GLM5_DSA_FULL_CUDA_GRAPH=1
BATCHGEN_GLM5_WHOLE_MODEL_CUDA_GRAPH=0
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93These flags select the full-DSA attention CUDA graph path plus the public CUDA graph serving interface. Whole-model CUDA graph remains separate from this release path.
BatchGen v1.0.10.post1
BatchGen v1.0.10.post1
What's New
- Makes GLM-5-FP8 and GLM-5.1-FP8 segmented CUDA graph serving configurable through server CLI flags, so
--enable-cuda-graphenables the DSA and MoE graph paths without requiring graph-specific environment variables. See #149. - Tightens the GLM-5 graph path contract so DSA and MoE graph requirements are propagated end-to-end from server startup through worker setup and model wrappers. See #149.
- Adds fixes for decode metadata stability, DSA GPU KV loading, DSA graph page-table refresh, FlashMLA replay metadata preparation, and captured-sequence-length handling in the segmented graph path. See #149.
Compatibility and Installation
- Python: 3.11+.
- CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
- Wheels: this release publishes a new
batchgenwheel and reuses the compatible Hopper dependency wheels fromv1.0.10. - Install: use the wheel assets from the GitHub release page for
v1.0.10.post1.
Notes
For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8, the recommended server-side CUDA graph flags are:
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93These flags select the segmented DSA and MoE graph path through the public server interface. Whole-model CUDA graph remains separate from this release path.
BatchGen v1.0.10
BatchGen v1.0.10
Highlights
- Adds GLM-5-FP8 segmented CUDA graph decode support for the DSA and MoE paths.
- Adds graph-safe DSA selected-KV handling with FlashMLA metadata passed as replay inputs.
- Adds fixed-capacity CUDA graph page-table storage so graph capture does not mutate active decode page tables.
- Adds GLM-5 MoE CUDA graph segment support and stricter graph state handling after model memory cleanup.
- Includes focused CUDA graph tests for DSA projection/selection, FlashMLA metadata, page-table stability, DSA graph replay, and MoE graph replay.
Packaging
batchgenis released as1.0.10.batchgen_kernelsis released as0.3.2+sm90a.- FlashAttention, FlashMLA, and DeepGEMM wheels are reused from the previous compatible release because their pinned versions did not change.
- This is a wheel release; no Docker image is published for this release.
Runtime notes
For the segmented GLM-5-FP8 graph path, enable segmented graph mode and the GLM-5 DSA/MoE graph flags. Whole-model CUDA graph remains separate from this release path.
Assets
The release includes BatchGen, BatchGen kernels, reused dependency wheels, and SHA256SUMS for verification.
v1.0.9.post5
BatchGen v1.0.9.post5
- Packages GLM-5 DSA Hadamard kernels into batchgen_kernels as AOT extensions to remove the production runtime-JIT failure surface.
- Bumps batchgen_kernels to 0.3.1.post3 and requires it from BatchGen.
- Builds the internal H20 batchgen_kernels wheel with explicit sm90a arch selection.
- Reuses unchanged dependency wheels from v1.0.9.post4.
Validated on two-node H20 GLM-5-FP8 after clearing test conda envs and Torch extension caches: installed post5 wheels in the persistent batchgen env, launched the server, submitted a batch, and confirmed DSA/Hadamard imports came from packaged site-packages .so files with no DSA/Hadamard Torch extension cache entries.
BatchGen v1.0.9.post4
Hotfix release for GLM-5 DSA startup on H20.
Changes:
- Fixes duplicate Torch JIT extension ownership for GLM-5 DSA Hadamard/RoPE kernels.
- Makes
batchgen_kernels.attention.dsa.indexerthe canonical loader with package-scopedbatchgen_dsa_*extension names. - Keeps
batchgen.other_kernels.hadamard_transformas a compatibility shim. - Requires
batchgen_kernels 0.3.1.post2from this release.
BatchGen v1.0.9.post3
Highlights
GLM-5-FP8 MoE perf release on top of v1.0.9.post2 (#142).
- Wire
act_quant_3dv2 into the WGMMA pipeline. The post1 release notes claimed anact_quant_3drewrite from(E,) + serial token loop→(E, mtp), one CTA per (expert, token), but the rewrite landed only intests/kernels/test_act_quant_3d_v2.pyand the standalonebatchgen_kernelsC extension — production (fp8_wgmma_pipeline.py's inline JIT module) still ran v1, leaving GLM-5-FP8 on a 16-CTA / 132-SM (~12% occupancy) path. post3 swaps the inline kernel for the v2 layout (grid=(E, mtp)) and updates the launcher accordingly. - CUDA-graph MoE capture is on by default for GLM-5.
BATCHGEN_GLM5_CUDA_GRAPHdefault flipped from"0"→"1".BATCHGEN_GLM5_CUDA_GRAPH=0is the new explicit opt-out (eager MoE). - Default
BATCHGEN_GLM5_MTP_BUCKETS = (128, 256, 512, 768)— the bucket set validated end-to-end this evening on H20 with 2K MMLU + 2K LongBench through the full 128K-decode regime. Empty string opts out to legacy singleton.
Validation: install + 2-node server launch + 2K+2K batch hot-patched against the post2 wheels reached Application startup complete, ran the prefill + decode regime, and submitted both batches without element-size-mismatch / SIGSEGV. post3 is the same source landed in the wheel.
Wheels
batchgen-1.0.9.post3— rebuilt on H20 node 0 from tagv1.0.9.post3.batchgen_kernels-0.3.1.post1— reused unchanged from v1.0.9.post1.flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post3"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post3-py3-none-any.whl"Behavior change for operators
If you previously launched GLM-5 servers without BATCHGEN_GLM5_CUDA_GRAPH=1, the post3 server will now capture MoE graphs by default and pre-allocate per-bucket 3D buffers for (128, 256, 512, 768). To restore the old eager-MoE behavior, set BATCHGEN_GLM5_CUDA_GRAPH=0 (or set BATCHGEN_GLM5_MTP_BUCKETS= to an empty string for the legacy singleton path).
BatchGen v1.0.9.post2
Highlights
Hotfix release for v1.0.9.post1. Fixes a packaging regression that broke the wheel-only install path on a fresh conda env. See #141.
batchgen/other_kernels/hadamard_transform/csrc/now ships in the wheel. v1.0.9.post1'sMANIFEST.inandsetup.pyonly globbedbatchgen/{core,external}source trees; the GLM-5 indexer's JIT-compiled hadamard / fused_rope_hadamard CUDA kernels need their 7 csrc files at runtime (hadamard_binding.cpp,fast_hadamard_transform_cuda.cu,fused_rope_hadamard*.{cpp,cu}, plus 3 headers). On a wheel-only install the JIT raisesFileNotFoundError, the indexer silently falls back to non-fused RoPE+Hadamard, the fallback returns the input dtype afterk_norm(which yields FP32 in torch 2.9), and the FP32 indexer K tensor tripshost_paged_kv_worker_view.h"K tensor element size does not match configuration" (aux KV is BF16) on the first prefill batch. Worker SIGSEGV with no actionable log line.- Warn loudly on kernel-import failure. WP2/WP4/WP5 indexer kernel imports + the hadamard-csrc imports now log at WARNING (was DEBUG / silent). The next time a packaging gap surfaces, it shows up at server-launch time instead of as an opaque later SIGSEGV.
Wheels
batchgen-1.0.9.post2— rebuilt on H20 node 0 from tagv1.0.9.post2.batchgen_kernels-0.3.1.post1— reused unchanged from v1.0.9.post1.flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post2"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post2-py3-none-any.whl"BatchGen v1.0.9.post1
Highlights
Patch release for GLM-5 / GLM-5.1 FP8 on top of v1.0.9. Quiets noisy decode/startup logs that were overwhelming long stress runs, and bumps the BatchGen Python package to 1.0.9.post1. See #139.
- Worker async load confirmation (
batchgen/batchgen_worker.py) — per-sequence[LOAD_CONFIRM] ... finalized uuid=...downgraded from info to debug. - GLM-5 MoE bucket fallback (
batchgen/models/glm/glm5/model.py) — oversize bucket fallback and GLM-5 hot-path banners downgraded from warning to debug. - DSA fused indexer JIT helper (
batchgen_kernels/attention/dsa/fused_indexer_kv_proj_cuda.py) — build-timeprints replaced with debug logging;load_inline(verbose=...)is now opt-in viaBATCHGEN_KERNEL_LOAD_TRACE=1. - 3D MoE dispatch wrapper (
batchgen/moe/dispatch_scatter_3d.py) — successful precompiled-kernel load banner downgraded from info to debug. - Core engine loader (
batchgen/models/engine_loader.py) — JIT fallback message routed through debug logging instead of stdout. - DSA fused indexer score (
batchgen_kernels/attention/dsa/fused_indexer_score.py) —softmax_scaleconstexpr + ReLU clamp;compute_head_gatesreverted to BF16 matmul + FP32 promotion to match the PyTorch fallback path.
Wheels
batchgen-1.0.9.post1— rebuilt on H20 node 0 from tagv1.0.9.post1.batchgen_kernels-0.3.1.post1— rebuilt on H20 node 0 from tagv1.0.9.post1(sm_90a only).flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post1"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post1-py3-none-any.whl"