Skip to content

Releases: batchgen-project/batchgen

BatchGen v1.0.11

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 16 Sep 23:49
19b18d3

Prebuilt Hopper (sm90a) wheel set for the reference env (py3.11, torch 2.9.0+cu128, CUDA 12.8): FA3, FlashMLA, DeepGEMM,
batchgen_kernels, batchgen. Enables install_deps.sh wheel fast-path.

BatchGen v1.0.10.post4

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 18 May 06:55

BatchGen v1.0.10.post4

What's New

  • Adds GLM whole-model CUDA graph support for the GLM-5 FP8 release path. Users can enable the production graph path with --enable-cuda-graph and configure bucket coverage with --cuda-graph-max-bucket-size and --cuda-graph-num-buckets; GLM-specific graph-selection environment variables are no longer required for the default path. See #155.
  • Captures configured GLM decode buckets during decode configuration instead of recapturing in the decode loop. Over-bucket or stale graph state uses eager decode fallback rather than hot-path graph capture. See #155.
  • Expands and hardens the GLM graph implementation across DSA, MoE, decoder-layer, and whole-model graph boundaries, including safer empty-rank handling, replay comparison coverage, graph-buffer lifecycle handling, and synchronized bucket capture. See #155 and #158.
  • Keeps GLM graph-path diagnostic logging quiet by default so normal decode runs avoid the previous verbose graph-path log stream. See #155.
  • Updates the GLM graph release compatibility floor to batchgen_kernels 0.3.3 and aligns the packaged kernel source version. See #156 and #157.
  • Reverts the MiniMax MoE decode buffer growth change from this release line. See #151 and #152.

Compatibility and Installation

  • Python: unchanged from the previous BatchGen post release.
  • CUDA/PyTorch: unchanged from the previous BatchGen post release.
  • Kernel package: GLM graph support in this release requires the matching batchgen_kernels 0.3.3 wheel.
  • Wheels: release assets attached to GitHub release v1.0.10.post4.
  • Install: use the wheel assets from the GitHub release page.

Notes

  • Docker image publication is intentionally outside the scope of this formal wheel release.

BatchGen v1.0.10.post2

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 07 May 11:54

BatchGen v1.0.10.post2

What's New

  • Adds the opt-in GLM-5-FP8 full-DSA CUDA graph path, which captures the full decode attention segment through projections, DSA selection/FlashMLA, KV writes, absorb, and final attention projection while preserving the existing public attention input/output API.
  • Reduces full-DSA CUDA graph retained HBM by sharing large scratch buffers across GLM-5 attention layers and releasing static buffers when graph buckets are dropped.
  • Keeps the previously released segmented DSA/MoE CUDA graph path available; full-DSA graph is explicitly gated so deployments can choose the safer segmented path or the new full-DSA path.

Compatibility and Installation

  • Python: 3.11+.
  • CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
  • Wheels: this release publishes a new batchgen wheel and reuses the unchanged Hopper dependency wheels from v1.0.10.post1.
  • Install: use the wheel assets from the GitHub release page for v1.0.10.post2.

Notes

For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8 with the full-DSA CUDA graph path, the recommended server-side flags are:

BATCHGEN_GLM5_DSA_FULL_CUDA_GRAPH=1
BATCHGEN_GLM5_WHOLE_MODEL_CUDA_GRAPH=0
--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93

These flags select the full-DSA attention CUDA graph path plus the public CUDA graph serving interface. Whole-model CUDA graph remains separate from this release path.

BatchGen v1.0.10.post1

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 07 May 03:13
7f41304

BatchGen v1.0.10.post1

What's New

  • Makes GLM-5-FP8 and GLM-5.1-FP8 segmented CUDA graph serving configurable through server CLI flags, so --enable-cuda-graph enables the DSA and MoE graph paths without requiring graph-specific environment variables. See #149.
  • Tightens the GLM-5 graph path contract so DSA and MoE graph requirements are propagated end-to-end from server startup through worker setup and model wrappers. See #149.
  • Adds fixes for decode metadata stability, DSA GPU KV loading, DSA graph page-table refresh, FlashMLA replay metadata preparation, and captured-sequence-length handling in the segmented graph path. See #149.

Compatibility and Installation

  • Python: 3.11+.
  • CUDA/PyTorch: CUDA 12.8+ with PyTorch 2.9.0+cu128.
  • Wheels: this release publishes a new batchgen wheel and reuses the compatible Hopper dependency wheels from v1.0.10.
  • Install: use the wheel assets from the GitHub release page for v1.0.10.post1.

Notes

For two-node NVIDIA H20 deployments serving GLM-5-FP8 or GLM-5.1-FP8, the recommended server-side CUDA graph flags are:

--enable-cuda-graph
--cuda-graph-max-bucket-size 64
--cuda-graph-num-buckets 7
--gpu-memory-frac 0.93

These flags select the segmented DSA and MoE graph path through the public server interface. Whole-model CUDA graph remains separate from this release path.

BatchGen v1.0.10

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 06 May 05:19
8f592e9

BatchGen v1.0.10

Highlights

  • Adds GLM-5-FP8 segmented CUDA graph decode support for the DSA and MoE paths.
  • Adds graph-safe DSA selected-KV handling with FlashMLA metadata passed as replay inputs.
  • Adds fixed-capacity CUDA graph page-table storage so graph capture does not mutate active decode page tables.
  • Adds GLM-5 MoE CUDA graph segment support and stricter graph state handling after model memory cleanup.
  • Includes focused CUDA graph tests for DSA projection/selection, FlashMLA metadata, page-table stability, DSA graph replay, and MoE graph replay.

Packaging

  • batchgen is released as 1.0.10.
  • batchgen_kernels is released as 0.3.2+sm90a.
  • FlashAttention, FlashMLA, and DeepGEMM wheels are reused from the previous compatible release because their pinned versions did not change.
  • This is a wheel release; no Docker image is published for this release.

Runtime notes

For the segmented GLM-5-FP8 graph path, enable segmented graph mode and the GLM-5 DSA/MoE graph flags. Whole-model CUDA graph remains separate from this release path.

Assets

The release includes BatchGen, BatchGen kernels, reused dependency wheels, and SHA256SUMS for verification.

v1.0.9.post5

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 29 Apr 06:27

BatchGen v1.0.9.post5

  • Packages GLM-5 DSA Hadamard kernels into batchgen_kernels as AOT extensions to remove the production runtime-JIT failure surface.
  • Bumps batchgen_kernels to 0.3.1.post3 and requires it from BatchGen.
  • Builds the internal H20 batchgen_kernels wheel with explicit sm90a arch selection.
  • Reuses unchanged dependency wheels from v1.0.9.post4.

Validated on two-node H20 GLM-5-FP8 after clearing test conda envs and Torch extension caches: installed post5 wheels in the persistent batchgen env, launched the server, submitted a batch, and confirmed DSA/Hadamard imports came from packaged site-packages .so files with no DSA/Hadamard Torch extension cache entries.

BatchGen v1.0.9.post4

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 28 Apr 14:43

Hotfix release for GLM-5 DSA startup on H20.

Changes:

  • Fixes duplicate Torch JIT extension ownership for GLM-5 DSA Hadamard/RoPE kernels.
  • Makes batchgen_kernels.attention.dsa.indexer the canonical loader with package-scoped batchgen_dsa_* extension names.
  • Keeps batchgen.other_kernels.hadamard_transform as a compatibility shim.
  • Requires batchgen_kernels 0.3.1.post2 from this release.

BatchGen v1.0.9.post3

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 28 Apr 08:36
e12eca3

Highlights

GLM-5-FP8 MoE perf release on top of v1.0.9.post2 (#142).

  • Wire act_quant_3d v2 into the WGMMA pipeline. The post1 release notes claimed an act_quant_3d rewrite from (E,) + serial token loop → (E, mtp), one CTA per (expert, token), but the rewrite landed only in tests/kernels/test_act_quant_3d_v2.py and the standalone batchgen_kernels C extension — production (fp8_wgmma_pipeline.py's inline JIT module) still ran v1, leaving GLM-5-FP8 on a 16-CTA / 132-SM (~12% occupancy) path. post3 swaps the inline kernel for the v2 layout (grid=(E, mtp)) and updates the launcher accordingly.
  • CUDA-graph MoE capture is on by default for GLM-5. BATCHGEN_GLM5_CUDA_GRAPH default flipped from "0" → "1". BATCHGEN_GLM5_CUDA_GRAPH=0 is the new explicit opt-out (eager MoE).
  • Default BATCHGEN_GLM5_MTP_BUCKETS = (128, 256, 512, 768) — the bucket set validated end-to-end this evening on H20 with 2K MMLU + 2K LongBench through the full 128K-decode regime. Empty string opts out to legacy singleton.

Validation: install + 2-node server launch + 2K+2K batch hot-patched against the post2 wheels reached Application startup complete, ran the prefill + decode regime, and submitted both batches without element-size-mismatch / SIGSEGV. post3 is the same source landed in the wheel.

Wheels

  • batchgen-1.0.9.post3 — rebuilt on H20 node 0 from tag v1.0.9.post3.
  • batchgen_kernels-0.3.1.post1 — reused unchanged from v1.0.9.post1.
  • flash_attn_3, flash_mla, deep_gemm — reused unchanged from v1.0.9.

Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).

Install:

RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post3"
pip install \
  "${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
  "${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen-1.0.9.post3-py3-none-any.whl"

Behavior change for operators

If you previously launched GLM-5 servers without BATCHGEN_GLM5_CUDA_GRAPH=1, the post3 server will now capture MoE graphs by default and pre-allocate per-bucket 3D buffers for (128, 256, 512, 768). To restore the old eager-MoE behavior, set BATCHGEN_GLM5_CUDA_GRAPH=0 (or set BATCHGEN_GLM5_MTP_BUCKETS= to an empty string for the legacy singleton path).

BatchGen v1.0.9.post2

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 28 Apr 06:49
036b78b

Highlights

Hotfix release for v1.0.9.post1. Fixes a packaging regression that broke the wheel-only install path on a fresh conda env. See #141.

  • batchgen/other_kernels/hadamard_transform/csrc/ now ships in the wheel. v1.0.9.post1's MANIFEST.in and setup.py only globbed batchgen/{core,external} source trees; the GLM-5 indexer's JIT-compiled hadamard / fused_rope_hadamard CUDA kernels need their 7 csrc files at runtime (hadamard_binding.cpp, fast_hadamard_transform_cuda.cu, fused_rope_hadamard*.{cpp,cu}, plus 3 headers). On a wheel-only install the JIT raises FileNotFoundError, the indexer silently falls back to non-fused RoPE+Hadamard, the fallback returns the input dtype after k_norm (which yields FP32 in torch 2.9), and the FP32 indexer K tensor trips host_paged_kv_worker_view.h "K tensor element size does not match configuration" (aux KV is BF16) on the first prefill batch. Worker SIGSEGV with no actionable log line.
  • Warn loudly on kernel-import failure. WP2/WP4/WP5 indexer kernel imports + the hadamard-csrc imports now log at WARNING (was DEBUG / silent). The next time a packaging gap surfaces, it shows up at server-launch time instead of as an opaque later SIGSEGV.

Wheels

  • batchgen-1.0.9.post2 — rebuilt on H20 node 0 from tag v1.0.9.post2.
  • batchgen_kernels-0.3.1.post1 — reused unchanged from v1.0.9.post1.
  • flash_attn_3, flash_mla, deep_gemm — reused unchanged from v1.0.9.

Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).

Install:

RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post2"
pip install \
  "${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
  "${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen-1.0.9.post2-py3-none-any.whl"

BatchGen v1.0.9.post1

Choose a tag to compare

@Andrewxu313 Andrewxu313 released this 27 Apr 16:48
ee1b556

Highlights

Patch release for GLM-5 / GLM-5.1 FP8 on top of v1.0.9. Quiets noisy decode/startup logs that were overwhelming long stress runs, and bumps the BatchGen Python package to 1.0.9.post1. See #139.

  • Worker async load confirmation (batchgen/batchgen_worker.py) — per-sequence [LOAD_CONFIRM] ... finalized uuid=... downgraded from info to debug.
  • GLM-5 MoE bucket fallback (batchgen/models/glm/glm5/model.py) — oversize bucket fallback and GLM-5 hot-path banners downgraded from warning to debug.
  • DSA fused indexer JIT helper (batchgen_kernels/attention/dsa/fused_indexer_kv_proj_cuda.py) — build-time prints replaced with debug logging; load_inline(verbose=...) is now opt-in via BATCHGEN_KERNEL_LOAD_TRACE=1.
  • 3D MoE dispatch wrapper (batchgen/moe/dispatch_scatter_3d.py) — successful precompiled-kernel load banner downgraded from info to debug.
  • Core engine loader (batchgen/models/engine_loader.py) — JIT fallback message routed through debug logging instead of stdout.
  • DSA fused indexer score (batchgen_kernels/attention/dsa/fused_indexer_score.py) — softmax_scale constexpr + ReLU clamp; compute_head_gates reverted to BF16 matmul + FP32 promotion to match the PyTorch fallback path.

Wheels

  • batchgen-1.0.9.post1 — rebuilt on H20 node 0 from tag v1.0.9.post1.
  • batchgen_kernels-0.3.1.post1 — rebuilt on H20 node 0 from tag v1.0.9.post1 (sm_90a only).
  • flash_attn_3, flash_mla, deep_gemm — reused unchanged from v1.0.9.

Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).

Install:

RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post1"
pip install \
  "${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
  "${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
  "${RELEASE_URL}/batchgen-1.0.9.post1-py3-none-any.whl"