BatchGen v1.0.9.post1
·
248 commits
to main
since this release
Highlights
Patch release for GLM-5 / GLM-5.1 FP8 on top of v1.0.9. Quiets noisy decode/startup logs that were overwhelming long stress runs, and bumps the BatchGen Python package to 1.0.9.post1. See #139.
- Worker async load confirmation (
batchgen/batchgen_worker.py) — per-sequence[LOAD_CONFIRM] ... finalized uuid=...downgraded from info to debug. - GLM-5 MoE bucket fallback (
batchgen/models/glm/glm5/model.py) — oversize bucket fallback and GLM-5 hot-path banners downgraded from warning to debug. - DSA fused indexer JIT helper (
batchgen_kernels/attention/dsa/fused_indexer_kv_proj_cuda.py) — build-timeprints replaced with debug logging;load_inline(verbose=...)is now opt-in viaBATCHGEN_KERNEL_LOAD_TRACE=1. - 3D MoE dispatch wrapper (
batchgen/moe/dispatch_scatter_3d.py) — successful precompiled-kernel load banner downgraded from info to debug. - Core engine loader (
batchgen/models/engine_loader.py) — JIT fallback message routed through debug logging instead of stdout. - DSA fused indexer score (
batchgen_kernels/attention/dsa/fused_indexer_score.py) —softmax_scaleconstexpr + ReLU clamp;compute_head_gatesreverted to BF16 matmul + FP32 promotion to match the PyTorch fallback path.
Wheels
batchgen-1.0.9.post1— rebuilt on H20 node 0 from tagv1.0.9.post1.batchgen_kernels-0.3.1.post1— rebuilt on H20 node 0 from tagv1.0.9.post1(sm_90a only).flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post1"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post1-py3-none-any.whl"