BatchGen v1.0.9.post2
·
251 commits
to main
since this release
Highlights
Hotfix release for v1.0.9.post1. Fixes a packaging regression that broke the wheel-only install path on a fresh conda env. See #141.
batchgen/other_kernels/hadamard_transform/csrc/now ships in the wheel. v1.0.9.post1'sMANIFEST.inandsetup.pyonly globbedbatchgen/{core,external}source trees; the GLM-5 indexer's JIT-compiled hadamard / fused_rope_hadamard CUDA kernels need their 7 csrc files at runtime (hadamard_binding.cpp,fast_hadamard_transform_cuda.cu,fused_rope_hadamard*.{cpp,cu}, plus 3 headers). On a wheel-only install the JIT raisesFileNotFoundError, the indexer silently falls back to non-fused RoPE+Hadamard, the fallback returns the input dtype afterk_norm(which yields FP32 in torch 2.9), and the FP32 indexer K tensor tripshost_paged_kv_worker_view.h"K tensor element size does not match configuration" (aux KV is BF16) on the first prefill batch. Worker SIGSEGV with no actionable log line.- Warn loudly on kernel-import failure. WP2/WP4/WP5 indexer kernel imports + the hadamard-csrc imports now log at WARNING (was DEBUG / silent). The next time a packaging gap surfaces, it shows up at server-launch time instead of as an opaque later SIGSEGV.
Wheels
batchgen-1.0.9.post2— rebuilt on H20 node 0 from tagv1.0.9.post2.batchgen_kernels-0.3.1.post1— reused unchanged from v1.0.9.post1.flash_attn_3,flash_mla,deep_gemm— reused unchanged from v1.0.9.
Build environment: CUDA 12.8, PyTorch 2.9.0+cu128, Python 3.11, NVIDIA H20 (SM 9.0).
Install:
RELEASE_URL="https://github.com/EfficientMoE/BatchGen/releases/download/v1.0.9.post2"
pip install \
"${RELEASE_URL}/flash_attn_3-3.0.0b1-cp39-abi3-linux_x86_64.whl" \
"${RELEASE_URL}/flash_mla-1.0.0+1408756-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/deep_gemm-2.1.1+c9f8b34-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen_kernels-0.3.1.post1-cp311-cp311-linux_x86_64.whl" \
"${RELEASE_URL}/batchgen-1.0.9.post2-py3-none-any.whl"