AITER v0.1.22
·
1 commit
to release/v0.1.22
since this release
AITER v0.1.22 - bi-weekly release
Scheduled release from release/v0.1.22.
Diff base: v0.1.21
Wheels
Prebuilt manylinux_2_28 wheels, GPU_ARCHS=gfx942;gfx950:
What's Changed
Changes
- [Triton/Gluon] Add config-aware repr to the rope and normalization kernels by @Boss2002n in #5099
- [Triton/Gluon] Add config-aware repr to the quant kernels by @Boss2002n in #5100
- [CI] CI: use app token for release refs by @gyohuangxin in #5198
- [CI] CI: respect docker login input in release builds by @gyohuangxin in #5201
- [Triton/Gluon] [HIP] [CK] Remove obsolete availability helpers by @coderfeli in #5116
- [FlyDSL] Add FlyDSL Radix-Select TopK Path to the Existing Per-Row Decode Interface by @lirui927 in #5011
- [CI] CI: set release publish repository context by @gyohuangxin in #5204
- [HIP] Add fused SiTUv2 activation + per-token FP8 quant kernel by @XiaobingSuper in #5081
- [Docs] docs: update README by @shengnxu in #5073
- [CK] [FlyDSL] Retune Kimi-K3 a16w4 MoE tile geometry by @amd-wsung102 in #5118
- [Config] tune a8w8 gemm with 64-step m for k3 by @gbyu-amd in #5197
- [CI] CI: add extended test workflow by @gyohuangxin in #4458
- [CI] Run extended test dispatch on internal runner by @gyohuangxin in #5205
- Fix stale TopK availability checks by @vorapolsiloai in #5210
- [CI] Extended tests client_payload change by @leo-automation in #5209
- [Triton/Gluon] Move the sage-attention launch params into the config tree by @Boss2002n in #5106
- [Triton/Gluon] moe_gemm_a4w4 num_warps 8 -> 4 for block_m != 16 tile by @nidal567 in #5189
- [Docs] Update docs for the single nested config layout by @Boss2002n in #5085
- [skills] Add kernel PR validation and structural D9 scanning (supersedes #4870) by @zhiding512 in #5142
- [AMD][DSV4] Fix(fmoe): fuse stage-1 fp8 quant on the heuristic FlyDSL fallback by @karverma-amd in #4994
- [Bugfix] Make
_fold_seqlen_indptrcudagraph-safe (avoid scalar H2D copy) by @micah-wil in #5202 - [FlyDSL] Relax GDR decode test mismatch tolerance by @xytpai in #5215
- [HIP] Extend fused QK norm for MiniMax-M3 by @weitliao in #5143
- Revert "[AMD][DSV4] Fix(fmoe): fuse stage-1 fp8 quant on the heuristic FlyDSL fallback" by @valarLip in #5227
- [Triton/Gluon] [CI] Add fused KDA decode kernel (conv1d + recurrence + gated RMSNorm) by @mengfei-jiang in #4712
- [HIP] fix(mla/metadata): scale the auto KV-split count with the machine width by @whx-sjtu in #5212
- [Triton/Gluon] Add
reprto DiT fused kernels + Coverround_intermediatein tests by @brunomazzottiamd in #5154 - [HIP] [FlyDSL] Refactor gfx950 A16W16 GEMM with centralized policy selection and tuning by @xytpai in #5145
- [Triton/Gluon] [GFX9] [GFX12] EP MOE changes by @k50112113 in #4500
- [FlyDSL] [CI] FlyDSL FMHA forward-prefill A16W16 kernel for gfx1250 (
fmha_fwd_prefill_m32x8) by @ruanjm in #5168 - [HIP] [Bugfix] Guard against negative expert ids in MoE sorting P0 by @xudonlyu in #4839
- [FlyDSL] perf(mqa-logits): build the FP4 prefill schedule in one kernel, not 25 torch ops by @valarLip in #5276
- [Config] configs: GLM-5.2 a8w8-bpreshuffle rows for the shapes a live engine dispatches by @ThomasNing in #5281
- [FlyDSL] Rewrite gfx950 HGEMM test to op_test standard by @xytpai in #5287
- [Config] tune(dsv4): gfx950 FP8 blockscale bpreshuffle for wq_b / wqkv_a by @lixiufei-leo in #5283
- [FlyDSL] [DSv4] Bound the FP4 MQA-logits store with a window-sized V# instead of a compare by @valarLip in #5285
- [Config] [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs by @karverma-amd in #5279
- [Triton/Gluon] Add FP8 block-wise quantization kernels by @WuLei-AMD in #5177
- [Triton/Gluon] Add vocab-parallel cross-entropy kernel by @WuLei-AMD in #5165
- [FlyDSL] Remove the unused gfx1250 d192 FMHA sibling kernel by @coderfeli in #5306
- [HIP] [CK] [MoE] Reject non-int32 index buffers in the topk kernels by @i-kosarev in #5255
- [FlyDSL] Migrate mixed MoE 2-stage LDS to fly shared storage by @xudoyuan in #5317
- review-pr: gates that can be checked, and 200 PRs of evidence about them by @zufayu in #5289
- [CI] fix auditwheel excludes: versioned ROCm SONAMEs were grafting ~890MB into the wheel by @amd-ruitang3 in #5318
- [Triton/Gluon] Add MXFP8 convert and fast-transpose kernels by @WuLei-AMD in #5203
- [HIP] [JIT] [Build] Stop stamping torch-free modules with torch's pybind11 ABI identity by @amd-ruitang3 in #5312
- fix(mega_moe): stop the reference clamping SwiGLU at swiglu_limit=0 by @JohnQinAMD in #5271
- [Triton/Gluon] chunk_kimi_delta_attn: accept a non-fp32 KDA state by @XiaobingSuper in #5249
- [Triton/Gluon] Fix ragged-K mask in batched A16WFP4 GEMM by @mjkvaak-amd in #4181
- [Triton/Gluon] [FlyDSL] feat(flydsl): Add HSTU Forward kernel by @damien-lejeune in #4441
- [FlyDSL] [JIT] refactor: move tuning config file for all2all by @JiaoliangYu in #5129
- [Triton/Gluon] [GFX950] Add MHA Gluon Kernel by @lucas-santos-amd in #4147
- [FlyDSL] Tune the GLM-5.2 decode shapes for gfx950 by @kyle-256 in #5328
- [FlyDSL] fix(flydsl): zero GDR decode graph padding output by @junna2016 in #5324
- [Triton/Gluon] [Conv2D] Add Conv2D configuration files for gfx1101 (RDNA3) and gfx1150 (RDNA3.5) by @Ragua1 in #5246
- [FlyDSL] feat(mega_moe): optimize fused stage1 and AOT bundles by @GwilliamHu in #5001
- [FlyDSL] Revert " feat(mega_moe): optimize fused stage1 and AOT bundle… by @coderfeli in #5345
- [FlyDSL] fix(flydsl): detach tensors before DLPack conversion by @xytpai in #5327
- [FlyDSL] [CI] [JIT] comm fused moe by @yifehuan in #4985
- [Triton/Gluon] split K correct OOB afp4wfp4 by @afriedri in #5261
- [HIP] [OPUS] fix: single-source opus's half-precision dtype spellings by @zufayu in #5329
- [FlyDSL] Megamoe restore 5001 aot fix by @GwilliamHu in #5346
- [HIP] Use designated initializers for FMHA forward arguments by @rocking5566 in #5359
- [ASM] [HIP] [OPUS] Add narrow-head (H<=32) gfx1250 MLA sparse-prefill kernels for DSv4 TP by @kaiyang-1 in #5228
- [Triton/Gluon] [CI] [Bugfix] Fix stale MHA config-utils import by @vorapolsiloai in #5352
- [OPUS] gfx1250: keep split-K off shapes its reduce cannot address by @demonsan in #5162
- [Config] Tune Kimi-K3 a8w8 bpreshuffle long-prefill shapes by @rebklee in #5360
- [Triton/Gluon] [HIP] Fix large tensor addressing-quant & inverse_rope by @yzhou103 in #5314
- [Triton/Gluon] [ASM] [HIP] MHA v4: fixes, refactor, new kernel, perf tweaks by @jcaraban in #5335
- [Triton/Gluon] add mla decode kernel name prefix by @Boss2002n in #5214
- [FlyDSL] opt-in a16wi4 gemm2 CShuffle epilog via CSV kernelName2 by @msaffari-amd in #5240
- [ASM] Use 64-bit offsets for weigths by @JohnNikolay84 in #5357
- [HIP] [CK] [FlyDSL] Low-M MXFP4 fused-MoE: padded-row launch bound, SiTUv2 stage1 fusion, BM16 inline-sort by @XiaobingSuper in #5300
- [Config] Retune gpt-oss N=2880,K=4096 off hipBLASLt so CUDAGraph capture succeeds by @valarLip in #5371
- [FlyDSL] replace ptr_rsrc and buffer ops low level api use by @coderfeli in #5303
- [Docs] Add PR-scope, comment and UT-hygiene rules to Copilot review instructions by @Boss2002n in #5369
- [Config] [Tune] Add GLM-5.3-Flash a8w8 blockscale GEMM configs for gfx942 by @jin-amd in #5343
- [Triton/Gluon] [MI350] Optimize FP8 MQA logits kernel by @cagrikymk in #5216
- [Triton/Gluon] [PERF] Optimize Triton unified attention prefill and decode by @vorapolsiloai in #4761
- [Triton/Gluon] Move fp4 GEMM gluon to _gluon_kernels by @vgokhale in #4997
- [Docs] Satya/torch free triton docs by @Boss2002n in #5383
- [Triton/Gluon] Use absolute imports under ops/triton and triton_tests by @Boss2002n in #5382
- [CI] fix: fix the conflict of cpu isolation and sglang cpu affinity by @siweiiiiii in #5365
- [HIP] Fix (custom_all_reduce): correct output RankData slot prediction during graph capture by @afriedri in #4974
- [Triton/Gluon] [Config] [GFX1250] DSR1 Triton Untuned GEMM kernels Tuning by @leonling-ll in #5349
- [OPUS] [JIT] [Bugfix] Make blob codegen cache publication transactional by @Ruye-aa in #4800
- [FlyDSL] add gather-gemm for a8w8 by @solinzby1 in #5207
- [Config] [FMoE] Tuned bf16 MoE for K2 horizon 375B by @a-sidorova in #5275
- [ASM] [gfx1250] mla v4 prefill: rebuild the sparse_pfl kernel by @junxiaguo in #5367
- [CK] [JIT] feat(cktile a8w8-bpreshuffle): add rowcol_wp_v2 as a selectable kernel type by @ThomasNing in #5280
- [Triton/Gluon] Add large-M/small-N RMSNorm backward specialization by @WuLei-AMD in #5305
- [FlyDSL] add flydsl flash attention fp8 support(dense, asymmetric head dimensions, variable sequence lengths (varlen), and KV splitting) by @binding7012 in #5326
- [HIP] [CK] perf(tuner): build the result frame once instead of concatenating per row by @ThomasNing in #5262
- [FlyDSL] Enable SwiGLU and tune MiniMax-M3 A16W4 MoE on gfx950 by @yifehuan in #5395
- [CI] CI: require label for extended tests by @gyohuangxin in #5411
- [FlyDSL] Fix HSTU reference import fallback by @gyohuangxin in #5394
- [Triton/Gluon] Supress aiter logger under pytest by @Boss2002n in #5390
- [Triton/Gluon] fused sigmoid implementation for vLLM by @omuhamma in #5188
- [Triton/Gluon] [Build] Drop deprecated kwargs by @Boss2002n in #5381
- [ASM] [gfx1250] MLA decode: rebuild 7 kernels with missing SCHED_MODE_… by @junxiaguo in #5402
- [HIP] fused_qk_norm_rope_group_quant opt in gfx1250 by @yzhou103 in #5226
- [FlyDSL] [gfx1250] Speed up the grouped-MoE route, quant, and gather-reduce epilogues by @lalala-sh in #5313
- [CI] CI: auto-update split test FILE_TIMES by @aiter-gh-app[bot] in #5131
- [HIP] [FlyDSL] [Bugfix][MLA] Correct final_lse in PS MLA prefill kernel for chunked prefill by @simondanielsson in #3606
- [FlyDSL] feat(comm-fused-moe): log activated runner configuration by @yifehuan in #5433
- [Triton/Gluon] Restore shared MHA backend parameterization by @vorapolsiloai in #5374
- [Triton/Gluon] Add MoE GEMM variants and weight-gradient kernel by @WuLei-AMD in #5310
- [FlyDSL] [gfx1201] Optimize causal Flash Attention and make V prefetch OOB-safe by @vlluvia in #5441
- [HIP] add env var to control perf co by @JaxChen29 in #5437
- [FlyDLSL] Widen gfx950 stable decode TopK gates by @lirui927 in #5442
- [Triton/Gluon] [FlyDSL] [CI] Add GDN MTP kernels: causal conv1d update and gated delta rule by @yiijin in #4907
- [FlyDSL] fix trunci AttributeError in qk_norm_rope TDM path by @jli-melchior in #5438
- [Config] [Perf] Tune DSV4 EP48 top-k 6 prefill shapes by @yhl-amd in #5459
- [Config] [tune] DSv4 a8w8 blockscale: add gfx950 configs for three MI355X shapes by @jiacao-amd in #4664
- [Triton/Gluon] [GFX12] A4W4 MOE tune/optimize by @k50112113 in #5416
- [Triton/Gluon] Moe routing optimizations by @lburzawa in #5038
- [Triton/Gluon] make W4A16 backend warning print it only once instead of on every call by @functionstackx in #5471
- [Triton/Gluon] Do not autotune for the FA port in unit tests by @Boss2002n in #5419
- [Fix] [Perf] Tune DSV4 EP48 stage2 for RCCL routing by @yhl-amd in #5467
- [Triton/Gluon] One shared helper for opt-in @triton.autotune config lists by @Boss2002n in #5414
- [Triton/Gluon] [OPUS] [FlyDSL] Gfx1250/microbench by @JiaoliangYu in #5391
- [FlyDSL] gather_kv_b_proj: lift the 4 GiB cache limit, and let callers ask before committing by @valarLip in #5475
- [HIP] fea: add LL, LL128 proto into gfx9 by @TennyWang1223 in #5272
- [HIP] [CK] [FlyDSL] [Kernel] Extend MXFP4 GEMM1 replacement to A4W4 by @fsx950223 in #4526
- [FlyDSL] some moe optimization by @yadaish in #5448
- [CK] Add tuned CK a8w8 blockscale GEMM configs for Gemma-4-31B FP8-block by @mustafayildirim in #5062
- [Triton/Gluon] attn_residual prefill prefix seperation by @yanxuer-999 in #5067
- [Config] Tune GLM5.2 per-token FP8 MoE shapes by @qichu-yun in #5144
New Contributors
- @amd-wsung102 made their first contribution in #5118
- @weitliao made their first contribution in #5143
- @whx-sjtu made their first contribution in #5212
- @lixiufei-leo made their first contribution in #5283
- @i-kosarev made their first contribution in #5255
- @damien-lejeune made their first contribution in #4441
- @Ragua1 made their first contribution in #5246
- @afriedri made their first contribution in #5261
- @rebklee made their first contribution in #5360
- @siweiiiiii made their first contribution in #5365
- @Ruye-aa made their first contribution in #4800
- @simondanielsson made their first contribution in #3606
- @vlluvia made their first contribution in #5441
- @functionstackx made their first contribution in #5471
- @mustafayildirim made their first contribution in #5062
Full Changelog: v0.1.21...v0.1.22