Skip to content

AITER v0.1.22

Choose a tag to compare

@github-actions github-actions released this 14 Sep 08:24
· 1 commit to release/v0.1.22 since this release
a7c3b96

AITER v0.1.22 - bi-weekly release

Scheduled release from release/v0.1.22.

Diff base: v0.1.21

Wheels

Prebuilt manylinux_2_28 wheels, GPU_ARCHS=gfx942;gfx950:

ROCm Python Wheel
ROCm 7.0 cp310 amd_aiter-0.1.22+rocm7.0.manylinux.2.28-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
ROCm 7.0 cp312 amd_aiter-0.1.22+rocm7.0.manylinux.2.28-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
ROCm 7.1 cp310 amd_aiter-0.1.22+rocm7.1.manylinux.2.28-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
ROCm 7.1 cp312 amd_aiter-0.1.22+rocm7.1.manylinux.2.28-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
ROCm 7.2 cp310 amd_aiter-0.1.22+rocm7.2.manylinux.2.28-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
ROCm 7.2 cp312 amd_aiter-0.1.22+rocm7.2.manylinux.2.28-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
What's Changed

Changes

  • [Triton/Gluon] Add config-aware repr to the rope and normalization kernels by @Boss2002n in #5099
  • [Triton/Gluon] Add config-aware repr to the quant kernels by @Boss2002n in #5100
  • [CI] CI: use app token for release refs by @gyohuangxin in #5198
  • [CI] CI: respect docker login input in release builds by @gyohuangxin in #5201
  • [Triton/Gluon] [HIP] [CK] Remove obsolete availability helpers by @coderfeli in #5116
  • [FlyDSL] Add FlyDSL Radix-Select TopK Path to the Existing Per-Row Decode Interface by @lirui927 in #5011
  • [CI] CI: set release publish repository context by @gyohuangxin in #5204
  • [HIP] Add fused SiTUv2 activation + per-token FP8 quant kernel by @XiaobingSuper in #5081
  • [Docs] docs: update README by @shengnxu in #5073
  • [CK] [FlyDSL] Retune Kimi-K3 a16w4 MoE tile geometry by @amd-wsung102 in #5118
  • [Config] tune a8w8 gemm with 64-step m for k3 by @gbyu-amd in #5197
  • [CI] CI: add extended test workflow by @gyohuangxin in #4458
  • [CI] Run extended test dispatch on internal runner by @gyohuangxin in #5205
  • Fix stale TopK availability checks by @vorapolsiloai in #5210
  • [CI] Extended tests client_payload change by @leo-automation in #5209
  • [Triton/Gluon] Move the sage-attention launch params into the config tree by @Boss2002n in #5106
  • [Triton/Gluon] moe_gemm_a4w4 num_warps 8 -> 4 for block_m != 16 tile by @nidal567 in #5189
  • [Docs] Update docs for the single nested config layout by @Boss2002n in #5085
  • [skills] Add kernel PR validation and structural D9 scanning (supersedes #4870) by @zhiding512 in #5142
  • [AMD][DSV4] Fix(fmoe): fuse stage-1 fp8 quant on the heuristic FlyDSL fallback by @karverma-amd in #4994
  • [Bugfix] Make _fold_seqlen_indptr cudagraph-safe (avoid scalar H2D copy) by @micah-wil in #5202
  • [FlyDSL] Relax GDR decode test mismatch tolerance by @xytpai in #5215
  • [HIP] Extend fused QK norm for MiniMax-M3 by @weitliao in #5143
  • Revert "[AMD][DSV4] Fix(fmoe): fuse stage-1 fp8 quant on the heuristic FlyDSL fallback" by @valarLip in #5227
  • [Triton/Gluon] [CI] Add fused KDA decode kernel (conv1d + recurrence + gated RMSNorm) by @mengfei-jiang in #4712
  • [HIP] fix(mla/metadata): scale the auto KV-split count with the machine width by @whx-sjtu in #5212
  • [Triton/Gluon] Add repr to DiT fused kernels + Cover round_intermediate in tests by @brunomazzottiamd in #5154
  • [HIP] [FlyDSL] Refactor gfx950 A16W16 GEMM with centralized policy selection and tuning by @xytpai in #5145
  • [Triton/Gluon] [GFX9] [GFX12] EP MOE changes by @k50112113 in #4500
  • [FlyDSL] [CI] FlyDSL FMHA forward-prefill A16W16 kernel for gfx1250 (fmha_fwd_prefill_m32x8) by @ruanjm in #5168
  • [HIP] [Bugfix] Guard against negative expert ids in MoE sorting P0 by @xudonlyu in #4839
  • [FlyDSL] perf(mqa-logits): build the FP4 prefill schedule in one kernel, not 25 torch ops by @valarLip in #5276
  • [Config] configs: GLM-5.2 a8w8-bpreshuffle rows for the shapes a live engine dispatches by @ThomasNing in #5281
  • [FlyDSL] Rewrite gfx950 HGEMM test to op_test standard by @xytpai in #5287
  • [Config] tune(dsv4): gfx950 FP8 blockscale bpreshuffle for wq_b / wqkv_a by @lixiufei-leo in #5283
  • [FlyDSL] [DSv4] Bound the FP4 MQA-logits store with a window-sized V# instead of a compare by @valarLip in #5285
  • [Config] [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs by @karverma-amd in #5279
  • [Triton/Gluon] Add FP8 block-wise quantization kernels by @WuLei-AMD in #5177
  • [Triton/Gluon] Add vocab-parallel cross-entropy kernel by @WuLei-AMD in #5165
  • [FlyDSL] Remove the unused gfx1250 d192 FMHA sibling kernel by @coderfeli in #5306
  • [HIP] [CK] [MoE] Reject non-int32 index buffers in the topk kernels by @i-kosarev in #5255
  • [FlyDSL] Migrate mixed MoE 2-stage LDS to fly shared storage by @xudoyuan in #5317
  • review-pr: gates that can be checked, and 200 PRs of evidence about them by @zufayu in #5289
  • [CI] fix auditwheel excludes: versioned ROCm SONAMEs were grafting ~890MB into the wheel by @amd-ruitang3 in #5318
  • [Triton/Gluon] Add MXFP8 convert and fast-transpose kernels by @WuLei-AMD in #5203
  • [HIP] [JIT] [Build] Stop stamping torch-free modules with torch's pybind11 ABI identity by @amd-ruitang3 in #5312
  • fix(mega_moe): stop the reference clamping SwiGLU at swiglu_limit=0 by @JohnQinAMD in #5271
  • [Triton/Gluon] chunk_kimi_delta_attn: accept a non-fp32 KDA state by @XiaobingSuper in #5249
  • [Triton/Gluon] Fix ragged-K mask in batched A16WFP4 GEMM by @mjkvaak-amd in #4181
  • [Triton/Gluon] [FlyDSL] feat(flydsl): Add HSTU Forward kernel by @damien-lejeune in #4441
  • [FlyDSL] [JIT] refactor: move tuning config file for all2all by @JiaoliangYu in #5129
  • [Triton/Gluon] [GFX950] Add MHA Gluon Kernel by @lucas-santos-amd in #4147
  • [FlyDSL] Tune the GLM-5.2 decode shapes for gfx950 by @kyle-256 in #5328
  • [FlyDSL] fix(flydsl): zero GDR decode graph padding output by @junna2016 in #5324
  • [Triton/Gluon] [Conv2D] Add Conv2D configuration files for gfx1101 (RDNA3) and gfx1150 (RDNA3.5) by @Ragua1 in #5246
  • [FlyDSL] feat(mega_moe): optimize fused stage1 and AOT bundles by @GwilliamHu in #5001
  • [FlyDSL] Revert " feat(mega_moe): optimize fused stage1 and AOT bundle… by @coderfeli in #5345
  • [FlyDSL] fix(flydsl): detach tensors before DLPack conversion by @xytpai in #5327
  • [FlyDSL] [CI] [JIT] comm fused moe by @yifehuan in #4985
  • [Triton/Gluon] split K correct OOB afp4wfp4 by @afriedri in #5261
  • [HIP] [OPUS] fix: single-source opus's half-precision dtype spellings by @zufayu in #5329
  • [FlyDSL] Megamoe restore 5001 aot fix by @GwilliamHu in #5346
  • [HIP] Use designated initializers for FMHA forward arguments by @rocking5566 in #5359
  • [ASM] [HIP] [OPUS] Add narrow-head (H<=32) gfx1250 MLA sparse-prefill kernels for DSv4 TP by @kaiyang-1 in #5228
  • [Triton/Gluon] [CI] [Bugfix] Fix stale MHA config-utils import by @vorapolsiloai in #5352
  • [OPUS] gfx1250: keep split-K off shapes its reduce cannot address by @demonsan in #5162
  • [Config] Tune Kimi-K3 a8w8 bpreshuffle long-prefill shapes by @rebklee in #5360
  • [Triton/Gluon] [HIP] Fix large tensor addressing-quant & inverse_rope by @yzhou103 in #5314
  • [Triton/Gluon] [ASM] [HIP] MHA v4: fixes, refactor, new kernel, perf tweaks by @jcaraban in #5335
  • [Triton/Gluon] add mla decode kernel name prefix by @Boss2002n in #5214
  • [FlyDSL] opt-in a16wi4 gemm2 CShuffle epilog via CSV kernelName2 by @msaffari-amd in #5240
  • [ASM] Use 64-bit offsets for weigths by @JohnNikolay84 in #5357
  • [HIP] [CK] [FlyDSL] Low-M MXFP4 fused-MoE: padded-row launch bound, SiTUv2 stage1 fusion, BM16 inline-sort by @XiaobingSuper in #5300
  • [Config] Retune gpt-oss N=2880,K=4096 off hipBLASLt so CUDAGraph capture succeeds by @valarLip in #5371
  • [FlyDSL] replace ptr_rsrc and buffer ops low level api use by @coderfeli in #5303
  • [Docs] Add PR-scope, comment and UT-hygiene rules to Copilot review instructions by @Boss2002n in #5369
  • [Config] [Tune] Add GLM-5.3-Flash a8w8 blockscale GEMM configs for gfx942 by @jin-amd in #5343
  • [Triton/Gluon] [MI350] Optimize FP8 MQA logits kernel by @cagrikymk in #5216
  • [Triton/Gluon] [PERF] Optimize Triton unified attention prefill and decode by @vorapolsiloai in #4761
  • [Triton/Gluon] Move fp4 GEMM gluon to _gluon_kernels by @vgokhale in #4997
  • [Docs] Satya/torch free triton docs by @Boss2002n in #5383
  • [Triton/Gluon] Use absolute imports under ops/triton and triton_tests by @Boss2002n in #5382
  • [CI] fix: fix the conflict of cpu isolation and sglang cpu affinity by @siweiiiiii in #5365
  • [HIP] Fix (custom_all_reduce): correct output RankData slot prediction during graph capture by @afriedri in #4974
  • [Triton/Gluon] [Config] [GFX1250] DSR1 Triton Untuned GEMM kernels Tuning by @leonling-ll in #5349
  • [OPUS] [JIT] [Bugfix] Make blob codegen cache publication transactional by @Ruye-aa in #4800
  • [FlyDSL] add gather-gemm for a8w8 by @solinzby1 in #5207
  • [Config] [FMoE] Tuned bf16 MoE for K2 horizon 375B by @a-sidorova in #5275
  • [ASM] [gfx1250] mla v4 prefill: rebuild the sparse_pfl kernel by @junxiaguo in #5367
  • [CK] [JIT] feat(cktile a8w8-bpreshuffle): add rowcol_wp_v2 as a selectable kernel type by @ThomasNing in #5280
  • [Triton/Gluon] Add large-M/small-N RMSNorm backward specialization by @WuLei-AMD in #5305
  • [FlyDSL] add flydsl flash attention fp8 support(dense, asymmetric head dimensions, variable sequence lengths (varlen), and KV splitting) by @binding7012 in #5326
  • [HIP] [CK] perf(tuner): build the result frame once instead of concatenating per row by @ThomasNing in #5262
  • [FlyDSL] Enable SwiGLU and tune MiniMax-M3 A16W4 MoE on gfx950 by @yifehuan in #5395
  • [CI] CI: require label for extended tests by @gyohuangxin in #5411
  • [FlyDSL] Fix HSTU reference import fallback by @gyohuangxin in #5394
  • [Triton/Gluon] Supress aiter logger under pytest by @Boss2002n in #5390
  • [Triton/Gluon] fused sigmoid implementation for vLLM by @omuhamma in #5188
  • [Triton/Gluon] [Build] Drop deprecated kwargs by @Boss2002n in #5381
  • [ASM] [gfx1250] MLA decode: rebuild 7 kernels with missing SCHED_MODE_… by @junxiaguo in #5402
  • [HIP] fused_qk_norm_rope_group_quant opt in gfx1250 by @yzhou103 in #5226
  • [FlyDSL] [gfx1250] Speed up the grouped-MoE route, quant, and gather-reduce epilogues by @lalala-sh in #5313
  • [CI] CI: auto-update split test FILE_TIMES by @aiter-gh-app[bot] in #5131
  • [HIP] [FlyDSL] [Bugfix][MLA] Correct final_lse in PS MLA prefill kernel for chunked prefill by @simondanielsson in #3606
  • [FlyDSL] feat(comm-fused-moe): log activated runner configuration by @yifehuan in #5433
  • [Triton/Gluon] Restore shared MHA backend parameterization by @vorapolsiloai in #5374
  • [Triton/Gluon] Add MoE GEMM variants and weight-gradient kernel by @WuLei-AMD in #5310
  • [FlyDSL] [gfx1201] Optimize causal Flash Attention and make V prefetch OOB-safe by @vlluvia in #5441
  • [HIP] add env var to control perf co by @JaxChen29 in #5437
  • [FlyDLSL] Widen gfx950 stable decode TopK gates by @lirui927 in #5442
  • [Triton/Gluon] [FlyDSL] [CI] Add GDN MTP kernels: causal conv1d update and gated delta rule by @yiijin in #4907
  • [FlyDSL] fix trunci AttributeError in qk_norm_rope TDM path by @jli-melchior in #5438
  • [Config] [Perf] Tune DSV4 EP48 top-k 6 prefill shapes by @yhl-amd in #5459
  • [Config] [tune] DSv4 a8w8 blockscale: add gfx950 configs for three MI355X shapes by @jiacao-amd in #4664
  • [Triton/Gluon] [GFX12] A4W4 MOE tune/optimize by @k50112113 in #5416
  • [Triton/Gluon] Moe routing optimizations by @lburzawa in #5038
  • [Triton/Gluon] make W4A16 backend warning print it only once instead of on every call by @functionstackx in #5471
  • [Triton/Gluon] Do not autotune for the FA port in unit tests by @Boss2002n in #5419
  • [Fix] [Perf] Tune DSV4 EP48 stage2 for RCCL routing by @yhl-amd in #5467
  • [Triton/Gluon] One shared helper for opt-in @triton.autotune config lists by @Boss2002n in #5414
  • [Triton/Gluon] [OPUS] [FlyDSL] Gfx1250/microbench by @JiaoliangYu in #5391
  • [FlyDSL] gather_kv_b_proj: lift the 4 GiB cache limit, and let callers ask before committing by @valarLip in #5475
  • [HIP] fea: add LL, LL128 proto into gfx9 by @TennyWang1223 in #5272
  • [HIP] [CK] [FlyDSL] [Kernel] Extend MXFP4 GEMM1 replacement to A4W4 by @fsx950223 in #4526
  • [FlyDSL] some moe optimization by @yadaish in #5448
  • [CK] Add tuned CK a8w8 blockscale GEMM configs for Gemma-4-31B FP8-block by @mustafayildirim in #5062
  • [Triton/Gluon] attn_residual prefill prefix seperation by @yanxuer-999 in #5067
  • [Config] Tune GLM5.2 per-token FP8 MoE shapes by @qichu-yun in #5144

New Contributors

Full Changelog: v0.1.21...v0.1.22