Skip to content

[Perf]optimize the flash attention fp8 performance for gfx950 - #950

Merged
coderfeli merged 4 commits into
ROCm:mainfrom
binding7012:binding_fa_fp8_gfx950_opt2
Aug 3, 2026
Merged

[Perf]optimize the flash attention fp8 performance for gfx950#950
coderfeli merged 4 commits into
ROCm:mainfrom
binding7012:binding_fa_fp8_gfx950_opt2

Conversation

@binding7012

Copy link
Copy Markdown
Contributor

Motivation

optimize the flash attention fp8 performance for gfx950 to close the gap between aiter asm

Test Result

batch=4, H=32, D=128, FP8, MI355X

image

@coderfeli
coderfeli merged commit ca99de5 into ROCm:main Aug 3, 2026
11 checks passed
xudoyuan added a commit that referenced this pull request Aug 3, 2026
Resolve flash_attn_utils.py conflict: keep main's #950 vectorized
_score_pair_sum / _scale_sub_score_pair, expressed in this branch's
DSL style (operators + fx.fma instead of fmath.fma / _fadd/_fsub/_fmul).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants