[inductor] Support atomic template reduction groups - #190822
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/190822
Note: Links to docs will display an error until the docs builds have been completed. ❌ 9 New Failures, 4 Cancelled JobsAs of commit 8c17d64 with merge base e56e791 ( NEW FAILURES - The following jobs have failed:
CANCELLED JOBS - The following jobs were cancelled. Please retry:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
AI-Review
Verified Issues
|
Allow template scheduling to claim a complete reduction chain as one fusion unit and retain the output identity needed by multi-output template kernels. This prevents partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. Test Plan: ```bash lintrunner -a ``` Authored with an AI assistant. ghstack-source-id: 1b35327 Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: 4893312 Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: ea479e6 Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: f127fe7 Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: 34892bf Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: 3af01e0 Pull-Request: #190822
> Allow template scheduling to claim a complete reduction chain as one atomic fusion unit and retain the template output identity needed by multi-output kernels. Prevent partial fusion from consuming layout or finalizer nodes before NVGEMM can validate the complete grouped pattern. > > Reject horizontal fusion involving NVGEMM boolean outputs and prioritize valid NVGEMM reduction chains ahead of ordinary Triton fusion candidates. > > Test Plan: > > ```bash > lintrunner -a > ``` > > Authored with an AI assistant. ghstack-source-id: 53be442 Pull-Request: #190822
Stack from ghstack (oldest at bottom):
cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @ipiszy @kadeng @muchulee8 @amjames @chauhang @aakhundov @coconutruben @jataylo