Enable NVRTC PCH reuse for CUTLASS providers - #11048
Draft
tpn wants to merge 1 commit into
Draft
Conversation
Create process-private NVRTC precompiled headers automatically and reuse them across compatible provider compilations. Preserve safe fallback and fork/concurrency isolation, and expose phase telemetry without changing the public primitive API. Signed-off-by: Trent Nelson <trent@trent.me>
Contributor
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This was referenced Aug 27, 2026
miscco
reviewed
Aug 28, 2026
| # PCH must register its child reset before cache registers its fork handlers. | ||
| # isort: off | ||
| from . import ( | ||
| _pch as _provider_pch, |
Contributor
There was a problem hiding this comment.
I am slightly scared that this might create hard to debug user issues if there is some bad include order somewhere
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this is needed
CUTLASS provider compilation repeatedly parses the same heavy CUDA/CUB header
preamble. This layer lets compatible NVRTC compilations in one process share
a process-managed precompiled header while preserving the existing provider
API and a safe non-PCH fallback.
What this PR includes
isolation, cleanup, and phase telemetry
The final stack layer owns AOT capture and publication.
Stack
6 of 7, based on PR #11047.
The final layer adds relocatable CUTLASS provider AOT packs.
Local validation
selection passed 87 tests
one PCH creation, one hit, and matching GPU output
imports,
compileall, changed-file pre-commit, andgit diff --checkpassedWheel and installed-package qualification, GitHub CI, and automated AI review
are deferred for now.