feat(nvidia): link vLLM moe_wna16_marlin_gemm - #912
Merged
Conversation
voltjia
force-pushed
the
feat/linked-moe-wna16-marlin-gemm
branch
from
August 9, 2026 11:10
3f89ad2 to
3173d65
Compare
voltjia
marked this pull request as ready for review
August 9, 2026 13:32
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
_moe_C::moe_wna16_marlin_gemmin slot 16.b_bias_or_none,a_scales, andthread_k/thread_n/blocks_per_sm; removeis_ep; rename the low-level type id tob_type_id.outconvention.Motivation
vLLM uses this Marlin grouped GEMM primitive in its Ampere-compatible fused MoE path. Linking the installed provider keeps the kernel in the provider DSO and avoids copying another vendor kernel into InfiniOps.
N/A - no linked issue.
Type of Change
feat- new feature / new operator / new platformfix- bug fixperf- performance improvement (no behavioral change)refactor- code restructuring without behavior changetest- adding or fixing tests onlydocs- documentation onlybuild/ci- build system or CI configurationchore- tooling, formatting, or other non-code changesPlatforms Affected
WITH_CPU)WITH_NVIDIA)WITH_ILUVATAR)WITH_METAX)WITH_CAMBRICON)WITH_MOORE)WITH_ASCEND)WITH_TORCH)Smoke Test Result
Focused Release/NDEBUG rebuild and smoke revalidation for commit
3173d651is in progress. The NVIDIA SSH endpoint became unavailable during setup; this section will be replaced with final command output rather than retaining the obsolete vLLM v0.10 result.The current-schema provider itself has already been probed on A100:
This runtime provider is vLLM 0.20.2, not the source alignment baseline; its registered schema is byte-for-byte equivalent to v0.26.0/current for this operator.
Test Results on Supported Platforms
Completed local validation
Benchmark / Performance Impact
N/A - this PR exposes the installed provider implementation and makes no performance claim.
Notes for Reviewers
568afb3aand was rechecked against main commit83ad767e.c_or_noneis output storage. InfiniOps requires a trailingout, passes it in the provider's second position, and verifies that the returned Tensor aliases it.thread_k/thread_n/blocks_per_smremain explicit-1arguments here; no convenience overload is added.float4_e2m1f,global_scale, and FP8 activation are rejected before dispatch because the current InfiniRT type system cannot represent the required FP8 tensors.fused_marlin_moecomposite is outside this PR and requires separate redesign.