Skip to content

ONNX Runtime CUDA Plugin EP 0.2.0

Choose a tag to compare

@tianleiwu tianleiwu released this 08 Oct 22:11

This release brings expanded quantized inference and attention support, kernel performance improvements, and
reliability fixes to the ONNX Runtime CUDA Plugin EP. Highlights include 2-bit MatMulNBits, INT4 paged KV caches,
speculative decoding improvements, compact recurrent-state updates, and safer CUDA Graph execution across sessions.

Please refer to the CUDA Plugin EP Quick Start for usage and installation guidance. The matched ONNX Runtime version is 1.30.0.

Highlights

Plugin Runtime and CUDA Graphs

  • Gave each plugin EP session its own device arena, preventing cross-session reuse of CUDA Graph buffers and honoring per-session arena options (#32807).
  • Added workspace-memory accounting and reporting for resource-constrained graph partitioning, including packed attention workspace recipes and estimates, and optional-input-aware shape handling (#31962, #32283, #32312, #32321).
  • Made GatherND capture-safe and kept CudaAsyncBuffer staging buffers alive across graph replay (#32120, #32121).
  • Fixed shared-cache scratch lifetime in CUDA MHA, reset plugin EP stream chunks before release, and corrected the user-stream test lifetime (#31968, #31983, #32449).
  • Fixed CUDA plugin device discovery on WSL, with improved hardware-identity matching and handling of devices missing from platform discovery (#32517).

Attention and KV Caching

  • Added bidirectional GroupQueryAttention support and an is_causal attribute to PagedAttention (#31704, #32225).
  • Added INT4 paged KV caches with per-channel scales, including FP16 XQA decode and speculative-decode paths. Packed INT4 cache storage uses half the bytes of INT8 cache storage; availability is controlled by the onnxruntime_USE_INT4_KV_CACHE build option, which defaults to enabled (#32515).
  • Enabled split-KV for paged FlashAttention decode and improved PagedAttention dispatch diagnostics (#32102, #32099).
  • Expanded paged XQA with group-size-6 support, head size 256, FP16-cache decode at head size 256, and native block tables for 128-token pages (#32108, #32229, #32263, #32127).
  • Added paged XQA speculative verification for metadata-bounded query groups of 2-8 tokens at head size 256 and group size 6, including ragged batches and INT8, FP8, and native FP16/BF16 KV caches (#32340).
  • Added the session.gqa_value_layout session option for applications using a BNHS GroupQueryAttention Value cache, with graph-inserted layout conversions; BNSH remains the default (#32139).

New Operator and Model Support

  • Added compact GatedDeltaNet with transition capture for replaying accepted speculative prefixes, plus BF16 support (#32282, #32307).
  • Added VarlenCausalConvWithState for packed, variable-length continuous batching, and compact state updates that avoid duplicating full convolution-state checkpoints for speculative prefixes (#32168, #32290).
  • Added CUDA kernels for the DeepSeek Engram contrib operators EngramGate and NGramHashMapping (#32268).
  • Registered BF16 ReduceMean kernels (#32326).

Quantized MatMul and MoE

  • Added 2-bit MatMulNBits support, including dequantization and specialized GEMV/batched paths, and fused floating-point-activation/integer-weight (fpA_intB) GEMM/GEMV kernels for eligible FP16/BF16 workloads (#32693, #32699).
  • Made compact fpA_intB kernels the default build configuration, extended compact MatMulNBits to BF16, and refined GEMV eligibility checks. The complete kernel matrix remains available through onnxruntime_USE_FPA_INTB_GEMM_FULL (#32324, #32721, #32338).
  • Enabled FP4 QMoE kernels by default in CUDA builds and added Windows support for Blackwell SM120 (#32096, #32163).
  • Added an opt-in FP8 DeepGEMM decode path for eligible SM90 QMoE workloads, controlled by ORT_QMOE_FP4_DEEPGEMM and disabled by default. This path is disabled on Windows because of upstream header requirements (#32122, #32485).
  • Bounded QMoE workspace with configurable row tiling and bounded FP8 weight-dequantization scratch by tiling over N (#32097, #32129).
  • Vectorized NVFP4 weight dequantization for prefill, tuned NVFP4 GEMV tiling for Qwen multi-token prediction, and extended speculative-decode GEMVs to 64 rows (#32128, #32140, #32289).
  • Tuned FP4/FP8 GEMV scheduling for 48-SM SM121 GPUs and improved FP8 GEMV residency for grids just past two blocks per SM (#32408, #32409, #32433).
  • Reduced expensive fpA_intB tactic-profiling work and added optional, size-gated M-row chunking for large MatMulNBits workloads. Chunking is disabled by default and can be configured with ep.cuda.matmul_nbits_m_chunk_size or ORT_MATMULNBITS_M_CHUNK_SIZE (#32758, #32810).

General Kernel Performance

  • Parallelized ArgMax/ArgMin over wide last axes, optimized wide-last-axis TopK, and accelerated low-lane INT64 CumSum (#32092, #32404, #32238).
  • Added a single-memcpy Slice fast path for contiguous subregions (#28902).
  • Passed Concat per-input metadata by value and avoided pinned buffers in CUDA Split and Concat fast paths (#32119, #32410).

Reliability and Correctness

  • Fixed ScatterElements reduction dispatch by element type and Abs signed-zero handling (#29879, #31477).
  • Handled zero-element inputs/bias in BiasGelu and FastGelu as no-ops and zero-sized outputs in CUDA random generator kernels (#31698, #31997).
  • Added bounds and shape validation for per-element Split sizes, 8-bit MatMulNBits g_idx, GatherElements counts, and ScatterND index depth (#29461, #31643, #32030, #32034).
  • Validated CUDA NMS mask sizes, QDQ element counts, per-channel ImageScaler bias, and Crop inputs (#32014, #32029, #32002, #32157).
  • Bounded RemovePadding sequence-token counts and Whisper beam-search cross-QK layer/head indices (#31994, #31998).
  • Hardened integer arithmetic in RotaryEmbedding, SparseAttention, CUDA reduction scans, and softmax offsets (#31995, #31996, #32137, #32330).

Build, Packaging, and Dependencies

  • Upgraded CUTLASS to 4.7, cuDNN frontend to 1.27, and Protobuf to 33.6 (#32111, #29906).
  • Updated CUDA architecture selections and plugin package-test pipelines, and adjusted Linux AArch64 build parallelism (#31989, #32072, #32165).
  • Fixed Windows CUDA 12.9 SM120 compilation and Windows ARM64 CUDA plugin packaging (#32114, #32355).
  • Fixed PagedAttention builds without FlashAttention and added the CUDA 13 CCCL include path to plugin builds (#32327, #32392).
  • Used authenticated package feeds, pinned GitHub Actions to full commit SHAs, and updated artifact upload/download actions (#32005, #32176, #32207, #32209).

Contributors

Thanks to the 17 human contributors who contributed to this release:

@apsonawane, @arnej27959, @baijumeswani, @chilo-ms, @danfiedler-msft, @DKAIN-py, @edgchen1, @eserscor, @javier-intel, @jiafatom, @justinchuby, @kunal-vaishnavi, @MohamedElashri, @Noperi0r, @sanaa-hamel-microsoft,
@tianleiwu, @titaiwangms


This release covers changes since CUDA Plugin EP v0.1.0 affecting CUDA Plugin EP code, CUDA kernels, build integration,
and packaging. The highlights focus on user-facing changes rather than listing every infrastructure-only update.

This summary was drafted with AI assistance from commit history and PR metadata.

Full Changelog: plugin-ep-cuda/v0.1.0...plugin-ep-cuda/v0.2.0