You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development velocity. Primitives are released as experimental and will evolve based on user feedback.
NOTE: Primitives is a transitional API until a CUDA Python-like solution is available.
Introduced the Task Scheduling framework. This provides static analysis of execution schedules for warp-specialized kernels. Compilation stops when known concurrency issues are detected. Also provides tools for visualizing resource/task dependencies and analyzing a kernel's schedule.
Improved compiler diagnostics. Register spills and use of local memory can now be reported at compile time with source line numbers. Preliminary support for detecting classes of NVVM synchronization and execution hazards at compile-time when using the Primitives API. Better reporting of compiler errors that previously did not include source line numbers.
This release has been tested against the following packages:
Add support to specify data movement strategy for each operand being loaded/stored.
C++
Add implementation of 2-kernel backward targeting at FP8 in FMHA example.
Added a backward fused multi-head attention benchmark with multi-precision (FP16/FP8), configurable batch/sequence/head sizes, variable-length and masking options, plus built-in correctness checks and runtime/throughput reporting.
The 2-kernel backward has approximately 25% improvement compared to 1-kernel implementation at FP8 without mask on Blackwell SM103 chip.
Support CUDA 12.6 and newer structured bindings headers in NVRTC.
Add fp32/fp16/bf16/e4m3/e5m2 -> e2m1 (FP4) in NumericArrayConverter.
Optimize the E2M1 -> FP16 LUT decode helpers _e2m1_to_half_x2 and _e2m1_to_half_x4 by merging mask before prmt.
Fix some issues:
Update streamk heuristic algorithm to optimize some kernels with mix cluster sizes.
Avoid integer-sequence get ambiguity in CuTe tuple algorithms.
Fix a TMA creation driver bug: detect if the first 128KiB is mapped in conservatively by checking if the tensor is compact, if so it is valid to flip the bit otherwise zero it.
Various improvements and fixes from the community and CUTLASS team. Thanks to everyone who submitted PRs!
Optimal code generation with CUDA toolkit versions 13.3.
This discussion was created from the release CUTLASS 4.7.0.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
CuTe DSL
Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development velocity. Primitives are released as experimental and will evolve based on user feedback.
NOTE: Primitives is a transitional API until a CUDA Python-like solution is available.
Introduced the Task Scheduling framework. This provides static analysis of execution schedules for warp-specialized kernels. Compilation stops when known concurrency issues are detected. Also provides tools for visualizing resource/task dependencies and analyzing a kernel's schedule.
Improved compiler diagnostics. Register spills and use of local memory can now be reported at compile time with source line numbers. Preliminary support for detecting classes of NVVM synchronization and execution hazards at compile-time when using the Primitives API. Better reporting of compiler errors that previously did not include source line numbers.
This release has been tested against the following packages:
CUTLASS Operator API
C++
_e2m1_to_half_x2and_e2m1_to_half_x4by merging mask before prmt.This discussion was created from the release CUTLASS 4.7.0.
All reactions