v0.1.14 #3135
LeiWang1999
announced in
Announcements
v0.1.14
#3135
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Highlights
T.alloc_reducerreworked into first-class deferred reduction epochs whose physical lowering is planned by layout inference (via a first-class PartialFragment layout). Adds loop-scoped epochs, conditional reducer finalization, and automatic vectorization of contiguous reducer updates.Language
T.unroll(explicit=True)for early explicit unrolling ([Refactor] RecycleT.unroll(explicit=True)for early explicit unrolling #2859)cluster_maskonT.tma_copy([Enhancement] Expose cluster_mask on T.tma_copy #2932)T.gemmtile dimensions with a clear message ([Bugfix] Reject symbolic T.gemm tile dimensions with a clear message #3113), validateT.gemmk_packarguments ([Bugfix] Validate T.gemm k_pack arguments #3094), reject non-positivearrive_countinalloc_barrier/alloc_cluster_barrier([Lang] Reject non-positive arrive_count in alloc_barrier and alloc_cluster_barrier #3112), rejectT.Parallelindexing of local buffers ([Analysis] RejectT.Parallelindexing of local buffers #3041), rejectbreakin fully expanded loops ([BugFix][Transform] Reject break in fully expanded loops (#3026) #3078)CUDA
tcgen05.allocarenas ([CUDA] Pack logical TMEM buffers into sharedtcgen05.allocarenas #2831); support half-subpartition (M=64) TMEM tiles intcgen05.ld/st([Layout] Support half-subpartition (M=64) TMEM tiles in tcgen05.ld/st #2880); fix ld/st segment pointer advancement in b32 columns ([BugFix][CUDA] Advance tcgen05 ld/st segment pointers in b32 columns #2952)int4x2/uint4x2codegen ([BugFix][CUDA] Support int4x2 and uint4x2 codegen #3036); 16-bit CUTLASS type overloads for fast-math,__ldg, and htan intrinsics ([CUDA] Add 16-bit overloads for CUTLASS fast-math functions #3097, [CUDA] Bridge half-style math intrinsics for 16-bit CUTLASS types #3077, [CUDA] Add __ldg overloads for 16-bit CUTLASS types #3028, [BugFix][CUDA] Provide htan overloads for fp16/bf16 tangent #2894)T.{reads,writes}forT.tma_{gather4,scatter4}([CUDA] RemoveT.{reads,writes}forT.tma_{gather4,scatter4}#3053); separate TMA atomic-add dtype support from layout encoding ([CUDA][TMA] Separate atomic-add dtype support from layout encoding #2846)ROCm and other backends
tl.sync_warpon HIP ([BugFix][ROCm] Emit a compiler barrier for tl.sync_warp on HIP #2872), reject sub-wavefront block sizes instead of crashing ([BugFix][ROCm] Reject sub-wavefront block sizes instead of crashing #2918), resolve versioned device properties in the HIP stub ([ROCm] Resolve versioned device properties in HIP stub #2919); emit#linedirectives for the HIP target ([CodeGen][ROCm] Emit #line directives for HIP target #3058)Compiler / Transform
T.Parallelloops ([BugFix] Always vectorizeT.Parallelloops #3121); scalarize Select in automatic vectorization ([BugFix] Scalarize Select in automatic vectorization #3060)VerifyBufferInit, a general buffer-initialization check ([Transform] Add VerifyBufferInit, a general buffer-initialization check #2956)#linedirectives from TIR spans ([CodeGen] Emit #line directives from TIR spans #3048)Runtime / JIT / Build
get_parent_localsframe self-reference leak ([Cherry-Pick][BugFix] Fix get_parent_locals frame self-reference leak #2934)project()([Bugfix][CMake] Fix reconfigure aborting in FindPipCUDAToolkit before project() #3102)Tooling / Ecosystem
What's Changed
tcgen05.allocarenas by @Rachmanino in [CUDA] Pack logical TMEM buffers into sharedtcgen05.allocarenas #2831T.unroll(explicit=True)for early explicit unrolling by @Yongqi-Zhuo in [Refactor] RecycleT.unroll(explicit=True)for early explicit unrolling #2859T.Parallelindexing of local buffers by @SiriusNEO in [Analysis] RejectT.Parallelindexing of local buffers #3041T.{reads,writes}forT.tma_{gather4,scatter4}by @Yongqi-Zhuo in [CUDA] RemoveT.{reads,writes}forT.tma_{gather4,scatter4}#3053_tir_packed_to_unsigned_convert_with_zeroscomputesq - zeroin the unsigned storage dtype, silently wrapping instead of dequantizing #2947) by @Junius-Wynn in [Bugfix][Quantize] Fix unsigned zero-point decode underflow (#2947) #3118T.loop_break()insideT.unroll(explicit=True)emits barebreak;outside any loop, so nvcc rejects the kernel instead of it compiling #3026) by @KellyFrog in [BugFix][Transform] Reject break in fully expanded loops (#3026) #3078T.Parallelloops by @Yongqi-Zhuo in [BugFix] Always vectorizeT.Parallelloops #3121New Contributors
Full Changelog: v0.1.13...v0.1.14
This discussion was created from the release v0.1.14.
All reactions