Repository navigation
v0.1.15 #3311
LeiWang1999
announced in
Announcements
v0.1.15
#3311
Replies: 1 comment
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Highlights
AutoScheduletoAutoWarpSpecialization#3185): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups.T.gemm_blockscaledsemantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support.enumerate,zip, comprehensions, and generator expressions.Ascend 950
tilelang.ascend.languageandtarget="ascend"for Huawei Ascend 950 (dav-3510).T.SimdVFandT.SimtVFregions.tvm_ffiand Cython execution, PyTorch NPU tensors and streams, and NPU profiling.See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.
CUDA
TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples ([Feature] [CUDA] Role-based automatic warp specialization #3059, [Refactor] [CUDA] RenameAutoScheduletoAutoWarpSpecialization#3185).T.fmaandT.fmul, and complete additional FP16/BF16 math bridges ([Language][CUDA] Add T.fma and T.fmul round-to-nearest intrinsics #3134, [CUDA] Complete 16-bit bridges for CUDA-lowered unary math #3132, [CUDA] Add 16-bit hpow and hfmod bridges #3163).Language and Compiler
for ... in range(...)([Frontend] Support Python iterables and comprehensions #3230).elseclauses ([Bugfix][Language] Support unary plus on PrimExpr in the eager frontend #3141, [BugFix] Bind loop targets independently of mutable scalar variables #3232, [Frontend] Remove warnings for immutable variable rebinding #3262, [Bugfix][Language] Reject loop else clauses in the eager frontend #3142).T.copy/T.async_copycoalesced_widthas a hint and clamp it to the achievable vector width ([Fix]Clamp T.copy/T.async_copy coalesced_width to achievable vector size instead of LOG(FATAL) #3246).ROCm, CPU, and Metal
Runtime, Build, and Tooling
uint64argument mappings ([BugFix][JIT] Allocate a dynamic-shape output that precedes its sizing input #3207, [JIT] Add missing uint64 argument type mappings #3229).Compatibility Notes
syncandgroupparameters fromT.Pipelined, andk_packfromT.gemm_sp([Cleanup][Language] Remove dead T.Pipelined(sync/group) and T.gemm_sp(k_pack) parameters #3202).T.gemm(k_pack=...)should importtilelang.rocm.language([Refactor][Language] Move backend-specific op hints into their owning dialects #3203).threads=explicitly inT.Kernel([Refactor][Language] Make T.Kernel target-neutral with dialect-owned launch annotations #3186).T.symbolicremains available as a deprecated alias; useT.dynamicfor new code ([Language] Restore deprecated T.symbolic alias #3216).TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically ([Cache] Store cached kernel params as JSON instead of cloudpickle #3143, [CUDA][Cache] Publish immutable binary cache directories #3177).Full Changelog: v0.1.14...v0.1.15
What's Changed
T.Parallelloops #3121 by @Yongqi-Zhuo in [Refactor] Restore LoopUnswitching test altered by #3121 #3124AutoScheduletoAutoWarpSpecializationby @Yongqi-Zhuo in [Refactor] [CUDA] RenameAutoScheduletoAutoWarpSpecialization#3185New Contributors
Full Changelog: v0.1.14...v0.1.15
This discussion was created from the release v0.1.15.
All reactions