Repository navigation
Highlights
- Native Huawei Ascend 950 support (#3308): an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming.
- Automatic CUDA warp specialization (#3059, #3185): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups.
- Unified block-scaled GEMM (#3237, #3257, #3284): common
T.gemm_blockscaledsemantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support. - More expressive Python frontend (#3230): compile-time iteration over Python iterables,
enumerate,zip, comprehensions, and generator expressions.
Ascend 950
- Add
tilelang.ascend.languageandtarget="ascend"for Huawei Ascend 950 (dav-3510). - Combine Cube GEMM and Vector computation in one kernel, with
T.SimdVFandT.SimtVFregions. - Support explicit UB/L1/L0 storage, tiled copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM.
- Add automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion.
- Integrate Bisheng compilation,
tvm_ffiand Cython execution, PyTorch NPU tensors and streams, and NPU profiling. - Include GEMM, DeepGEMM-style kernels, FlashAttention forward/backward, RMSNorm, and FP8 quantization examples.
See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.
CUDA
- Enable role-based automatic warp specialization with
TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples (#3059, #3185). - Extend SM120 block-scaled GEMM to fragment-resident A operands, row-major scale fragments, and odd per-warp atom grids; diagnose conflicting scale-fragment layouts (#3257, #3284).
- Add round-to-nearest
T.fmaandT.fmul, and complete additional FP16/BF16 math bridges (#3134, #3132, #3163). - Preserve packed FP8 vector copies and fuse exact FP4-to-FP8 conversions through FP32 (#3276, #3204).
- Fix WGMMA/UMMA K-panel strides for operand layouts and insert async-proxy fences for sparse MMA (#2965, #3139).
- Correct atomic vectorization for invariant or non-contiguous destinations; keep shared FP32 atomics scalar on SM90 (#3129, #3219, #3238).
- Fix NVRTC warp reductions, kernel-body assertions, boolean bitwise negation, and non-constant int4/uint4 broadcasts; restrict 256-bit global loads/stores to SM100+ (#3260, #3206, #3228, #3114, #3248).
- Add an FP8 sparse MLA forward example for DeepSeek V3.2 on Hopper (#3224).
Language and Compiler
- Make kernel launch encoding target-neutral, with launch options and operation hints owned by their backend dialects (#3186, #3203).
- Support Python compile-time iteration and comprehensions while preserving device-loop semantics for
for ... in range(...)(#3230). - Support unary plus on symbolic expressions, fix loop-variable binding, remove warnings for immutable rebinding, and reject unsupported loop
elseclauses (#3141, #3232, #3262, #3142). - Treat
T.copy/T.async_copycoalesced_widthas a hint and clamp it to the achievable vector width (#3246). - Preserve dynamic reduction tail guards, fix parallel-loop lowering with let inlining disabled, and restore async-copy lowering with partitioned layouts (#3294, #3269, #3278).
- Improve reducer layout planning and symbolic layout validation; reject unsupported AllReduce thread strides and reduction NaN-propagation dtypes (#3171, #3233, #3266, #3273).
- Fix NaN-propagating clamp lowering and eliminate unused bindings during simplification (#3205, #3293).
ROCm, CPU, and Metal
- ROCm: support the DeepSeek V3.2 Top-K selector, wave64 Hadamard transforms, and GLM-5.3 k-pool examples, including Top-K index transformation (#3147, #3154, #3254).
- Expand ROCm validation for attention kernels and portable examples, and document pip installation (#3146, #3148, #3165, #3168).
- CPU: compute scalar GEMM products in the accumulator dtype and legalize BF16 arithmetic (#2917, #3201).
- Metal: support 32-bit integer atomic add and respect GEMM buffer-region offsets (#3211, #3209).
Runtime, Build, and Tooling
- Store cached kernel parameters as JSON instead of cloudpickle (#3143).
- Publish CUDA binaries and metadata atomically in immutable cache directories, preventing readers from observing partially published entries (#3177).
- Fix dynamic-output allocation when the sizing input follows the output, and add missing
uint64argument mappings (#3207, #3229). - Add wall-clock benchmarking and MPS-compatible timing helpers (#3234).
- Improve CUDA/ROCm compiler discovery and library loading for symlinked installations (#2839, #3166, #3227).
- Enable optimization for default single-config native builds and restore effective Windows wheel-build caching (#3191, #3133, #3305).
- Validate built wheels on GPU runners and make performance regressions fail CI (#3167, #3182).
Compatibility Notes
- Remove the unused
syncandgroupparameters fromT.Pipelined, andk_packfromT.gemm_sp(#3202). - ROCm kernels using
T.gemm(k_pack=...)should importtilelang.rocm.language(#3203). - Kernels querying thread extents during tracing must specify
threads=explicitly inT.Kernel(#3186). T.symbolicremains available as a deprecated alias; useT.dynamicfor new code (#3216).- Remove
TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically (#3143, #3177).
Full Changelog: v0.1.14...v0.1.15
What's Changed
- [CI] Cap torch<2.14 for CUDA tests until flash-attn supports torch 2.14 by @LeiWang1999 in #3136
- Bump transformers from 5.5.0 to 5.10.1 in /examples/bitnet-1.58b by @dependabot[bot] in #3131
- [CI] Fix slow Windows wheel builds by making ccache effective by @LeiWang1999 in #3133
- [Language][CUDA] Add T.fma and T.fmul round-to-nearest intrinsics by @LeiWang1999 in #3134
- [CUDA] Complete 16-bit bridges for CUDA-lowered unary math by @Chennesxu in #3132
- [Refactor] Restore LoopUnswitching test altered by #3121 by @Yongqi-Zhuo in #3124
- [BugFix][CUDA] Support non-constant int4/uint4 broadcast by @jjppp in #3114
- [BugFix][CPU] Compute scalar GEMM products in accum dtype instead of input dtype by @Dino1844 in #2917
- [BugFix][Hopper][Blackwell] Read the WGMMA/UMMA K-panel stride from the layout instead of the operand extent by @bigSheep123 in #2965
- [BugFix][Carver] Gate _legalize_info fallback on sm_version by @mocusez in #3145
- [ROCm] Validate block-causal attention kernels by @andyluo7 in #3146
- [ROCm] Validate attention sink kernels by @andyluo7 in #3148
- [ROCm] Support DeepSeek-V3.2 Top-K selector by @andyluo7 in #3147
- [Bugfix][Language] Reject loop else clauses in the eager frontend by @rishabhsinha17 in #3142
- [Example] Fix hadamard reference precision on TF32 and add pytest entry for CI by @Guan-jeans in #3126
- [BugFix][Carver] Compare sm_version numerically in plan_rasterization by @mocusez in #3144
- [Misc] Expose .agents skills to Claude Code via .claude/skills symlinks by @LeiWang1999 in #3161
- [Language] Cache dtype.as_torch and demote storage-dtype fallback logs to debug by @LeiWang1999 in #3160
- [Refactor][CUDA] Replace per-lane-count vector constructors with variadic packers by @LeiWang1999 in #3158
- [CUDA] Add 16-bit hpow and hfmod bridges by @Chennesxu in #3163
- [BugFix][ROCm] Resolve hipcc robustly instead of trusting bare PATH lookup by @LeiWang1999 in #3166
- [CI] Validate built wheels on GPU runners in the Dist workflow by @LeiWang1999 in #3167
- [Doc] Document pip installation on AMD GPUs (ROCm) by @LeiWang1999 in #3168
- [Bugfix][Language] Support unary plus on PrimExpr in the eager frontend by @rishabhsinha17 in #3141
- [CUDA] Add async-proxy fences for sparse MMA intrinsics by @ZenAlexa in #3139
- [Feature] [CUDA] Role-based automatic warp specialization by @Yongqi-Zhuo in #3059
- [CUDA] Resolve symlinked nvcc before deriving CUDA_HOME by @morluto in #2839
- [Language] Remove deprecated T.symbolic alias by @SiriusNEO in #3122
- [ROCm] Support wave64 Hadamard transforms by @andyluo7 in #3154
- [Cache] Store cached kernel params as JSON instead of cloudpickle by @rishabhsinha17 in #3143
- [BugFix][Transform] Keep shared FP32 atomics scalar on SM90 by @ZenAlexa in #3129
- [Layout] Consider scalar reducer plans in register-count search by @LeiWang1999 in #3171
- [Transform] Add snapshot mode to FreshenMutableReads by @LeiWang1999 in #3174
- [CUDA][Cache] Publish immutable binary cache directories by @LeiWang1999 in #3177
- [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in #3181
- [CI] Fail the perf bot on a regression instead of only reporting it by @cklxx in #3182
- [Refactor] [CUDA] Rename
AutoScheduletoAutoWarpSpecializationby @Yongqi-Zhuo in #3185 - [Refactor][Language] Make T.Kernel target-neutral with dialect-owned launch annotations by @LeiWang1999 in #3186
- [Cleanup][Language] Remove dead T.Pipelined(sync/group) and T.gemm_sp(k_pack) parameters by @LeiWang1999 in #3202
- [Refactor][Language] Move backend-specific op hints into their owning dialects by @LeiWang1999 in #3203
- [BugFix][CPU] Legalize BF16 arithmetic in the CPU pipeline by @anerli in #3201
- [Language] Restore deprecated T.symbolic alias by @LeiWang1999 in #3216
- [Refactor][JIT] Replace callee-allocated-output target sniffing with a backend capability flag by @LeiWang1999 in #3218
- [Transform] Restore bind-before-use order when merging constraint sets by @LeiWang1999 in #3221
- [BugFix][Metal] Respect GEMM buffer region offsets by @anerli in #3209
- [Runtime] Fix library loading for symlink installs by @sepcnt in #3227
- [CUDA] Fix boolean bitwise negation codegen by @sepcnt in #3228
- [Example] FP8 sparse MLA forward for DeepSeek V3.2 on Hopper by @xuebozhang525-alt in #3224
- [Profiler] Split timing helpers and add wall-clock benchmarking by @SiriusNEO in #3234
- [CPU] Import the CPU dialect in CPU tests by @penguin-wwy in #3236
- [BugFix][Vectorize] Keep atomic_add scalar for invariant/non-contiguous destinations by @Dino1844 in #3219
- [Transform][CUDA] Plan atomic vector widths from destination addresses by @LeiWang1999 in #3238
- [Do not review][Op][Language][CUDA] Add common block-scaled GEMM semantics and backend dispatch by @LeiWang1999 in #3237
- [JIT] Add missing uint64 argument type mappings by @sepcnt in #3229
- [BugFix] Restore symbolic loop-layout injectivity proof; reject layouts on symbolic shared tiles by @sepcnt in #3233
- [ROCm] Run portable example validation in CI by @andyluo7 in #3165
- [BugFix][CUDA] Only emit 256-bit global load/store on SM100+ targets by @penguin-wwy in #3248
- [Metal] Support 32-bit integer atomic add by @anerli in #3211
- [Testing] Drop duplicated codegen-smoke tests; fix fastmath self-comparison assertions by @penguin-wwy in #3252
- [ROCm] Add GLM-5.3 k-pool Top-K transform by @andyluo7 in #3254
- [Frontend] Support Python iterables and comprehensions by @sepcnt in #3230
- [Frontend] Remove warnings for immutable variable rebinding by @LJC00118 in #3262
- [BugFix] Reject non-power-of-two AllReduce thread strides by @LeiWang1999 in #3266
- [NVRTC] Fix warp reduction compilation by @sepcnt in #3260
- [BugFix][CUDA] Lower a kernel-body assert to a device-legal check by @xy200303 in #3206
- [BugFix][JIT] Allocate a dynamic-shape output that precedes its sizing input by @xy200303 in #3207
- [CUDA] Fuse exact FP4 to FP8 conversion through FP32 by @ZenAlexa in #3204
- [Fix]Clamp T.copy/T.async_copy coalesced_width to achievable vector size instead of LOG(FATAL) by @edragain2nd in #3246
- [BugFix] Bind loop targets independently of mutable scalar variables by @sepcnt in #3232
- [BugFix][Language] Lower NaN-propagating clamp through device templates by @xy200303 in #3205
- [Fix] Reject unsupported reduction NaN propagation dtypes by @ZenAlexa in #3273
- [BugFix] Fix parallel loop lowering with let inlining disabled by @penguin-wwy in #3269
- [Transform][CUDA] Fix async copy lowering with partitioned layouts by @LeiWang1999 in #3278
- [CUDA] Keep FP8 vector copies packed by @LeiWang1999 in #3276
- [CUDA] Support SM120 block-scaled GEMM fragments and odd warp atom grids by @sepcnt in #3257
- [CUDA] Reject conflicting SM120 scale fragment layouts by @LeiWang1999 in #3284
- [Build] Optimize default single-config builds by @KellyFrog in #3191
- [Fix][TVM] Update TVM to preserve dynamic reduction tail guards by @LeiWang1999 in #3294
- [BugFix] Eliminate unused bindings in Simplify by @LJC00118 in #3293
- [Build][Windows] Restore ccache hits for clang-cl wheel builds by @LeiWang1999 in #3305
- [Public Release 9/30] Introduce Ascend 950 backend by @SiriusNEO in #3308
- [Release] Bump version into 0.1.15 by @LeiWang1999 in #3309
New Contributors
- @Dino1844 made their first contribution in #2917
- @mocusez made their first contribution in #3145
- @rishabhsinha17 made their first contribution in #3142
- @Guan-jeans made their first contribution in #3126
- @anerli made their first contribution in #3201
- @xuebozhang525-alt made their first contribution in #3224
- @xy200303 made their first contribution in #3206
Full Changelog: v0.1.14...v0.1.15