Skip to content

v0.1.15

Latest

Choose a tag to compare

@LeiWang1999 LeiWang1999 released this 30 Sep 00:29
· 34 commits to main since this release
a35f8dd

Highlights

  • Native Huawei Ascend 950 support (#3308): an end-to-end NPU backend with native code generation, automatic Cube/Vector scheduling and synchronization, and mixed SIMD/SIMT programming.
  • Automatic CUDA warp specialization (#3059, #3185): an opt-in role-based scheduler that assigns TMA loads, MMA computation, TMA stores, and worker operations to specialized warp groups.
  • Unified block-scaled GEMM (#3237, #3257, #3284): common T.gemm_blockscaled semantics with dedicated backend dispatch, improved SM100 instruction selection, and expanded SM120 fragment support.
  • More expressive Python frontend (#3230): compile-time iteration over Python iterables, enumerate, zip, comprehensions, and generator expressions.

Ascend 950

  • Add tilelang.ascend.language and target="ascend" for Huawei Ascend 950 (dav-3510).
  • Combine Cube GEMM and Vector computation in one kernel, with T.SimdVF and T.SimtVF regions.
  • Support explicit UB/L1/L0 storage, tiled copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM.
  • Add automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion.
  • Integrate Bisheng compilation, tvm_ffi and Cython execution, PyTorch NPU tensors and streams, and NPU profiling.
  • Include GEMM, DeepGEMM-style kernels, FlashAttention forward/backward, RMSNorm, and FP8 quantization examples.

See the Ascend 950 guide for installation and usage. Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects.

CUDA

  • Enable role-based automatic warp specialization with TL_ENABLE_AUTO_WARP_SPECIALIZATION: "role_based"; add GEMM and FlashAttention examples (#3059, #3185).
  • Extend SM120 block-scaled GEMM to fragment-resident A operands, row-major scale fragments, and odd per-warp atom grids; diagnose conflicting scale-fragment layouts (#3257, #3284).
  • Add round-to-nearest T.fma and T.fmul, and complete additional FP16/BF16 math bridges (#3134, #3132, #3163).
  • Preserve packed FP8 vector copies and fuse exact FP4-to-FP8 conversions through FP32 (#3276, #3204).
  • Fix WGMMA/UMMA K-panel strides for operand layouts and insert async-proxy fences for sparse MMA (#2965, #3139).
  • Correct atomic vectorization for invariant or non-contiguous destinations; keep shared FP32 atomics scalar on SM90 (#3129, #3219, #3238).
  • Fix NVRTC warp reductions, kernel-body assertions, boolean bitwise negation, and non-constant int4/uint4 broadcasts; restrict 256-bit global loads/stores to SM100+ (#3260, #3206, #3228, #3114, #3248).
  • Add an FP8 sparse MLA forward example for DeepSeek V3.2 on Hopper (#3224).

Language and Compiler

  • Make kernel launch encoding target-neutral, with launch options and operation hints owned by their backend dialects (#3186, #3203).
  • Support Python compile-time iteration and comprehensions while preserving device-loop semantics for for ... in range(...) (#3230).
  • Support unary plus on symbolic expressions, fix loop-variable binding, remove warnings for immutable rebinding, and reject unsupported loop else clauses (#3141, #3232, #3262, #3142).
  • Treat T.copy / T.async_copy coalesced_width as a hint and clamp it to the achievable vector width (#3246).
  • Preserve dynamic reduction tail guards, fix parallel-loop lowering with let inlining disabled, and restore async-copy lowering with partitioned layouts (#3294, #3269, #3278).
  • Improve reducer layout planning and symbolic layout validation; reject unsupported AllReduce thread strides and reduction NaN-propagation dtypes (#3171, #3233, #3266, #3273).
  • Fix NaN-propagating clamp lowering and eliminate unused bindings during simplification (#3205, #3293).

ROCm, CPU, and Metal

  • ROCm: support the DeepSeek V3.2 Top-K selector, wave64 Hadamard transforms, and GLM-5.3 k-pool examples, including Top-K index transformation (#3147, #3154, #3254).
  • Expand ROCm validation for attention kernels and portable examples, and document pip installation (#3146, #3148, #3165, #3168).
  • CPU: compute scalar GEMM products in the accumulator dtype and legalize BF16 arithmetic (#2917, #3201).
  • Metal: support 32-bit integer atomic add and respect GEMM buffer-region offsets (#3211, #3209).

Runtime, Build, and Tooling

  • Store cached kernel parameters as JSON instead of cloudpickle (#3143).
  • Publish CUDA binaries and metadata atomically in immutable cache directories, preventing readers from observing partially published entries (#3177).
  • Fix dynamic-output allocation when the sizing input follows the output, and add missing uint64 argument mappings (#3207, #3229).
  • Add wall-clock benchmarking and MPS-compatible timing helpers (#3234).
  • Improve CUDA/ROCm compiler discovery and library loading for symlinked installations (#2839, #3166, #3227).
  • Enable optimization for default single-config native builds and restore effective Windows wheel-build caching (#3191, #3133, #3305).
  • Validate built wheels on GPU runners and make performance regressions fail CI (#3167, #3182).

Compatibility Notes

  • Remove the unused sync and group parameters from T.Pipelined, and k_pack from T.gemm_sp (#3202).
  • ROCm kernels using T.gemm(k_pack=...) should import tilelang.rocm.language (#3203).
  • Kernels querying thread extents during tracing must specify threads= explicitly in T.Kernel (#3186).
  • T.symbolic remains available as a deprecated alias; use T.dynamic for new code (#3216).
  • Remove TILELANG_CACHE_VERIFY_HASH; binary artifact hash verification is now mandatory. Legacy cache formats are rebuilt automatically (#3143, #3177).

Full Changelog: v0.1.14...v0.1.15

What's Changed

  • [CI] Cap torch<2.14 for CUDA tests until flash-attn supports torch 2.14 by @LeiWang1999 in #3136
  • Bump transformers from 5.5.0 to 5.10.1 in /examples/bitnet-1.58b by @dependabot[bot] in #3131
  • [CI] Fix slow Windows wheel builds by making ccache effective by @LeiWang1999 in #3133
  • [Language][CUDA] Add T.fma and T.fmul round-to-nearest intrinsics by @LeiWang1999 in #3134
  • [CUDA] Complete 16-bit bridges for CUDA-lowered unary math by @Chennesxu in #3132
  • [Refactor] Restore LoopUnswitching test altered by #3121 by @Yongqi-Zhuo in #3124
  • [BugFix][CUDA] Support non-constant int4/uint4 broadcast by @jjppp in #3114
  • [BugFix][CPU] Compute scalar GEMM products in accum dtype instead of input dtype by @Dino1844 in #2917
  • [BugFix][Hopper][Blackwell] Read the WGMMA/UMMA K-panel stride from the layout instead of the operand extent by @bigSheep123 in #2965
  • [BugFix][Carver] Gate _legalize_info fallback on sm_version by @mocusez in #3145
  • [ROCm] Validate block-causal attention kernels by @andyluo7 in #3146
  • [ROCm] Validate attention sink kernels by @andyluo7 in #3148
  • [ROCm] Support DeepSeek-V3.2 Top-K selector by @andyluo7 in #3147
  • [Bugfix][Language] Reject loop else clauses in the eager frontend by @rishabhsinha17 in #3142
  • [Example] Fix hadamard reference precision on TF32 and add pytest entry for CI by @Guan-jeans in #3126
  • [BugFix][Carver] Compare sm_version numerically in plan_rasterization by @mocusez in #3144
  • [Misc] Expose .agents skills to Claude Code via .claude/skills symlinks by @LeiWang1999 in #3161
  • [Language] Cache dtype.as_torch and demote storage-dtype fallback logs to debug by @LeiWang1999 in #3160
  • [Refactor][CUDA] Replace per-lane-count vector constructors with variadic packers by @LeiWang1999 in #3158
  • [CUDA] Add 16-bit hpow and hfmod bridges by @Chennesxu in #3163
  • [BugFix][ROCm] Resolve hipcc robustly instead of trusting bare PATH lookup by @LeiWang1999 in #3166
  • [CI] Validate built wheels on GPU runners in the Dist workflow by @LeiWang1999 in #3167
  • [Doc] Document pip installation on AMD GPUs (ROCm) by @LeiWang1999 in #3168
  • [Bugfix][Language] Support unary plus on PrimExpr in the eager frontend by @rishabhsinha17 in #3141
  • [CUDA] Add async-proxy fences for sparse MMA intrinsics by @ZenAlexa in #3139
  • [Feature] [CUDA] Role-based automatic warp specialization by @Yongqi-Zhuo in #3059
  • [CUDA] Resolve symlinked nvcc before deriving CUDA_HOME by @morluto in #2839
  • [Language] Remove deprecated T.symbolic alias by @SiriusNEO in #3122
  • [ROCm] Support wave64 Hadamard transforms by @andyluo7 in #3154
  • [Cache] Store cached kernel params as JSON instead of cloudpickle by @rishabhsinha17 in #3143
  • [BugFix][Transform] Keep shared FP32 atomics scalar on SM90 by @ZenAlexa in #3129
  • [Layout] Consider scalar reducer plans in register-count search by @LeiWang1999 in #3171
  • [Transform] Add snapshot mode to FreshenMutableReads by @LeiWang1999 in #3174
  • [CUDA][Cache] Publish immutable binary cache directories by @LeiWang1999 in #3177
  • [CI] [pre-commit.ci] autoupdate by @pre-commit-ci[bot] in #3181
  • [CI] Fail the perf bot on a regression instead of only reporting it by @cklxx in #3182
  • [Refactor] [CUDA] Rename AutoSchedule to AutoWarpSpecialization by @Yongqi-Zhuo in #3185
  • [Refactor][Language] Make T.Kernel target-neutral with dialect-owned launch annotations by @LeiWang1999 in #3186
  • [Cleanup][Language] Remove dead T.Pipelined(sync/group) and T.gemm_sp(k_pack) parameters by @LeiWang1999 in #3202
  • [Refactor][Language] Move backend-specific op hints into their owning dialects by @LeiWang1999 in #3203
  • [BugFix][CPU] Legalize BF16 arithmetic in the CPU pipeline by @anerli in #3201
  • [Language] Restore deprecated T.symbolic alias by @LeiWang1999 in #3216
  • [Refactor][JIT] Replace callee-allocated-output target sniffing with a backend capability flag by @LeiWang1999 in #3218
  • [Transform] Restore bind-before-use order when merging constraint sets by @LeiWang1999 in #3221
  • [BugFix][Metal] Respect GEMM buffer region offsets by @anerli in #3209
  • [Runtime] Fix library loading for symlink installs by @sepcnt in #3227
  • [CUDA] Fix boolean bitwise negation codegen by @sepcnt in #3228
  • [Example] FP8 sparse MLA forward for DeepSeek V3.2 on Hopper by @xuebozhang525-alt in #3224
  • [Profiler] Split timing helpers and add wall-clock benchmarking by @SiriusNEO in #3234
  • [CPU] Import the CPU dialect in CPU tests by @penguin-wwy in #3236
  • [BugFix][Vectorize] Keep atomic_add scalar for invariant/non-contiguous destinations by @Dino1844 in #3219
  • [Transform][CUDA] Plan atomic vector widths from destination addresses by @LeiWang1999 in #3238
  • [Do not review][Op][Language][CUDA] Add common block-scaled GEMM semantics and backend dispatch by @LeiWang1999 in #3237
  • [JIT] Add missing uint64 argument type mappings by @sepcnt in #3229
  • [BugFix] Restore symbolic loop-layout injectivity proof; reject layouts on symbolic shared tiles by @sepcnt in #3233
  • [ROCm] Run portable example validation in CI by @andyluo7 in #3165
  • [BugFix][CUDA] Only emit 256-bit global load/store on SM100+ targets by @penguin-wwy in #3248
  • [Metal] Support 32-bit integer atomic add by @anerli in #3211
  • [Testing] Drop duplicated codegen-smoke tests; fix fastmath self-comparison assertions by @penguin-wwy in #3252
  • [ROCm] Add GLM-5.3 k-pool Top-K transform by @andyluo7 in #3254
  • [Frontend] Support Python iterables and comprehensions by @sepcnt in #3230
  • [Frontend] Remove warnings for immutable variable rebinding by @LJC00118 in #3262
  • [BugFix] Reject non-power-of-two AllReduce thread strides by @LeiWang1999 in #3266
  • [NVRTC] Fix warp reduction compilation by @sepcnt in #3260
  • [BugFix][CUDA] Lower a kernel-body assert to a device-legal check by @xy200303 in #3206
  • [BugFix][JIT] Allocate a dynamic-shape output that precedes its sizing input by @xy200303 in #3207
  • [CUDA] Fuse exact FP4 to FP8 conversion through FP32 by @ZenAlexa in #3204
  • [Fix]Clamp T.copy/T.async_copy coalesced_width to achievable vector size instead of LOG(FATAL) by @edragain2nd in #3246
  • [BugFix] Bind loop targets independently of mutable scalar variables by @sepcnt in #3232
  • [BugFix][Language] Lower NaN-propagating clamp through device templates by @xy200303 in #3205
  • [Fix] Reject unsupported reduction NaN propagation dtypes by @ZenAlexa in #3273
  • [BugFix] Fix parallel loop lowering with let inlining disabled by @penguin-wwy in #3269
  • [Transform][CUDA] Fix async copy lowering with partitioned layouts by @LeiWang1999 in #3278
  • [CUDA] Keep FP8 vector copies packed by @LeiWang1999 in #3276
  • [CUDA] Support SM120 block-scaled GEMM fragments and odd warp atom grids by @sepcnt in #3257
  • [CUDA] Reject conflicting SM120 scale fragment layouts by @LeiWang1999 in #3284
  • [Build] Optimize default single-config builds by @KellyFrog in #3191
  • [Fix][TVM] Update TVM to preserve dynamic reduction tail guards by @LeiWang1999 in #3294
  • [BugFix] Eliminate unused bindings in Simplify by @LJC00118 in #3293
  • [Build][Windows] Restore ccache hits for clang-cl wheel builds by @LeiWang1999 in #3305
  • [Public Release 9/30] Introduce Ascend 950 backend by @SiriusNEO in #3308
  • [Release] Bump version into 0.1.15 by @LeiWang1999 in #3309

New Contributors

Full Changelog: v0.1.14...v0.1.15