xtc-v0.3.1
XTC Release xtc-v0.3.1
We're pleased to announce XTC version xtc-v0.3.1. Thanks to all contributors.
Important features
- MLIR consumer fusion
addsfuse_consumer_at()support to the MLIR tensor backend. This enables
consumers such as ReLU operations to be fused into matrix multiplications and
convolutions. - Producer and consumer fusion in schedule descriptions
makes fusion transformations available through the scheduling description DSL,
with coverage for both MLIR and TVM. - TVM split schedule support
allows partitioned TVM loops to be tiled and vectorized by rebasing loops with
non-zero minima while preserving their original index offsets. - Selectable exploration module output types
adds the--module-typeoption to choose shared-library, archive, or C-source
output and fixes TVM archive linking outside the current directory.
Important fixes
- Fix incorrectly calculated TVM tiling factors
derives split factors from adjacent tile sizes instead of cumulative sizes,
preventing oversized outer factors and incorrect generated schedules. - Fix schedules with an omitted interchange
restores the default tiling interchange order when no explicit interchange is
provided, avoiding aKeyErrorduring schedule construction.
Changes
- explore: support selectable module output types by @guillon
- schedules: added producer and consumer fusion to descript by @liamsemeria
- mlir: ease the interoperability with a local LLVM checkout by @qaco
- [schedule][mlir] consumer fusion by @liamsemeria
- tests: fixed duplicate dump file names by @liamsemeria
- tvm: support split schedules by @guillon
- schedule: fix issue with missing interchange by @guillon
- tvm: fix incorrectly calculated tiling factors by @guillon
New contributors
There were no first-time contributors in this release.
Comments
Still remaining issues, to be created as issues:
- performance not in line for the MLIR backend w/ or w/o --use-tensor w.r.t. TVM in particular on conv2d
- still extremely long compilation times in MLIR backends w/ or w/o --use-tensor
Small search (20 trials) on conv2d for 4 threads on a laptop Intel Core i5-1135G7 (11th Gen), 4 cores / 8 threads, 2.40-4.20 GHz turbo:
- TVM 90.04% peak (4 compiles/sec),
- MLIR 3.36% peak !! (4 secs/compile !!)
- MLIR --use-tensor 28.50% peak ! (3 secs/compile !!)
$ uv pip install 'xtc-tools[tvm,mlir]==0.3.1'
$ for backend in tvm mlir "mlir --use-tensor"; do uv run loop-explore --threads 4 --peak-flops 65.6e9 --operator conv2d --backends $backend --strategy tile8dv --trials 20; done
conv2d evaluate: 100%|██████████████████████████████████████████████████| 20/20 [00:11<00:00, 1.75it/s]
conv2d compile: 100%|███████████████████████████████████████████████████| 20/20 [00:05<00:00, 3.99it/s]
conv2d execute: 100%|███████████████████████████████████████████████████| 20/20 [00:06<00:00, 3.11it/s]
Schedule: tvm: [1; 1; 1; 16; 1; 1; 1; 7; 1; 2; 1; 16; 1; 7; 1; 1]: time: 0.50 msecs, peak perf: 90.04%
conv2d evaluate: 100%|██████████████████████████████████████████████████| 20/20 [01:29<00:00, 4.49s/it]
conv2d compile: 100%|███████████████████████████████████████████████████| 20/20 [01:23<00:00, 4.17s/it]
conv2d execute: 100%|███████████████████████████████████████████████████| 20/20 [00:06<00:00, 3.07it/s]
Schedule: mlir: [1; 1; 1; 16; 1; 1; 1; 7; 1; 2; 1; 16; 1; 7; 1; 1]: time: 13.39 msecs, peak perf: 3.36%
conv2d evaluate: 100%|██████████████████████████████████████████████████| 20/20 [01:17<00:00, 3.85s/it]
conv2d compile: 100%|███████████████████████████████████████████████████| 20/20 [01:10<00:00, 3.55s/it]
conv2d execute: 100%|███████████████████████████████████████████████████| 20/20 [00:06<00:00, 3.29it/s]
Schedule: mlir: [1; 1; 1; 2; 14; 1; 16; 7; 1; 1; 1; 64; 7; 1; 3; 0]: time: 1.58 msecs, peak perf: 28.50%