Releases
v0.11.0
Compare
Sorry, something went wrong.
No results found
laggui
released this
06 Oct 18:45
What's Changed
Add tiled layout (#1329 ) @Sublime12
Bump version to 0.11.0-pre.1 (#1326 ) @laggui
feat: streaming: add stream priority hint, wire into CUDA backend (#1324 ) @lilith
perf: lighter cpu runtime (#1330 ) @marcantoinem
feat: Add support for VK_EXT_shader_long_vector (#1188 ) @wingertge
mega-refactor: Totally change the frontend to enable references, among other things (#1322 ) @wingertge
Add f64 support back for CUDA. (#1321 ) @vaijira
fix: Ensure constant block args are generated in the correct block (#1332 ) @wingertge
Add workgroupUniformLoad primitive (#1327 ) @ArthurBrussee
fix: shared memories allocation calculation were wrong (#1337 ) @marcantoinem
make renderdoc optional (does not compile on mac os) (#1339 ) @louisfd
Metal: fix atomic syntax (#1340 ) @louisfd
Metal: missing space caused compilation errors (#1342 ) @louisfd
Feat/obfuscation (#1341 ) @nathanielsimard
fix metal compilation (#1343 ) @louisfd
fix: Fix more compilation issues on metal (#1347 ) @wingertge
feat(runtime): add memory_usage_total aggregating all streams (#1333 ) @ArthurBrussee
Refactor/wgpu compilers (#1346 ) @nathanielsimard
Improve CPU backend by removing need to flush, add atomics test and add atomic to cubecl-cpu (#1345 ) @marcantoinem
fix: CubeOr expand methods computed AND instead of OR (#1350 ) @LucaCappelletti94
fix: short-circuit || and && in #[cube] code (#1348 ) @LucaCappelletti94
refactor: Add a lifetime to views (#1344 ) @wingertge
fix: Assign/init_mut (#1355 ) @wingertge
fix: cubecl-cpu synchronisation problem (#1356 ) @marcantoinem
fix: skip kernel launch when cube count is zero (#1349 ) @LucaCappelletti94
fix: Fix atomics on CUDA again (#1354 ) @wingertge
Refactor/readme (#1358 ) @nathanielsimard
fix: re-export wgpu::Backend from cubecl-wgpu (#1357 ) @zhan-wei-919
refactor(runtime): aggregate memory_usage + memory_cleanup in client (#1360 ) @ArthurBrussee
Tiling as a composable view layout (#1362 ) @louisfd
fix(wgpu): bound live timestamp query sets per device (Metal counter-sample-buffer exhaustion) (#1361 ) @AdrianEddy
fix shared memory bytes (#1367 ) @louisfd
fix: reject unsupported kernel argument types instead of panicking (#1373 ) @LucaCappelletti94
fix: don't panic when calling a non-path expression in a kernel (#1372 ) @LucaCappelletti94
fixed typo check and bitwise test errors (#1369 ) @Andy2887
fix: don't panic when const-folding an unfoldable const expression (#1374 ) @LucaCappelletti94
book: constants (plane) (#1320 ) @Redhawk18
feat: add expm1 lowering across backends (#1301 ) @shinaoka
refactor: Simplify Variable to align it with existing IRs (#1378 ) @wingertge
Feat/bytes improvements (#1379 ) @nathanielsimard
Fix imports (#1383 ) @nathanielsimard
Fix handling one tuple (#1382 ) @akiradeveloper
fixes return Self::Scalar in impl_unary_func_scalar_out (Fixes #1283 ) (#1338 ) @ethqnol
Feat/fix abusive allocation in CPU backend (#1385 ) @marcantoinem
feat: Coordinates CoordsDynI type (#1386 ) @SamuelBelanger
Add read_lazy method for non-WASM targets in ComputeClient (#1392 ) @jwric
add as read to view mut (#1393 ) @louisfd
Cpu dump improvement (#1394 ) @marcantoinem
add citation (#1335 ) @Redhawk18
burn element trait related changes for e4m3, e5m2 (#1389 ) @skewballfox
Fix: returning panic payload in device channel tasks to caller and re-raising panic for blocking APIs (#1376 ) @Andy2887
feat: metal backend (#1175 ) @dcvz
fix(metal): attribute launch errors to the issuing stream (#1403 ) @dcvz
fix(metal): validate shared memory limit before pipeline creation (#1405 ) @dcvz
fixed cargo xtask validate fails on Linux (#1397 ) @Andy2887
refactor: replace dirs with etcetera (drops MPL-2.0 option-ext) (#1395 ) @shimwell
Feat/new CPU runtime (#1400 ) @marcantoinem
Peak device throughput (#1408 ) @ThierryCantin-Demers
disaggregate-oob-consts (#1409 ) @nathanielsimard
Fix/cpu index unit (#1411 ) @louisfd
fix(cpu): reserve shared memory from a dedicated pool (#1412 ) @louisfd
Feat/dynamic memory pool config (#1417 ) @nathanielsimard
Feat/hip captured graph (#1415 ) @nathanielsimard
update wgpu version (#1416 ) @Charles23R
Fix/graph safety (#1419 ) @nathanielsimard
Disable persistent tune cache option (#1423 ) @SamuelBelanger
Feat/autotune throughput (#1422 ) @SamuelBelanger
fix(cpu): parallelize sync_cube and clamp test cube dims to core count (#1424 ) @louisfd
Cpu: Allow more threads than cores (#1426 ) @louisfd
Wmma compile error (#1427 ) @nathanielsimard
add cross-stream input bindings pinning wgpu (#1434 ) @Charles23R
fix atomic import (#1441 ) @Charles23R
Allow nested tuple destructure (#1420 ) @akiradeveloper
feat: add AsIndex fallible dimension index conversion (#1445 ) @mattisonchao
chore: use tracel-llvm version 22.1.4-5 (#1447 ) @laggui
fix: add missing cubecl-metal publish (#1448 ) @laggui
Feat/autotune observability (#1437 ) @SamuelBelanger
CubeCL Environment: multiplatform HPC runtime API (#1435 ) @nathanielsimard
Fix initialize memory failure with error handling + fallback (#1454 ) @nathanielsimard
Invalidate the kernel cache when a kernel changes (#1455 ) @nathanielsimard
fix(quant): map QuantParam::UE4M3 to FloatKind::E4M3 (#1450 ) @ThierryCantin-Demers
Feat/dry run environment (#1457 ) @nathanielsimard
fix(runtime): check PendingDropQueue during gpu writes to prevent u32 overflow (#1456 ) @for-lack-of-a-better-name-j
Adaptive autotune scheduler with early elimination (#1449 ) @SamuelBelanger @nathanielsimard
feat(runtime): DeviceProperties::identity (#1465 ) @nathanielsimard
feat(tune): roofline bounds from a Work amount and per-resource thresholds (#1462 ) @SamuelBelanger
fix(float): correct the e4m3 and e5m2 limits and exponent constants (#1472 ) @nathanielsimard
feature(cubecl-cuda): add cuda graph capture and replay (#1469 ) @michaelvsinko
fix(cuda): reclaim and retry a failed device reserve (#1468 ) @laurigates
fix(common): give device runner threads a shutdown path (#1470 ) @nathanielsimard
fix(runtime): classify driver out-of-memory apart from BufferTooBig (#1474 ) @ThierryCantin-Demers
Add a two-level quantization level (#1453 ) @nathanielsimard
fix(common): call ceil through num_traits so round_up builds on no_std (#1475 ) @ThierryCantin-Demers
fix(cpp): materialize constant bitcasts as lvalues (#1477 ) @jcwal1516
fix(common): release the caller's borrows before waking it in run_scoped (#1478 ) @ThierryCantin-Demers
fix(spirv): name the FPEncoding variants added in tracel-rspirv 0.13.3 (#1484 ) @louisfd
fix(cubecl-hip): implement sync_warp instead of emitting #error (#1483 ) @nathanielsimard
feat(common): add ComptimeFloat and Ratio for hashable comptime kernel params (#1482 ) @SamuelBelanger
chore: version Cargo.lock (#1485 ) @ThierryCantin-Demers
chore: bump version to 0.11.0-pre.2 (#1489 ) @laggui
fix(publish): add missing cubecl-environment (#1490 ) @laggui
fix(publish): add missing cubecl-std requirement to cubecl-metal (#1492 ) @laggui
Feat/throughput memory curve (#1493 ) @louisfd
feat(common): support negative numerators in Ratio (#1491 ) @SamuelBelanger
fix Mutex imports in cubecl-metal (#1481 ) @olukowski
refactor: Make everything Pliron (#1402 ) @wingertge
Fix memory read throughput runner (#1495 ) @wingertge
Reorder passes to avoid CondBrOp panic (#1496 ) @wingertge
Fix features for cubecl-spirv (#1497 ) @wingertge
Suffixes the llvm.is.fpclass intrinsic with the type to fix crash in cubek-reduce (#1498 ) @wingertge
Add missing dynamic extract lowering for CPU (#1499 ) @wingertge
Clean up dead regions before lowering branches to avoid weird crash in DCE (#1500 ) @wingertge
Fix ue8m0 display (#1501 ) @wingertge
Apply the global scale in the quantized view for two level quant (#1479 ) @ThierryCantin-Demers
Improve memory management and memory pools (#1494 ) @nathanielsimard
Add SROA pass @wingertge
Cache metadata info uniforms in the wgpu runtime (#1504 ) @nathanielsimard
Zero-init arrays (#1511 ) @wingertge
Fix MSL codegen (#1507 ) @louisfd
Fix/pliron codegen regressions (#1508 ) @nathanielsimard
Add version check to ensure required features are supported in MSL (#1512 ) @wingertge
Fix atomics and elect on CUDA and WGSL (#1517 ) @wingertge
Fix Pliron timer panic on WebAssembly (#1519 ) @syl20bnr
Refactor/comptime option launch derive (#1513 ) @ThierryCantin-Demers
Fix/unroll initializer type (#1515 ) @louisfd
Update to syn 3.0 (#1521 ) @wingertge
Declare wasm-bindgen-futures for WASM runtime (#1518 ) @syl20bnr
Fix cubecl-runtime build for wasm32 (#1524 ) @ArthurBrussee
Sync before the load in workgroup_uniform_load (#1525 ) @ArthurBrussee
fix broken merge (#1528 ) @wingertge
Add two-level caching for CUDA and HIP (#1526 ) @wingertge
Wgpu graph capture (#1505 ) @nathanielsimard
Replace QuantLevel with per-tensor and per-block scale levels (#1514 ) @ThierryCantin-Demers
Feat/quant lookup (#1532 ) @louisfd
Fix/cpu shared memory (#1533 ) @nathanielsimard
feat(cpu): implement thread affinity and topology on Windows (#1537 ) @ThierryCantin-Demers
ci: pin the shared actions to v11 (#1535 ) @ThierryCantin-Demers
Add non-pointer MMA dialect (#1530 ) @wingertge
Fix/plane reductions and emitter errors (#1520 ) @nathanielsimard
feat(quant): Software fp8 conversion on every backend, native where the hardware has it (#1534 ) @ThierryCantin-Demers
chore: tracel github actions v11 (#1539 ) @syl20bnr
chore: fix clippy lints (#1541 ) @laggui
refactor: split cubecl-cpu in two to extract cubecl-llvm (#1542 ) @marcantoinem
fix(metal): correct MSL capability detection (#1540 ) @laggui
fix(wgpu): keep the query set current names out of cleanup (#1543 ) @nathanielsimard
chore: update to xtask v5 (#1544 ) @syl20bnr
fix(cpu): yield between idle polls instead of spinning (#1545 ) @ThierryCantin-Demers
feat(bench): write-only memory probe and roofline scoring for benchmarks (#1546 ) @ThierryCantin-Demers
fix(unroll): unroll dynamic vector extract and insert (#1547 ) @nathanielsimard
chore: bump version to 0.11.0-pre.3 (#1555 ) @laggui
fix(cubecl-llvm): add missing publish and use released tracel-llvm-bundler (#1556 ) @laggui
fix(runtime): lazy error handling on the logical stream (#1548 ) @nathanielsimard
Perf/add fma detection (#1550 ) @marcantoinem
feat(cpu): report the device last level cache size (#1551 ) @ThierryCantin-Demers
fix(throughput): probe a working set at its own footprint (#1552 ) @ThierryCantin-Demers
fix(throughput): walk the last window of every cycle (#1554 ) @ThierryCantin-Demers
Fix/e2m1 software codec (#1557 ) @nathanielsimard
fix(llvm): name an overloaded intrinsic for the type it is overloaded on (#1561 ) @ThierryCantin-Demers
feat(cpu): Vectorize exp,log,sin,cos,tanh with polynomial approximation (#1553 ) @ThierryCantin-Demers
Feat/erased sink (#1563 ) @nathanielsimard
Feat/execution observer (#1573 ) @nathanielsimard
Refactor/runtimes (#1568 ) @nathanielsimard
fix(ir): an scf.if reports its condition as its result to SCCP (#1574 ) @nathanielsimard
fix(llvm): give ue8m0 a type, and report it on the CPU backend (#1565 ) @ThierryCantin-Demers
rename the fetch_update deprecrated method to remove miri warnings (#1569 ) @ThierryCantin-Demers
fix(cpp): an accumulator fragment names no layout (#1570 ) @ThierryCantin-Demers
fix(plane): document the shuffle up and down boundary as undefined (#1564 ) @ThierryCantin-Demers
fix(tests): stop f16 tolerating an error of eight in assert_equals_approx (#1567 ) @ThierryCantin-Demers
Add data flow solver (#1572 ) @wingertge
feat: add complex support and validation (#1300 ) @shinaoka
Feat/add amdgpu backend (#1571 ) @marcantoinem
fix(throughput): stop the memory curve measuring cache, and collapse its modes (#1559 ) @ThierryCantin-Demers
Llvm amdgpu fixes (#1577 ) @nathanielsimard
fix(core): pin a scalar's WithScalar to itself (#1578 ) @ThierryCantin-Demers
Refactor/amdgpu llvm review (#1582 ) @nathanielsimard
Revert "rename the fetch_update deprecrated method to remove miri warnings (#…" (#1583 ) @ThierryCantin-Demers
feat(profiling): device timing for the HIP and CUDA backends (#1580 ) @nathanielsimard
feat: add dp4a intrinsic for packed int8×4 DOT accumulate (#1471 ) @saivishwak @wingertge
Chore/reduce dependencies (#1587 ) @nathanielsimard
fix: stop advertising tensor cores on dies that have none (#1588 ) @ThierryCantin-Demers
fix macos llvm (#1586 ) @louisfd
Fix/throughput probe rotation (#1589 ) @nathanielsimard
Add fine-grained memory clobbers for gpu_asm (#1576 ) @wingertge
fix(cpp): emit casts using destination type (#1549 ) @y7nieSEl5
feat: update tracel llvm (#1593 ) @marcantoinem
test(runtime): coordinate the profile tests with the refusal toggle (#1579 ) @ThierryCantin-Demers
fix(cuda): stop advertising MMA and tensor cores on dies that have none (#1591 ) @ThierryCantin-Demers
fix(throughput): measure the peaks the hardware can carry, and say why a device has none (#1598 ) @ThierryCantin-Demers
fix(wgpu): close a device profile on its last compute pass (#1592 ) @ThierryCantin-Demers
fix(spirv, wgsl): emit real fmas and fuse a product into the subtraction consuming it (#1562 ) @ThierryCantin-Demers
Refactor/runtime erasure (#1590 ) @nathanielsimard
refactor: ComputeClient and ComputeServer are Client and Server (#1603 ) @nathanielsimard
fix(topology): a cube's units are contiguous in ABSOLUTE_POS (#1595 ) @ThierryCantin-Demers
feat(device): make Device a device in its own right (#1604 ) @nathanielsimard
Feat/metadata tiling (#1605 ) @louisfd
Fix/ordered tune groups (#1615 ) @nathanielsimard
feat(tune): an eviction the tuner runs before every measured sample (#1621 ) @nathanielsimard
Feat/metadata logical shape (#1622 ) @louisfd
fix(cpu): report the host CPU as the device name (#1623 ) @ThierryCantin-Demers
fix: add integer powi lowering (#1624 ) @laggui
fix(throughput): stop reporting ceilings the hardware beats (#1607 ) @ThierryCantin-Demers
feat(zspace): the stride and shape mutators refuse a storage-tiled metadata (#1627 ) @louisfd
fix(ir): decline the self-operand folds when the result is a vector (#1626 ) @ThierryCantin-Demers
feat: Uniformity analysis (#1620 ) @wingertge
fix: don't compile AMDGPU for MacOS (#1597 ) @marcantoinem
fix(cpp): lower CUDA complex casts with cuComplex helpers (#1581 ) @shinaoka
fix(cuda): brace the ldmatrix destination when the vector is 1 wide (#1596 ) @michaelvsinko
fix(llvm): gate the last AmdGpu match arm on the amdgpu feature (#1631 ) @louisfd
fix(llvm): gate the AMDGPU arm of lang_tag (#1630 ) @ThierryCantin-Demers
fix(wgpu): a window with no timing reports the absence, not a zero (#1632 ) @louisfd
Fix/profile/absence is not zero (#1634 ) @louisfd
fix format (#1637 ) @ThierryCantin-Demers
fix(llvm): keep amdgpu out of cubecl-llvm's default features (#1638 ) @ThierryCantin-Demers
fix: metal kernel corruption (#1636 ) @Marc-AnthonyG
Refactor/crate layout (#1606 ) @nathanielsimard
fix(runtime): disambiguate kernel entrypoint names and avoid launch-path allocation (#1625 ) (#1633 ) @Sadik00789
fix audit (#1644 ) @ThierryCantin-Demers
ci: pin the miri job to nightly-2026-09-09 (#1640 ) @ThierryCantin-Demers
feat(cpu): evaluate f16 arithmetic in f32 where the host has none (#1629 ) @ThierryCantin-Demers
chore: bump version to 0.11.0-pre.4 (#1648 ) @laggui
fix(runtime): saturate FlushingPolicyState counters to prevent add-overflow panic (#1359 ) (#1413 ) @teddytennant
device fence (#1654 ) @louisfd
fix(server): skip copy-layout validation for empty transfers (#1660 ) @laggui
feat: MemorySSA (#1656 ) @wingertge
fix(throughput): probe memory outside the installed pools, and fail when it cannot be allocated (#1659 ) @SamuelBelanger
feat(wgpu): write host data straight into mapped storage on integrated Vulkan GPUs (#1651 ) @SamuelBelanger
chore: update to pliron 0.18 (#1662 ) @laggui
fix launch overhead (#1664 ) @louisfd
Feat/cuda llvm backend (#1609 ) @nathanielsimard
Fix/hip runtime on no hip platform (#1649 ) @syl20bnr
Feat/profiling (#1666 ) @nathanielsimard
fix(runtime): refuse collectives on a runtime with no device transport (#1653 ) @ThierryCantin-Demers
feat(ir): report the physical card behind a device (#1652 ) @ThierryCantin-Demers
fix(llvm): build the bitcode linker shim for NVPTX too (#1667 ) @ThierryCantin-Demers
feat: state the device's capacity in MemoryDeviceProperties (#1672 ) @nathanielsimard
fix(opt): walk memory phis on an explicit stack (#1669 ) @SamuelBelanger
chore: use tracel-rspirv and plrion-spirv 0.15.0 (#1681 ) @laggui
refactor e2m1 (#1682 ) @louisfd
Feat/environment records (#1680 ) @nathanielsimard
feat: fix hip test on machine without ROCm support and bump llvm (#1679 ) @syl20bnr
fix(ci): publish LLVM before CUDA (#1683 ) @laggui
Refactor/turso persistence (#1677 ) @jwric
fix(llvm): emit the PTX version the CUDA driver loads (#1684 ) @SamuelBelanger
fix(throughput): release the probe pool lock before the device call (#1678 ) @ThierryCantin-Demers
fix(throughput): probe the device with nothing else running on it (#1686 ) @ThierryCantin-Demers
feat(environment): open an environment database at its path (#1688 ) @nathanielsimard
Perf/llvm gpu parity (#1687 ) @nathanielsimard
Remove NoMemoryEffect from sync ops (#1696 ) @wingertge
Feat/adaptive pool (#1685 ) @nathanielsimard
Fix some wgpu-msl tests downstream (#1701 ) @louisfd
docs: update cubecl docs (#1700 ) @marcantoinem
Feat/tune plan record (#1702 ) @nathanielsimard
fix(macros): make kernel expansion deterministic (#1705 ) @jwric
fix(wgpu): preserve float literal precision in WGSL codegen (#1708 ) @laggui
feat(cuda): default to sync allocation and make allocator configurable (#1709 ) @laggui
fix(llvm): the GPU targets read PLANE_POS (#1714 ) @nathanielsimard
fix(core): warn on constant shared barrier elections (#1639 ) @arcusbuilds
feat(zspace): a tiling holds up to 16 pieces a dim (#1723 ) @louisfd
feat(runtime): let a client ask whether its stream is capturing a graph (#1720 ) @nathanielsimard
feat(cpu): report the host's real register file, and size IO by its own width (#1661 ) @ThierryCantin-Demers
docs(ir): cut the register file comments down to what the code cannot say (#1724 ) @ThierryCantin-Demers
fix(runtime): keep the capture record a projection of the stream states (#1728 ) @nathanielsimard
Feat/tuner elected (#1727 ) @nathanielsimard
fix(llvm): clamp unsigned divisors so LSR can reduce addresses behind them (#1731 ) @SamuelBelanger
fix: correct signed remainder semantics and add NaN-propagating extrema (#1732 ) @laggui
fix(cuda): handle zero-size buffers without device allocation (#1733 ) @laggui
Fix/spirv storage visibility (#1737 ) @nathanielsimard
fix(wgpu): reject unsupported explicit MSL runtime (#1734 ) @laggui
chore(lint): allow redundant_field_names pending upstream Clippy fix (#1741 ) @laggui
feat(wgpu): add fallible initialization and external device registration (#1740 ) @laggui
fix(no-std): restore fetch_update for portable-atomic compatibility (#1742 ) @laggui
feat: add ability for device to probe usage (#1735 ) @Marc-AnthonyG
feat: Sparse backward, dense forward and backward, etc (#1743 ) @wingertge
Perf/e2m1 pair decode (#1746 ) @louisfd
fix(runtime): a panic while expanding a kernel fails its launch (#1747 ) @louisfd
Feat/sync point errors (#1719 ) @Charles23R
fix(std): keep a reshape out of the perpendicular copy (#1671 ) @Liberxue
fix(cuda): offer m16n8k8 from Turing, not Volta (#1717 ) @Liberxue
fix(throughput): answer a declined key without taking the device (#1710 ) @Liberxue
Fix/metal profile window drains (#1751 ) @louisfd
fix(llvm): enable TF32 tensor cores on CUDA (#1750 ) @laggui
fix(runtime): skip PageUpdateForbidden backtrace under serializable (#1755 ) @antimora
fix(wgpu): restore WGSL execution on WASM (#1619 ) @ArthurBrussee
feat(tune): stop an ordered group once its members stop beating its best (#1756 ) @nathanielsimard
Feat/llvm native bf16 (#1759 ) @nathanielsimard
fix(llvm): promise nothing about a buffer some cube writes (#1761 ) @nathanielsimard
chore: use tracel-rspirv and pliron-spirv 0.16.0 (#1766 ) @laggui
fix(llvm): mangle count-bits intrinsic names by type (#1768 ) @nathanielsimard
chore: update version to 0.11.0 (#1769 ) @laggui
You can’t perform that action at this time.