Skip to content

0.9: Affine analytic cost model, pad-to-tile scheduling, better tiling for convolutions

Choose a tag to compare

@lukstafi lukstafi released this 03 Aug 14:06
· 3464 commits to master since this release
b339bca

From the README:

  • 0.9: Program search and optimization.
    • Native affine program analysis and a Schedule.op_legality oracle; an analytic roofline cost model that picks untuned defaults and pre-filters the autotune beam (advisory throughout).
    • Deterministic split reductions, pad-to-tile scheduling (PADTO), static partitioning, epilogue fusion, and launch-time symbolic extents.
    • Convolution schedule families: implicit GEMM via packing Stage, blocked tile flavors, epilogue twins, compacting strided-row staging, and clamped-window pooling.
    • Mixed precision: a precision-assignment policy, master weights with cast twins, dynamic loss scaling with a fused on-device gate, tf32 matmuls, and forward-only reduced precision by load-time conversion.
    • Liveness-based buffer aliasing; a packed uniform that is total over shapes and now backs default parameter initialization.
    • Search survivability: typed candidate-failure containment across Metal/CUDA/HIP, HIP scratch pre-validation, no unparallelized GPU dispatches, and a decline census that accounts for every refusal.
    • A cross-machine benchmark sweep on Metal, CUDA and HIP, with the reports checked in under benchmarks/.