0.9: Affine analytic cost model, pad-to-tile scheduling, better tiling for convolutions
From the README:
- 0.9: Program search and optimization.
- Native affine program analysis and a
Schedule.op_legalityoracle; an analytic roofline cost model that picks untuned defaults and pre-filters the autotune beam (advisory throughout). - Deterministic split reductions, pad-to-tile scheduling (PADTO), static partitioning, epilogue fusion, and launch-time symbolic extents.
- Convolution schedule families: implicit GEMM via packing
Stage, blocked tile flavors, epilogue twins, compacting strided-row staging, and clamped-window pooling. - Mixed precision: a precision-assignment policy, master weights with cast twins, dynamic loss scaling with a fused on-device gate, tf32 matmuls, and forward-only reduced precision by load-time conversion.
- Liveness-based buffer aliasing; a packed
uniformthat is total over shapes and now backs default parameter initialization. - Search survivability: typed candidate-failure containment across Metal/CUDA/HIP, HIP scratch pre-validation, no unparallelized GPU dispatches, and a decline census that accounts for every refusal.
- A cross-machine benchmark sweep on Metal, CUDA and HIP, with the reports checked in under
benchmarks/.
- Native affine program analysis and a