Skip to content

1.0: Branch-and-bound on schedules and schedule-quality follow-through

Choose a tag to compare

@lukstafi lukstafi released this 13 Aug 14:08
· 3092 commits to master since this release
b362145

See CHANGES.md for details and ROADMAP.md for a more thorough overview than here. From the roadmap:

v1.0 — August 13, 2026 (released)

Theme: Advanced compiler tiers and schedule-quality follow-through

Advanced compiler tiers:

  • Branch-and-bound schedule inference (#514) — Ir.Schedule_space as a refinement tree over partial schedules, Cost_model.completion_floor is an admissible downward bound.
  • Inlining as a first-class, searchable schedule decision (#554, #555, #560).
  • CUDA tensor-core profile completeness (#481), software-pipelined double-buffered staging (#487), CUDA/HIP graph capture of the fissioned step (#488), and budget-driven rematerialization on top of the liveness planner (#498).
  • CPU reduced precision: 16-bit storage with f32 compute (#517) and native fp16 arithmetic (#516).

Schedule quality — the gpt2_mini arc (done):

  • The v0.9 sweep left gpt2_mini 72x off torch CUDA; #569 took the tuned step 107.4 → 52.4 ms on CUDA and 45.6 → 25.4 ms on HIP. #528 made tensor cores reachable on transformer workloads at all.
  • CUDA bf16 mma timed scalar-fallback code under an mma label.
  • Search-plumbing consolidation.

Frontend, configuration and diagnostics (done):

  • Shape inference: "close down when known" narrowed to the leaf-tensor rule it always was, with stretch requesting use-site resolution by name (#544).
  • Config profiles reproducible / performance with picker-inherited precedence (#559); config startup chatter moved off stdout (#581).
  • Site-targeted materialization: the capability exists twice over and ships nowhere measured (#558).

Infrastructure:

  • CI moved to OCaml 5.5, Windows off the per-PR path onto a twice-weekly schedule, and the GPU backends onto a daily cross-machine sweep (tools/sweep.sh).