1.0: Branch-and-bound on schedules and schedule-quality follow-through
See CHANGES.md for details and ROADMAP.md for a more thorough overview than here. From the roadmap:
v1.0 — August 13, 2026 (released)
Theme: Advanced compiler tiers and schedule-quality follow-through
Advanced compiler tiers:
- Branch-and-bound schedule inference (#514) —
Ir.Schedule_spaceas a refinement tree over partial schedules,Cost_model.completion_flooris an admissible downward bound. - Inlining as a first-class, searchable schedule decision (#554, #555, #560).
- CUDA tensor-core profile completeness (#481), software-pipelined double-buffered staging (#487), CUDA/HIP graph capture of the fissioned step (#488), and budget-driven rematerialization on top of the liveness planner (#498).
- CPU reduced precision: 16-bit storage with f32 compute (#517) and native fp16 arithmetic (#516).
Schedule quality — the gpt2_mini arc (done):
- The v0.9 sweep left
gpt2_mini72x off torch CUDA; #569 took the tuned step 107.4 → 52.4 ms on CUDA and 45.6 → 25.4 ms on HIP. #528 made tensor cores reachable on transformer workloads at all. - CUDA bf16 mma timed scalar-fallback code under an mma label.
- Search-plumbing consolidation.
Frontend, configuration and diagnostics (done):
- Shape inference: "close down when known" narrowed to the leaf-tensor rule it always was, with
stretchrequesting use-site resolution by name (#544). - Config profiles
reproducible/performancewith picker-inherited precedence (#559); config startup chatter moved off stdout (#581). - Site-targeted materialization: the capability exists twice over and ships nowhere measured (#558).
Infrastructure:
- CI moved to OCaml 5.5, Windows off the per-PR path onto a twice-weekly schedule, and the GPU backends onto a daily cross-machine sweep (
tools/sweep.sh).