1.0 -- Advanced compiler tiers and schedule-quality follow-through #583
lukstafi
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
See CHANGES.md for details and ROADMAP.md for a more thorough overview than here. From the roadmap:
v1.0 — August 13, 2026 (released)
Theme: Advanced compiler tiers and schedule-quality follow-through
Advanced compiler tiers:
Ir.Schedule_spaceas a refinement tree over partial schedules,Cost_model.completion_flooris an admissible downward bound.Schedule quality — the
gpt2_miniarc (done):gpt2_mini72x off torch CUDA; Companion-coverage rule pins five FFN-class kernels to 1.3% of fp32 peak — 70% of the gpt2_mini step, ~2.4x measured floor #569 took the tuned step 107.4 → 52.4 ms on CUDA and 45.6 → 25.4 ms on HIP. Autotune sketch seeding: rank-3 (batched/attention) matmul sites are never proposed, so gpt2_mini cannot reach tensor cores even after #521 #528 made tensor cores reachable on transformer workloads at all.Frontend, configuration and diagnostics (done):
stretchrequesting use-site resolution by name (Shape inference: 'close down when known' for operations, empty batch for parameters — use-site row resolution is a leaf-tensor rule applied too widely #544).reproducible/performancewith picker-inherited precedence (Config profiles: reproducible vs performance preset bundles, with picker-inherited precedence #559); config startup chatter moved off stdout (Config startup chatter prints to stdout, polluting every ocannl-linked executable's redirectable output #581).Infrastructure:
tools/sweep.sh).This discussion was created from the release 1.0 -- Advanced compiler tiers and schedule-quality follow-through.
All reactions