Submit work through per-task immediate command lists - #610
Draft
michel2323 wants to merge 1 commit into
Draft
Conversation
Every kernel launch, copy, and fill used to create a fresh command list, submit it to the task's command queue, and drop the reference — leaving destruction of thousands of driver objects (lists, command buffers, heaps) to finalizer timing. Under launch storms that garbage is what pushes the driver into allocation failure, where NEO's error handling is at its worst (the scratch path aborts outright). It is also pure overhead: the per-dispatch list costs ~8x in submission latency. Replace the per-dispatch machinery with a per-task oneStream holding one in-order asynchronous immediate command list: appends submit directly, and the garbage source disappears entirely. Level Zero >= 1.9 is required; there is no fallback submission path. oneMKL work still needs a real command queue for SYCL interop, so each stream lazily creates a companion queue — a separate execution stream, which makes the previously implicit ordering between Julia kernels and oneMKL calls explicit: sycl_queue drains the immediate list before handing out the SYCL queue (Julia -> MKL), and a dirty flag makes the next Julia-side submission drain the companion queue (MKL -> Julia). FFT plans capture their queue at construction, so their _exec! methods apply the boundary themselves. The LTS drain-before-free machinery follows the shape change: the queue registry becomes a stream registry, draining both the immediate list and the companion queue before a buffer referenced by in-flight work is freed; immediate lists get the same bounded-drain finalizer as queues. The sync-each-submission workaround now host-synchronizes the list after each append. KA.priority! swaps the task's stream for one with the requested priority. The scratch hedge moves to the stream, and remains on the explicit-queue compatibility path (@oneapi queue=...), which still submits through a per-dispatch list. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WSFxSBXtckG3BVAT12wYkf
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## scratch-hedge #610 +/- ##
=================================================
+ Coverage 78.80% 79.05% +0.25%
=================================================
Files 50 50
Lines 3505 3571 +66
=================================================
+ Hits 2762 2823 +61
- Misses 743 748 +5 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every kernel launch, copy, and fill used to create a fresh
ZeCommandList, submit it to the task's command queue, and drop the reference — leaving destruction of thousands of driver objects (lists, command buffers, heaps) to finalizer timing. Under launch storms that garbage is what pushes the driver into allocation failure, where NEO's error handling is at its worst (see #609). It is also pure overhead: ~8x in per-dispatch submission latency (measured with a microbenchmark on PVC; formal numbers will be attached before undrafting).This PR replaces the per-dispatch machinery with a per-task
oneStreamholding one in-order asynchronous immediate command list: appends submit directly, and the per-dispatch garbage class disappears entirely. Measured on Aurora (PVC, LTS): a pressure loop that accumulates 10,003 live command lists in 2,000 iterations on current main accumulates zero with this PR, and a reproducer that reliably dies at the driver's allocation wall on main (6/6 runs, OOM inretry_reclaim) survives 6/6 with the fix, including the first spilling submission.Design notes:
sycl_queuehost-synchronizes the immediate list before handing out the SYCL queue (Julia → MKL), and a dirty flag makes the next Julia-side submission drain the companion queue (MKL → Julia; one Bool load on the fast path). FFT plans capture their queue at construction, so their_exec!methods apply the boundary themselves. New interleave tests (broadcast → gemm → broadcast with no intermediate synchronization, 100 iterations, vs CPU reference) cover this.zeDriverGetApiVersion: the Aurora LTS driver reports API 1.6 while fully implementing in-order immediate lists (they are DPC++'s production submission path on PVC). For the same reason the code avoids 1.9-only loader entrypoints such aszeCommandListIsImmediate. Drivers that genuinely reject the creation get a clear error.ONEAPI_SYNC_EACH_SUBMISSIONnow host-synchronizes the list after each append. The dropped-tail driver bug it works around was only ever observed on the queue-submission path; whether it exists at all on immediate lists is left for a future oversubscription re-test (the knob is kept).KA.priority!swaps the task's stream for one with the requested priority. Explicit@oneapi queue=...still submits through a per-dispatch list, and the Guard NEO's scratch allocation with a spill-triggered drain #609 hedge covers both paths (per-stream and per-queue high-water marks).Full test suite on Aurora LTS (PVC,
ONEAPI_LTS=1, 12 workers): 10,900 pass; the single failure is a pre-existing host-side AVX512-FP16 Float16 issue unrelated to this change (fails identically on main). Also green: KernelAbstractions testsuite (2,218), the new level-zero immediate-list block, FFT interleave, and a two-task alloc/free churn test exercising the cross-task drain.🤖 Generated with Claude Code
https://claude.ai/code/session_01WSFxSBXtckG3BVAT12wYkf