Should Models Be Structured to Fit Under Frame Budgets when running ML work on display GPU? (DirectML / OpenVINO) #32556
Replies: 2 comments
|
This is an interesting use case. I believe the distinction here is between inference latency, asynchronous execution, and GPU scheduling/preemption. 1. Is this a runtime bug or expected behavior?A 16.67 ms frame budget does not necessarily mean every inference operation must complete within 16.67 ms. However, if inference is executed synchronously on the rendering thread, or if the rendering pipeline waits for GPU work to complete, a 30–50 ms inference can certainly cause visible stuttering. Even when inference runs asynchronously, both workloads may still compete for GPU execution resources, memory bandwidth, and scheduling time. DirectML provides GPU acceleration through DirectX 12, but that does not automatically guarantee that inference workloads will be preempted at arbitrary points to satisfy a GUI frame deadline. The exact scheduling and preemption behavior depends on the GPU, driver, operating system, and execution configuration. 2. Should models be structured to fit within the frame budget?Not necessarily. I would distinguish between two requirements:
For example, a model taking 40 ms could still be useful in a 60 FPS application if inference runs independently and the rendering thread never blocks waiting for its result. However, if the application requires fresh inference results every frame, then the model's latency and throughput become much more important. 3. Approaches worth investigatingA. Decouple inference from rendering Run inference on a dedicated worker thread or asynchronous pipeline. The rendering loop should consume the latest available result rather than synchronously waiting for inference completion. For OpenVINO, asynchronous inference requests ( Documentation: https://docs.openvino.ai/2024/openvino-workflow/running-inference/integrate-openvino-with-your-application/inference-request.html B. Avoid processing stale frames If inference takes longer than one frame, consider dropping outdated input frames instead of building an inference queue. For real-time desktop applications, processing the newest available frame is often more useful than processing every frame sequentially. C. Profile GPU execution separately from CPU execution Measure:
This can help determine whether the stuttering originates from CPU blocking, GPU contention, synchronization, or another bottleneck. D. Consider workload isolation If the machine has multiple GPUs, test running inference on a different adapter from the display GPU. For DirectML, the execution provider supports selecting a device through its This is not always possible on laptops or integrated-GPU systems, but it can help establish whether shared GPU resources are the primary cause. 4. One important distinction regarding OpenVINOOpenVINO's asynchronous inference API allows the host application to continue executing while inference is in progress. However, asynchronous execution alone does not guarantee GPU preemption or a particular scheduling priority for rendering workloads. Its GPU execution behavior also depends on the underlying device and driver. My recommendation: Before restructuring the model to guarantee that every operation completes within 16 ms, I would first test whether moving inference completely off the rendering thread and avoiding synchronization in the frame loop eliminates the stuttering. If the GPU itself remains saturated and rendering misses deadlines, then model optimization, reduced inference frequency, workload isolation, or a different execution strategy may be necessary. The key question is whether your application requires inference results every frame, or whether it can tolerate inference running asynchronously and updating the UI whenever a new result becomes available. If this answer helped or pointed you in the right direction, I'd appreciate it if you could mark it as the accepted answer so it's easier for others with the same issue to find. Also, if you found my contribution useful, I'd appreciate it if you could check out my GitHub profile, follow me, and star any repositories you find interesting. GitHub: https://github.com/Advait251206 |
|
Not a runtime bug. Windows can only preempt GPU work at the granularity the driver reports (IDXGIAdapter2::GetDesc2 → ComputePreemptionGranularity). On most consumer GPUs that's dispatch boundary at best, so once a 30 ms conv or matmul dispatch has started, your render queue waits for it no matter what priority it has. ORT doesn't add anything on top of that. What actually helps, in the order I'd try:
So yes, if you stay on the display GPU, keeping individual ops under the frame time is the real requirement. |
Uh oh!
There was an error while loading. Please reload this page.
Hi everyone,
We are developing a real-time desktop application targeting a strict 60 FPS render loop (16.67ms frame budget). The app runs ONNX Runtime models alongside our GUI rendering pipeline on the primary display GPU using DirectML and OpenVINO.
When running inference some of the operations take a long time 30 to 50ms and the GUI stutters, is this a bug in the runtime that stops the operations being preempted by GUI work or is it expected that we need to only run operations that take under the frame budget of 16ms. I can't find any conformation only if this is expected or not, thank you.
All reactions