Pre-submission Checklist
GPU Hardware
Intel Arc a770
DRI Devices Information
ls -ls /dev/dri/*
0 crw-rw----+ 1 root video 226, 0 24. Jul 15:50 /dev/dri/card0
0 crw-rw----+ 1 root video 226, 1 24. Jul 15:54 /dev/dri/card1
0 crw-rw-rw- 1 root video 226, 128 24. Jul 15:50 /dev/dri/renderD128
0 crw-rw-rw- 1 root video 226, 129 24. Jul 15:54 /dev/dri/renderD129
/dev/dri/by-path:
insgesamt 0
0 lrwxrwxrwx 1 root root 8 24. Jul 15:54 pci-0000:03:00.0-card -> ../card1
0 lrwxrwxrwx 1 root root 13 24. Jul 15:54 pci-0000:03:00.0-render -> ../renderD129
0 lrwxrwxrwx 1 root root 8 24. Jul 15:50 pci-0000:0a:00.0-card -> ../card0
0 lrwxrwxrwx 1 root root 13 24. Jul 15:50 pci-0000:0a:00.0-render -> ../renderD128
ls -la /dev/dri/by-path
insgesamt 0
drwxr-xr-x 2 root root 120 24. Jul 15:54 ./
drwxr-xr-x 3 root root 140 24. Jul 15:54 ../
lrwxrwxrwx 1 root root 8 24. Jul 15:54 pci-0000:03:00.0-card -> ../card1
lrwxrwxrwx 1 root root 13 24. Jul 15:54 pci-0000:03:00.0-render -> ../renderD129
lrwxrwxrwx 1 root root 8 24. Jul 15:50 pci-0000:0a:00.0-card -> ../card0
lrwxrwxrwx 1 root root 13 24. Jul 15:50 pci-0000:0a:00.0-render -> ../renderD128
GPU Detailed Information (lspci output)
03:00.0 VGA compatible controller: Intel Corporation DG2 [Arc A770] (rev 08) (prog-if 00 [VGA controller])
Subsystem: Sparkle Computer Co., Ltd. Device 3937
Control: I/O+ Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx-
Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- SERR- <PERR- INTx-
Latency: 0, Cache Line Size: 64 bytes
Interrupts: unknown pin routed to IRQ 97, MSI(X) routed to IRQ 97
IOMMU group: 12
Region 0: Memory at fb000000 (64-bit, non-prefetchable) [size=16M]
Region 2: Memory at 7800000000 (64-bit, prefetchable) [size=16G]
Expansion ROM at fc000000 [disabled] [size=2M]
Capabilities: [40] Vendor Specific Information: Intel Capabilities v1
CapA: Peg60Dis- Peg12Dis- Peg11Dis- Peg10Dis- PeLWUDis- DmiWidth=x4
EccDis- ForceEccEn- VTdDis- DmiG2Dis- PegG2Dis- DDRMaxSize=Unlimited
1NDis- CDDis- DDPCDis- X2APICEn- PDCDis- IGDis- CDID=0 CRID=0
DDROCCAP- OCEn- DDRWrtVrefEn+ DDR3LEn+
CapB: ImguDis- OCbySSKUCap- OCbySSKUEn- SMTCap- CacheSzCap 0x0
SoftBinCap- DDR3MaxFreqWithRef100=Disabled PegG3Dis-
PkgTyp- AddGfxEn- AddGfxCap- PegX16Dis- DmiG3Dis- GmmDis-
DDR3MaxFreq=2932MHz LPDDR3En-
Capabilities: [70] Express (v2) Endpoint, IntMsgNum 0
DevCap: MaxPayload 128 bytes, PhantFunc 0, Latency L0s <64ns, L1 <1us
ExtTag+ AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 0W TEE-IO-
DevCtl: CorrErr- NonFatalErr- FatalErr- UnsupReq-
RlxdOrd+ ExtTag+ PhantFunc- AuxPwr- NoSnoop+ FLReset-
MaxPayload 128 bytes, MaxReadReq 128 bytes
DevSta: CorrErr- NonFatalErr- FatalErr- UnsupReq- AuxPwr- TransPend-
LnkCap: Port #0, Speed 2.5GT/s, Width x1, ASPM L0s L1, Exit Latency L0s <64ns, L1 <1us
ClockPM- Surprise- LLActRep- BwNot- ASPMOptComp+
LnkCtl: ASPM L1 Enabled; RCB 64 bytes, LnkDisable- CommClk-
ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- FltModeDis-
LnkSta: Speed 2.5GT/s, Width x1
TrErr- Train- SlotClk- DLActive- BWMgmt- ABWMgmt-
DevCap2: Completion Timeout: Range B, TimeoutDis+ NROPrPrP- LTR+
10BitTagComp+ 10BitTagReq+ OBFF Not Supported, ExtFmt+ EETLPPrefix-
EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit-
FRS- TPHComp- ExtTPHComp-
AtomicOpsCap: 32bit- 64bit- 128bitCAS-
DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis-
AtomicOpsCtl: ReqEn-
IDOReq- IDOCompl- LTR- EmergencyPowerReductionReq-
10BitTagReq- OBFF Disabled, EETLPPrefixBlk-
LnkCap2: Supported Link Speeds: 2.5GT/s, Crosslink- Retimer- 2Retimers- DRS-
LnkCtl2: Target Link Speed: 2.5GT/s, EnterCompliance- SpeedDis-
Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS-
Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot
LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete- EqualizationPhase1-
EqualizationPhase2- EqualizationPhase3- LinkEqualizationRequest-
Retimer- 2Retimers- CrosslinkRes: unsupported, FltMode-
Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable+ 64bit+
Address: 00000000fee00000 Data: 0000
Masking: 00000000 Pending: 00000000
Capabilities: [d0] Power Management version 3
Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0+,D1-,D2-,D3hot+,D3cold-)
Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME-
Capabilities: [100 v1] Alternative Routing-ID Interpretation (ARI)
ARICap: MFVC- ACS-, Next Function: 0
ARICtl: MFVC- ACS-, Function Group: 0
Capabilities: [420 v1] Physical Resizable BAR
BAR 2: current size: 16GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB 16GB
Capabilities: [400 v1] Latency Tolerance Reporting
Max snoop latency: 0ns
Max no snoop latency: 0ns
Kernel driver in use: xe
Kernel modules: i915, xe
Driver Version
26.27.39122.11
Installed GPU Driver Packages
dpkg --list | grep -iE "igc|gmm|opencl|level-zero|fc|ocloc|libze"
ii clinfo 3.0.25.02.14-1 amd64 Query OpenCL system information
ii intel-igc-cm 1.0.176.54074-102924.04 amd64 Intel(R) C for Metal Compiler -- CM Frontend lib
ii intel-igc-core 1.0.17193.4 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-core-2 2.38.2 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-opencl 1.0.17193.4 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-opencl-2 2.38.2 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-level-zero-gpu-raytracing 1.0.0-71u24.04 amd64 oneAPI Level Zero Ray Tracing Support
ii intel-ocloc 26.27.39122.11-0 amd64 Tool for managing Intel Compute GPU device binary format
ii intel-ocloc-dbgsym 26.27.39122.11-0 amd64 debug symbols for intel-ocloc
rc intel-oneapi-runtime-dpcpp-sycl-opencl-cpu-2024 2024.2.1-1079 amd64 Intel® CPU Runtime for OpenCL(TM) Applications runtime
rc intel-oneapi-runtime-openmp-opencl-shared 2025.3.3-30 amd64 Intel(R) OpenMP and OpenCL shared files for runtime package
rc intel-oneapi-runtime-openmp-opencl-shared-2024 2024.2.1-1079 amd64 Intel(R) OpenMP and OpenCL shared files for runtime package
ii intel-opencl-icd 26.27.39122.11-0 amd64 Intel graphics compute runtime for OpenCL
ii intel-opencl-icd-dbgsym 26.27.39122.11-0 amd64 debug symbols for intel-opencl-icd
ii libclc-19 1:19.1.7-22 all OpenCL C language implementation - platform support
ii libclc-19-dev 1:19.1.7-22 all OpenCL C language implementation - development files
ii libemail-date-format-perl 1.008-1 all Module to generate RFC-2822-valid date strings
ii libigc1 1.0.17791.16-103224.04 amd64 Intel graphics compiler for OpenCL -- core libs
ii libigc2:amd64 2.28.4-4 amd64 Intel graphics compiler for OpenCL -- core libs
ii libigdfcl2:amd64 2.28.4-4 amd64 Intel graphics compiler for OpenCL -- OpenCL library
ii libigdgmm-dev:amd64 22.10.0+ds1-1 amd64 Intel Graphics Memory Management Library -- development files
ii libigdgmm12:amd64 22.10.0 amd64 Intel Graphics Memory Management Library -- shared library
ii libze-dev:amd64 1.28.2-2 amd64 oneAPI Level Zero -- development files
ii libze-intel-gpu1 26.27.39122.11-0 amd64 Intel(R) Graphics Compute Runtime for oneAPI Level Zero.
ii libze-intel-gpu1-dbgsym 26.27.39122.11-0 amd64 debug symbols for libze-intel-gpu1
ii libze1:amd64 1.28.2-2 amd64 oneAPI Level Zero -- share libraries
ii libzen0t64:amd64 0.4.41-4+b1 amd64 ZenLib C++ utility library -- runtime
rc mesa-opencl-icd:amd64 24.1.3-2 amd64 free implementation of the OpenCL API -- ICD runtime
ii ocl-icd-libopencl1:amd64 2.3.4-1+b1 amd64 Generic OpenCL ICD Loader
ii ocl-icd-libopencl1:i386 2.3.4-1+b1 i386 Generic OpenCL ICD Loader
ii ocl-icd-opencl-dev:amd64 2.3.4-1+b1 amd64 OpenCL development files
ii opencl-c-headers 3.02025.07.22-2 all OpenCL (Open Computing Language) C header files
ii opencl-clhpp-headers 3.0~2025.07.22-1 all C++ headers for OpenCL development
conda list | grep -Ei "intel|oneapi|dpcpp|mkl|sycl|level|xpu"
dpcpp-cpp-rt 2026.0.0 pypi_0 pypi
intel-cmplr-lib-rt 2026.0.0 pypi_0 pypi
intel-cmplr-lib-ur 2026.0.0 pypi_0 pypi
intel-cmplr-lic-rt 2026.0.0 pypi_0 pypi
intel-opencl-rt 2026.0.0 pypi_0 pypi
intel-openmp 2026.0.0 pypi_0 pypi
intel-pti 0.17.0 pypi_0 pypi
intel-sycl-rt 2026.0.0 pypi_0 pypi
mkl 2026.0.0 pypi_0 pypi
mkl-dpcpp 2025.0.1 pypi_0 pypi
oneccl-bind-pt 2.7.0+xpu pypi_0 pypi
onemkl-license 2026.0.0 pypi_0 pypi
onemkl-sycl-blas 2026.0.0 pypi_0 pypi
onemkl-sycl-datafitting 2025.0.1 pypi_0 pypi
onemkl-sycl-dft 2026.0.0 pypi_0 pypi
onemkl-sycl-lapack 2026.0.0 pypi_0 pypi
onemkl-sycl-rng 2026.0.0 pypi_0 pypi
onemkl-sycl-sparse 2026.0.0 pypi_0 pypi
onemkl-sycl-stats 2025.0.1 pypi_0 pypi
onemkl-sycl-vm 2025.0.1 pypi_0 pypi
pytorch-triton-xpu 3.5.0 pypi_0 pypi
torch 2.13.0+xpu pypi_0 pypi
torchaudio 2.11.0+xpu pypi_0 pypi
torchvision 0.28.0+xpu pypi_0 pypi
triton-xpu 3.7.2 pypi_0 pypi
ldd $(python -c "import torch; print(torch._C.file)") | grep -Ei "sycl|mkl|dnnl|ze|level|igc"
libsycl.so.9 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libsycl.so.9 (0x00007fc0dfe00000)
libtorch-xpu-ops-sycltla.so => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/libtorch-xpu-ops-sycltla.so (0x00007fc0d4800000)
libmkl_sycl_blas.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_blas.so.6 (0x00007fc0d0a00000)
libmkl_sycl_dft.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_dft.so.6 (0x00007fc0ce600000)
libmkl_sycl_lapack.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_lapack.so.6 (0x00007fc0cca00000)
libmkl_intel_lp64.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_intel_lp64.so.3 (0x00007fc0cb800000)
libmkl_core.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_core.so.3 (0x00007fc0c7000000)
libmkl_gnu_thread.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_gnu_thread.so.3 (0x00007fc0c5400000)
Driver Installation Details
initially followed https://dgpu-docs.intel.com/installation-guides/installing-packages-from-the-intel-ppa.html
https://repositories.intel.com/gpu/ubuntu
use https://github.com/intel/compute-runtime/releases when pytorch needed newer releases
Linux Distribution
Other (please specify below)
Other Linux Distribution
Devuan testing (based on Debian without systemd)
Kernel Version & Boot Parameters
uname -r
7.1.3+deb14-amd64
cat /proc/cmdline
BOOT_IMAGE=/vmlinuz-7.1.3+deb14-amd64 root=/dev/mapper/crypt-root ro quiet i915.force_probe=!56a0 xe.force_probe=56a0 i915.enable_hangcheck=1 xe.enable_hangcheck=1 net.ifnames=0 biosdevname=0
lsmod | grep -E 'i915|xe'
xe 4571136 4
intel_vsec 28672 1 xe
drm_gpuvm 57344 1 xe
drm_gpusvm_helper 61440 1 xe
configfs 65536 2 xe
gpu_sched 69632 2 amdgpu,xe
drm_ttm_helper 20480 3 amdgpu,xe
drm_buddy 12288 2 amdgpu,xe
drm_exec 12288 3 drm_gpuvm,amdgpu,xe
ttm 159744 3 amdgpu,drm_ttm_helper,xe
drm_suballoc_helper 24576 2 amdgpu,xe
drm_display_helper 311296 2 amdgpu,xe
cec 81920 3 drm_display_helper,amdgpu,xe
drm_client_lib 16384 2 amdgpu,xe
drm_kms_helper 266240 5 drm_display_helper,amdgpu,drm_ttm_helper,drm_client_lib,xe
drm 925696 27 gpu_sched,drm_panel_backlight_quirks,drm_kms_helper,drm_exec,drm_gpuvm,drm_suballoc_helper,drm_display_helper,drm_buddy,amdgpu,drm_gpusvm_helper,drm_ttm_helper,drm_client_lib,xe,ttm,amdxcp
i2c_algo_bit 16384 2 amdgpu,xe
video 81920 2 amdgpu,xe
Actual Behavior
AI Summary:
With Intel Compute Runtime 26.14.37833.4 and newer (IGC 2.32.x+), PyTorch 2.13.0+xpu fails on Intel Arc A770 (DG2) when creating oneDNN GPU primitives.
The initial failure occurs in torch.nn.functional.scaled_dot_product_attention(), which throws:
RuntimeError: could not create a primitive
with oneDNN reporting:
CL_OUT_OF_HOST_MEMORY
This is misleading, as GPU memory usage is low and the failure occurs even with very small tensors.
When running ComfyUI, this initially appears as std::future_error: Broken promise during the text encoder. After patching ComfyUI to force SDPBackend.MATH for SDPA (which successfully works around the attention failure), execution proceeds further but then fails in torch.nn.functional.linear() through the comfy_kitchen FP8 dispatch path with the same could not create a primitive error.
This indicates the issue is not limited to SDPA but affects multiple oneDNN GPU primitive types. The Python environment, PyTorch version, and application remained unchanged during testing. Only the Intel Compute Runtime/IGC packages were changed. Rolling back to Compute Runtime 26.09.37435.1 (IGC 2.30.1) resolves all failures and PyTorch 2.13.0+xpu runs normally.
Expected Behavior
AI Summary:
PyTorch should successfully create and execute oneDNN GPU primitives on Intel Arc A770 (DG2) without errors. torch.nn.functional.scaled_dot_product_attention(), torch.nn.functional.linear(), and downstream applications such as ComfyUI should execute normally, as they do when using Intel Compute Runtime 26.09.37435.1 (IGC 2.30.1). No RuntimeError: could not create a primitive or CL_OUT_OF_HOST_MEMORY errors should occur for valid workloads.
Reproduction Rate
Always reproduces - 100%
Steps to Reproduce
Install >= 26.14.37833.4
start comfy
run workflow
could not create a primitive
Is this a regression?
Last Known Working Driver Version
26.09.37435.1
First Known Failing Driver Version
26.14.37833.4
API Call Logs
No response
strace Logs
No response
System Logs / dmesg Output
No response
Backtrace (if crash or hang occurred)
No response
Source Code / Reproducer
AI:
Minimal PyTorch reproducer
import torch
print(torch.version)
q = torch.randn((1, 16, 16, 128), device="xpu")
k = torch.randn((1, 16, 16, 128), device="xpu")
v = torch.randn((1, 16, 16, 128), device="xpu")
torch.nn.functional.scaled_dot_product_attention(q, k, v)
Expected:
Returns a tensor.
Actual:
RuntimeError: could not create a primitive
Additional reproducer
The issue is also reproducible in ComfyUI using PyTorch 2.13.0+xpu on Intel Arc A770 (DG2).
Initially the failure occurs during the text encoder with std::future_error: Broken promise, which originates from torch.nn.functional.scaled_dot_product_attention().
After patching ComfyUI to force SDPBackend.MATH for SDPA, the attention error disappears, but execution later fails in torch.nn.functional.linear() (via comfy_kitchen FP8 torch_dispatch) with the same RuntimeError: could not create a primitive.
This demonstrates that the issue affects multiple oneDNN GPU primitives and is not limited to SDPA.
Command Line / Application Details
No response
oneAPI Version (if applicable)
No response
Screenshots / Video
No response
Additional Notes
AI Summary:
Additional testing performed:
- The issue was reproduced across multiple Intel Compute Runtime releases by changing only the system runtime packages while keeping the Python environment, PyTorch (2.13.0+xpu), ComfyUI, and workload unchanged.
- Intel Compute Runtime 26.09.37435.1 (IGC 2.30.1) works correctly. Intel Compute Runtime 26.14.37833.4 and all newer releases tested reproduce the failure.
- The reported
CL_OUT_OF_HOST_MEMORY error does not appear to reflect actual GPU memory exhaustion. The failure occurs with very small tensors (e.g. sequence length 16), and forcing SDPBackend.MATH allows the same tensors to execute successfully.
ONEDNN_GRAPH=0 did not change the behavior.
- The failure is not limited to
scaled_dot_product_attention(). After working around the SDPA failure, torch.nn.functional.linear() also fails with RuntimeError: could not create a primitive, suggesting a broader issue in oneDNN GPU primitive creation or compilation on DG2 with newer Intel Compute Runtime / IGC versions.
- The issue has been reproduced on Devuan Testing using an Intel Arc A770 (DG2).
Pre-submission Checklist
GPU Hardware
Intel Arc a770
DRI Devices Information
ls -ls /dev/dri/*
0 crw-rw----+ 1 root video 226, 0 24. Jul 15:50 /dev/dri/card0
0 crw-rw----+ 1 root video 226, 1 24. Jul 15:54 /dev/dri/card1
0 crw-rw-rw- 1 root video 226, 128 24. Jul 15:50 /dev/dri/renderD128
0 crw-rw-rw- 1 root video 226, 129 24. Jul 15:54 /dev/dri/renderD129
/dev/dri/by-path:
insgesamt 0
0 lrwxrwxrwx 1 root root 8 24. Jul 15:54 pci-0000:03:00.0-card -> ../card1
0 lrwxrwxrwx 1 root root 13 24. Jul 15:54 pci-0000:03:00.0-render -> ../renderD129
0 lrwxrwxrwx 1 root root 8 24. Jul 15:50 pci-0000:0a:00.0-card -> ../card0
0 lrwxrwxrwx 1 root root 13 24. Jul 15:50 pci-0000:0a:00.0-render -> ../renderD128
ls -la /dev/dri/by-path
insgesamt 0
drwxr-xr-x 2 root root 120 24. Jul 15:54 ./
drwxr-xr-x 3 root root 140 24. Jul 15:54 ../
lrwxrwxrwx 1 root root 8 24. Jul 15:54 pci-0000:03:00.0-card -> ../card1
lrwxrwxrwx 1 root root 13 24. Jul 15:54 pci-0000:03:00.0-render -> ../renderD129
lrwxrwxrwx 1 root root 8 24. Jul 15:50 pci-0000:0a:00.0-card -> ../card0
lrwxrwxrwx 1 root root 13 24. Jul 15:50 pci-0000:0a:00.0-render -> ../renderD128
GPU Detailed Information (lspci output)
03:00.0 VGA compatible controller: Intel Corporation DG2 [Arc A770] (rev 08) (prog-if 00 [VGA controller])
Subsystem: Sparkle Computer Co., Ltd. Device 3937
Control: I/O+ Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx-
Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- SERR- <PERR- INTx-
Latency: 0, Cache Line Size: 64 bytes
Interrupts: unknown pin routed to IRQ 97, MSI(X) routed to IRQ 97
IOMMU group: 12
Region 0: Memory at fb000000 (64-bit, non-prefetchable) [size=16M]
Region 2: Memory at 7800000000 (64-bit, prefetchable) [size=16G]
Expansion ROM at fc000000 [disabled] [size=2M]
Capabilities: [40] Vendor Specific Information: Intel Capabilities v1
CapA: Peg60Dis- Peg12Dis- Peg11Dis- Peg10Dis- PeLWUDis- DmiWidth=x4
EccDis- ForceEccEn- VTdDis- DmiG2Dis- PegG2Dis- DDRMaxSize=Unlimited
1NDis- CDDis- DDPCDis- X2APICEn- PDCDis- IGDis- CDID=0 CRID=0
DDROCCAP- OCEn- DDRWrtVrefEn+ DDR3LEn+
CapB: ImguDis- OCbySSKUCap- OCbySSKUEn- SMTCap- CacheSzCap 0x0
SoftBinCap- DDR3MaxFreqWithRef100=Disabled PegG3Dis-
PkgTyp- AddGfxEn- AddGfxCap- PegX16Dis- DmiG3Dis- GmmDis-
DDR3MaxFreq=2932MHz LPDDR3En-
Capabilities: [70] Express (v2) Endpoint, IntMsgNum 0
DevCap: MaxPayload 128 bytes, PhantFunc 0, Latency L0s <64ns, L1 <1us
ExtTag+ AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 0W TEE-IO-
DevCtl: CorrErr- NonFatalErr- FatalErr- UnsupReq-
RlxdOrd+ ExtTag+ PhantFunc- AuxPwr- NoSnoop+ FLReset-
MaxPayload 128 bytes, MaxReadReq 128 bytes
DevSta: CorrErr- NonFatalErr- FatalErr- UnsupReq- AuxPwr- TransPend-
LnkCap: Port #0, Speed 2.5GT/s, Width x1, ASPM L0s L1, Exit Latency L0s <64ns, L1 <1us
ClockPM- Surprise- LLActRep- BwNot- ASPMOptComp+
LnkCtl: ASPM L1 Enabled; RCB 64 bytes, LnkDisable- CommClk-
ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- FltModeDis-
LnkSta: Speed 2.5GT/s, Width x1
TrErr- Train- SlotClk- DLActive- BWMgmt- ABWMgmt-
DevCap2: Completion Timeout: Range B, TimeoutDis+ NROPrPrP- LTR+
10BitTagComp+ 10BitTagReq+ OBFF Not Supported, ExtFmt+ EETLPPrefix-
EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit-
FRS- TPHComp- ExtTPHComp-
AtomicOpsCap: 32bit- 64bit- 128bitCAS-
DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis-
AtomicOpsCtl: ReqEn-
IDOReq- IDOCompl- LTR- EmergencyPowerReductionReq-
10BitTagReq- OBFF Disabled, EETLPPrefixBlk-
LnkCap2: Supported Link Speeds: 2.5GT/s, Crosslink- Retimer- 2Retimers- DRS-
LnkCtl2: Target Link Speed: 2.5GT/s, EnterCompliance- SpeedDis-
Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS-
Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot
LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete- EqualizationPhase1-
EqualizationPhase2- EqualizationPhase3- LinkEqualizationRequest-
Retimer- 2Retimers- CrosslinkRes: unsupported, FltMode-
Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable+ 64bit+
Address: 00000000fee00000 Data: 0000
Masking: 00000000 Pending: 00000000
Capabilities: [d0] Power Management version 3
Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0+,D1-,D2-,D3hot+,D3cold-)
Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME-
Capabilities: [100 v1] Alternative Routing-ID Interpretation (ARI)
ARICap: MFVC- ACS-, Next Function: 0
ARICtl: MFVC- ACS-, Function Group: 0
Capabilities: [420 v1] Physical Resizable BAR
BAR 2: current size: 16GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB 16GB
Capabilities: [400 v1] Latency Tolerance Reporting
Max snoop latency: 0ns
Max no snoop latency: 0ns
Kernel driver in use: xe
Kernel modules: i915, xe
Driver Version
26.27.39122.11
Installed GPU Driver Packages
dpkg --list | grep -iE "igc|gmm|opencl|level-zero|fc|ocloc|libze"
ii clinfo 3.0.25.02.14-1 amd64 Query OpenCL system information
ii intel-igc-cm 1.0.176.54074-1029
24.04 amd64 Intel(R) C for Metal Compiler -- CM Frontend libu24.04 amd64 oneAPI Level Zero Ray Tracing Supportii intel-igc-core 1.0.17193.4 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-core-2 2.38.2 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-opencl 1.0.17193.4 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-igc-opencl-2 2.38.2 amd64 Intel(R) Graphics Compiler for OpenCL(TM)
ii intel-level-zero-gpu-raytracing 1.0.0-71
ii intel-ocloc 26.27.39122.11-0 amd64 Tool for managing Intel Compute GPU device binary format
ii intel-ocloc-dbgsym 26.27.39122.11-0 amd64 debug symbols for intel-ocloc
rc intel-oneapi-runtime-dpcpp-sycl-opencl-cpu-2024 2024.2.1-1079 amd64 Intel® CPU Runtime for OpenCL(TM) Applications runtime
rc intel-oneapi-runtime-openmp-opencl-shared 2025.3.3-30 amd64 Intel(R) OpenMP and OpenCL shared files for runtime package
rc intel-oneapi-runtime-openmp-opencl-shared-2024 2024.2.1-1079 amd64 Intel(R) OpenMP and OpenCL shared files for runtime package
ii intel-opencl-icd 26.27.39122.11-0 amd64 Intel graphics compute runtime for OpenCL
ii intel-opencl-icd-dbgsym 26.27.39122.11-0 amd64 debug symbols for intel-opencl-icd
ii libclc-19 1:19.1.7-22 all OpenCL C language implementation - platform support
ii libclc-19-dev 1:19.1.7-22 all OpenCL C language implementation - development files
ii libemail-date-format-perl 1.008-1 all Module to generate RFC-2822-valid date strings
ii libigc1 1.0.17791.16-1032
24.04 amd64 Intel graphics compiler for OpenCL -- core libs2025.07.22-2 all OpenCL (Open Computing Language) C header filesii libigc2:amd64 2.28.4-4 amd64 Intel graphics compiler for OpenCL -- core libs
ii libigdfcl2:amd64 2.28.4-4 amd64 Intel graphics compiler for OpenCL -- OpenCL library
ii libigdgmm-dev:amd64 22.10.0+ds1-1 amd64 Intel Graphics Memory Management Library -- development files
ii libigdgmm12:amd64 22.10.0 amd64 Intel Graphics Memory Management Library -- shared library
ii libze-dev:amd64 1.28.2-2 amd64 oneAPI Level Zero -- development files
ii libze-intel-gpu1 26.27.39122.11-0 amd64 Intel(R) Graphics Compute Runtime for oneAPI Level Zero.
ii libze-intel-gpu1-dbgsym 26.27.39122.11-0 amd64 debug symbols for libze-intel-gpu1
ii libze1:amd64 1.28.2-2 amd64 oneAPI Level Zero -- share libraries
ii libzen0t64:amd64 0.4.41-4+b1 amd64 ZenLib C++ utility library -- runtime
rc mesa-opencl-icd:amd64 24.1.3-2 amd64 free implementation of the OpenCL API -- ICD runtime
ii ocl-icd-libopencl1:amd64 2.3.4-1+b1 amd64 Generic OpenCL ICD Loader
ii ocl-icd-libopencl1:i386 2.3.4-1+b1 i386 Generic OpenCL ICD Loader
ii ocl-icd-opencl-dev:amd64 2.3.4-1+b1 amd64 OpenCL development files
ii opencl-c-headers 3.0
ii opencl-clhpp-headers 3.0~2025.07.22-1 all C++ headers for OpenCL development
conda list | grep -Ei "intel|oneapi|dpcpp|mkl|sycl|level|xpu"
dpcpp-cpp-rt 2026.0.0 pypi_0 pypi
intel-cmplr-lib-rt 2026.0.0 pypi_0 pypi
intel-cmplr-lib-ur 2026.0.0 pypi_0 pypi
intel-cmplr-lic-rt 2026.0.0 pypi_0 pypi
intel-opencl-rt 2026.0.0 pypi_0 pypi
intel-openmp 2026.0.0 pypi_0 pypi
intel-pti 0.17.0 pypi_0 pypi
intel-sycl-rt 2026.0.0 pypi_0 pypi
mkl 2026.0.0 pypi_0 pypi
mkl-dpcpp 2025.0.1 pypi_0 pypi
oneccl-bind-pt 2.7.0+xpu pypi_0 pypi
onemkl-license 2026.0.0 pypi_0 pypi
onemkl-sycl-blas 2026.0.0 pypi_0 pypi
onemkl-sycl-datafitting 2025.0.1 pypi_0 pypi
onemkl-sycl-dft 2026.0.0 pypi_0 pypi
onemkl-sycl-lapack 2026.0.0 pypi_0 pypi
onemkl-sycl-rng 2026.0.0 pypi_0 pypi
onemkl-sycl-sparse 2026.0.0 pypi_0 pypi
onemkl-sycl-stats 2025.0.1 pypi_0 pypi
onemkl-sycl-vm 2025.0.1 pypi_0 pypi
pytorch-triton-xpu 3.5.0 pypi_0 pypi
torch 2.13.0+xpu pypi_0 pypi
torchaudio 2.11.0+xpu pypi_0 pypi
torchvision 0.28.0+xpu pypi_0 pypi
triton-xpu 3.7.2 pypi_0 pypi
ldd $(python -c "import torch; print(torch._C.file)") | grep -Ei "sycl|mkl|dnnl|ze|level|igc"
libsycl.so.9 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libsycl.so.9 (0x00007fc0dfe00000)
libtorch-xpu-ops-sycltla.so => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/libtorch-xpu-ops-sycltla.so (0x00007fc0d4800000)
libmkl_sycl_blas.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_blas.so.6 (0x00007fc0d0a00000)
libmkl_sycl_dft.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_dft.so.6 (0x00007fc0ce600000)
libmkl_sycl_lapack.so.6 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_sycl_lapack.so.6 (0x00007fc0cca00000)
libmkl_intel_lp64.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_intel_lp64.so.3 (0x00007fc0cb800000)
libmkl_core.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_core.so.3 (0x00007fc0c7000000)
libmkl_gnu_thread.so.3 => /home/user/miniconda3/envs/comfy/lib/python3.11/site-packages/torch/lib/../../../../libmkl_gnu_thread.so.3 (0x00007fc0c5400000)
Driver Installation Details
initially followed https://dgpu-docs.intel.com/installation-guides/installing-packages-from-the-intel-ppa.html
https://repositories.intel.com/gpu/ubuntu
use https://github.com/intel/compute-runtime/releases when pytorch needed newer releases
Linux Distribution
Other (please specify below)
Other Linux Distribution
Devuan testing (based on Debian without systemd)
Kernel Version & Boot Parameters
uname -r
7.1.3+deb14-amd64
cat /proc/cmdline
BOOT_IMAGE=/vmlinuz-7.1.3+deb14-amd64 root=/dev/mapper/crypt-root ro quiet i915.force_probe=!56a0 xe.force_probe=56a0 i915.enable_hangcheck=1 xe.enable_hangcheck=1 net.ifnames=0 biosdevname=0
lsmod | grep -E 'i915|xe'
xe 4571136 4
intel_vsec 28672 1 xe
drm_gpuvm 57344 1 xe
drm_gpusvm_helper 61440 1 xe
configfs 65536 2 xe
gpu_sched 69632 2 amdgpu,xe
drm_ttm_helper 20480 3 amdgpu,xe
drm_buddy 12288 2 amdgpu,xe
drm_exec 12288 3 drm_gpuvm,amdgpu,xe
ttm 159744 3 amdgpu,drm_ttm_helper,xe
drm_suballoc_helper 24576 2 amdgpu,xe
drm_display_helper 311296 2 amdgpu,xe
cec 81920 3 drm_display_helper,amdgpu,xe
drm_client_lib 16384 2 amdgpu,xe
drm_kms_helper 266240 5 drm_display_helper,amdgpu,drm_ttm_helper,drm_client_lib,xe
drm 925696 27 gpu_sched,drm_panel_backlight_quirks,drm_kms_helper,drm_exec,drm_gpuvm,drm_suballoc_helper,drm_display_helper,drm_buddy,amdgpu,drm_gpusvm_helper,drm_ttm_helper,drm_client_lib,xe,ttm,amdxcp
i2c_algo_bit 16384 2 amdgpu,xe
video 81920 2 amdgpu,xe
Actual Behavior
AI Summary:
With Intel Compute Runtime 26.14.37833.4 and newer (IGC 2.32.x+), PyTorch 2.13.0+xpu fails on Intel Arc A770 (DG2) when creating oneDNN GPU primitives.
The initial failure occurs in torch.nn.functional.scaled_dot_product_attention(), which throws:
RuntimeError: could not create a primitive
with oneDNN reporting:
CL_OUT_OF_HOST_MEMORY
This is misleading, as GPU memory usage is low and the failure occurs even with very small tensors.
When running ComfyUI, this initially appears as std::future_error: Broken promise during the text encoder. After patching ComfyUI to force SDPBackend.MATH for SDPA (which successfully works around the attention failure), execution proceeds further but then fails in torch.nn.functional.linear() through the comfy_kitchen FP8 dispatch path with the same could not create a primitive error.
This indicates the issue is not limited to SDPA but affects multiple oneDNN GPU primitive types. The Python environment, PyTorch version, and application remained unchanged during testing. Only the Intel Compute Runtime/IGC packages were changed. Rolling back to Compute Runtime 26.09.37435.1 (IGC 2.30.1) resolves all failures and PyTorch 2.13.0+xpu runs normally.
Expected Behavior
AI Summary:
PyTorch should successfully create and execute oneDNN GPU primitives on Intel Arc A770 (DG2) without errors. torch.nn.functional.scaled_dot_product_attention(), torch.nn.functional.linear(), and downstream applications such as ComfyUI should execute normally, as they do when using Intel Compute Runtime 26.09.37435.1 (IGC 2.30.1). No RuntimeError: could not create a primitive or CL_OUT_OF_HOST_MEMORY errors should occur for valid workloads.
Reproduction Rate
Always reproduces - 100%
Steps to Reproduce
Install >= 26.14.37833.4
start comfy
run workflow
could not create a primitive
Is this a regression?
Last Known Working Driver Version
26.09.37435.1
First Known Failing Driver Version
26.14.37833.4
API Call Logs
No response
strace Logs
No response
System Logs / dmesg Output
No response
Backtrace (if crash or hang occurred)
No response
Source Code / Reproducer
AI:
Minimal PyTorch reproducer
import torch
print(torch.version)
q = torch.randn((1, 16, 16, 128), device="xpu")
k = torch.randn((1, 16, 16, 128), device="xpu")
v = torch.randn((1, 16, 16, 128), device="xpu")
torch.nn.functional.scaled_dot_product_attention(q, k, v)
Expected:
Returns a tensor.
Actual:
RuntimeError: could not create a primitive
Additional reproducer
The issue is also reproducible in ComfyUI using PyTorch 2.13.0+xpu on Intel Arc A770 (DG2).
Initially the failure occurs during the text encoder with std::future_error: Broken promise, which originates from torch.nn.functional.scaled_dot_product_attention().
After patching ComfyUI to force SDPBackend.MATH for SDPA, the attention error disappears, but execution later fails in torch.nn.functional.linear() (via comfy_kitchen FP8 torch_dispatch) with the same RuntimeError: could not create a primitive.
This demonstrates that the issue affects multiple oneDNN GPU primitives and is not limited to SDPA.
Command Line / Application Details
No response
oneAPI Version (if applicable)
No response
Screenshots / Video
No response
Additional Notes
AI Summary:
Additional testing performed:
CL_OUT_OF_HOST_MEMORYerror does not appear to reflect actual GPU memory exhaustion. The failure occurs with very small tensors (e.g. sequence length 16), and forcingSDPBackend.MATHallows the same tensors to execute successfully.ONEDNN_GRAPH=0did not change the behavior.scaled_dot_product_attention(). After working around the SDPA failure,torch.nn.functional.linear()also fails withRuntimeError: could not create a primitive, suggesting a broader issue in oneDNN GPU primitive creation or compilation on DG2 with newer Intel Compute Runtime / IGC versions.