Post your Arc results #1
Replies: 3 comments 6 replies
SYCL Flash Attention on Arc 140T (Xe-LPG): oneDNN SDPA returns wrong results for quantized KV with permuted viewsSummaryOn Intel Arc 140T (Xe-LPG+, Arrow Lake-H), the SYCL backend's oneDNN SDPA path returns numerically incorrect results when the KV cache is quantized, the layout is non-contiguous ( Separately, MTP + ngram-mod speculative decoding on SYCL leaks device memory and crashes with Hardware / Software
|
| oneDNN version | ERR |
|---|---|
| oneAPI bundled (3.11) | 1.26 |
| 3.13.3 (source build) | 1.33 |
| 3.13.3 | 1.28 |
| 3.13.3 | 1.48 |
| 3.13.3 | 1.18 |
| 3.13.3 | 1.23 |
Upgrading oneDNN from 3.11 to 3.13.3 did not fix the bug. Error stays in the 1.2–1.5 range, roughly 2500× the 0.0005 tolerance.
llama-bench Results
Scope: Ternary-Bonsai 2 27B PTQ1_0 only. These numbers apply to this specific model on this specific fork. They should not be extrapolated to other models or other builds.
| KV | ONEDNN | MKL | pp512 | tg128 |
|---|---|---|---|---|
| f16 | 0 | 0 | 167.39 ± 0.89 | 7.90 ± 0.03 |
| q8_0 | 0 | 0 | 170.46 ± 1.33 | 7.85 ± 0.03 |
| q8_0 | 1 | 1 | 174.03 ± 0.06 | 7.88 ± 0.01 |
| f16 | 1 | 1 | 212.10 ± 0.84 | 7.93 ± 0.04 |
The F16 + oneDNN row is the fastest correct configuration for this model. F16 removes the quantized-KV precondition, so the broken kernel path is not taken.
The q8_0 + oneDNN row is fast but produces wrong attention output per the test-backend-ops failure above. It should not be used for correctness-sensitive work.
Server Findings
MTP memory leak
llama-server with --spec-type draft-mtp,ngram-mod crashes at ggml-sycl.cpp:2853 with UR_RESULT_ERROR_OUT_OF_RESOURCES on the first task launch. This persists with GGML_SYCL_ENABLE_VMM=0. It is separate from the FA bug.
VMM
GGML_SYCL_ENABLE_VMM=0 is required to avoid an earlier OUT_OF_RESOURCES failure at graph reserve. This aligns with the VMM disable-by-default behavior in PR ggml-org#24673 and PR ggml-org#27689.
ngram-mod behavior
ngram-mod only fires when the model echoes a long contiguous block from the prompt. In one session:
ngram-mod: #calls(b,g,a) = 3 356 2
#gen drafts = 2
#acc drafts = 2
#gen tokens = 512
#acc tokens = 301
mean acc len = 151.50
Hit rate 2/356 calls. When it fires, it carries 300+ tokens in 2 drafts. When it does not, MTP alone carries generation at ~4–7 t/s.
ngram-mod window size
| n_max | Behavior |
|---|---|
| 256 | Catches long verbatim copies; mean accepted length up to 36 |
| 48 | Catches short targeted edits; mean accepted length up to 25, acceptance up to 89% |
Neither wins universally. A middle value (128) may capture both patterns.
Conclusion
- The oneDNN SDPA path is numerically broken on Xe-LPG for quantized KV +
permute=[0,2,1,3]+kv_view=1. Matches issue SYCL: flash attention returns a silently wrong answer for a quantised, non-contiguous KV view (Arc 140T, Xe-LPG) ggml-org/llama.cpp#27769, not fixed by oneDNN 3.13.3. - F16 KV avoids the bug and is the fastest correct configuration for this model on this fork.
- MTP + ngram-mod on SYCL has a separate memory leak causing
OUT_OF_RESOURCES. GGML_SYCL_ENABLE_VMM=0is required to avoid an earlier allocation failure.
Generated with DeepSeek <3
|
System info: Context size:
In most local models I have tried, this task end up getting brute forced and often is in an infinite loop. Result from the Terminal:
|
|
Thanks for testing, and for including the full log. 55 t/s sustained across almost 6,000 generated tokens on Windows is a solid result.
The "4 tokens" isn't wrong, it's prompt caching. Your prompt was 381 tokens and 377 of them were already in the server's cache from the previous turn (f_sim_best = 1.000), so only 4 new ones needed processing. That conversation used about 6,350 tokens of context (n_tokens = 6353). The limit is set by -c, which is 131,072 (128K) in run-bonsai.bat, and the server prints it at startup as n_ctx.
A couple of questions so I can read the numbers properly: did you use run-bonsai.bat as is, or your own command? The .bat starts with thinking turned off, so the model does its working in the answer itself. And did it get the timer puzzle right (it's 9 minutes: start both, flip the 7 when it runs out at 7, then flip it again when the 4 runs out at 8, and it finishes at 9)?
…Sent from my Galaxy
-------- Original message --------
From: PotatoNut ***@***.***>
Date: 30/9/26 10:05 pm (GMT+10:00)
To: "Torchit1/llama.cpp" ***@***.***>
Cc: Jesse Symons ***@***.***>, Author ***@***.***>
Subject: Re: [Torchit1/llama.cpp] Post your Arc results (Discussion #1)
System info:
Asrock B580 12 GB - Graphics Driver 32.0.101.9030
CPU: Ryzen 7 7700
RAM 32 GB 6000 mt/s cl 30
OS Windows 11
Context size:
Unknown, the prompt incorrectly marks this text as 4 tokens:
You have two sand timers, which can show 4 minutes and 7 minutes respectively. Use both the sand timers(at a time or one after other or any other combination) and measure a time of 9 minutes.
In most local models I have tried, this task end up getting brute forced and often is in an infinite loop.
Result from the Terminal:
9.32.264.104 I slot release: id 0 | task 391 | stop processing: n_tokens = 6385, truncated = 0
9.32.264.147 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 1.000 (> 0.100 thold), f_keep = 0.059
9.32.383.713 I slot launch_slot_: id 0 | task 637 | processing task, is_child = 0
9.35.515.056 I slot print_timing: id 0 | task 637 | n_gen = 168, tg = 55.00 t/s, tg_3s = 55.33 t/s
9.38.545.317 I slot print_timing: id 0 | task 637 | n_gen = 332, tg = 54.56 t/s, tg_3s = 54.12 t/s
9.41.549.918 I slot print_timing: id 0 | task 637 | n_gen = 518, tg = 56.99 t/s, tg_3s = 61.91 t/s
9.44.593.857 I slot print_timing: id 0 | task 637 | n_gen = 691, tg = 56.95 t/s, tg_3s = 56.83 t/s
9.47.642.063 I slot print_timing: id 0 | task 637 | n_gen = 862, tg = 56.78 t/s, tg_3s = 56.10 t/s
9.50.655.480 I slot print_timing: id 0 | task 637 | n_gen = 1029, tg = 56.56 t/s, tg_3s = 55.42 t/s
9.53.671.737 I slot print_timing: id 0 | task 637 | n_gen = 1202, tg = 56.67 t/s, tg_3s = 57.36 t/s
9.56.720.684 I slot print_timing: id 0 | task 637 | n_gen = 1350, tg = 55.65 t/s, tg_3s = 48.54 t/s
9.59.748.650 I slot print_timing: id 0 | task 637 | n_gen = 1515, tg = 55.52 t/s, tg_3s = 54.49 t/s
10.02.756.850 I slot print_timing: id 0 | task 637 | n_gen = 1680, tg = 55.45 t/s, tg_3s = 54.85 t/s
10.05.764.248 I slot print_timing: id 0 | task 637 | n_gen = 1866, tg = 56.03 t/s, tg_3s = 61.85 t/s
10.09.312.509 I slot print_timing: id 0 | task 637 | n_gen = 2056, tg = 55.79 t/s, tg_3s = 53.55 t/s
10.12.340.111 I slot print_timing: id 0 | task 637 | n_gen = 2231, tg = 55.94 t/s, tg_3s = 57.80 t/s
10.15.891.599 I slot print_timing: id 0 | task 637 | n_gen = 2419, tg = 55.70 t/s, tg_3s = 52.94 t/s
10.18.932.694 I slot print_timing: id 0 | task 637 | n_gen = 2581, tg = 55.54 t/s, tg_3s = 53.27 t/s
10.21.969.807 I slot print_timing: id 0 | task 637 | n_gen = 2752, tg = 55.59 t/s, tg_3s = 56.30 t/s
10.25.025.671 I slot print_timing: id 0 | task 637 | n_gen = 2917, tg = 55.49 t/s, tg_3s = 53.99 t/s
10.28.509.225 I slot print_timing: id 0 | task 637 | n_gen = 3102, tg = 55.35 t/s, tg_3s = 53.11 t/s
10.31.516.060 I slot print_timing: id 0 | task 637 | n_gen = 3289, tg = 55.69 t/s, tg_3s = 62.19 t/s
10.34.572.797 I slot print_timing: id 0 | task 637 | n_gen = 3469, tg = 55.85 t/s, tg_3s = 58.89 t/s
10.37.605.463 I slot print_timing: id 0 | task 637 | n_gen = 3622, tg = 55.60 t/s, tg_3s = 50.45 t/s
10.40.637.741 I slot print_timing: id 0 | task 637 | n_gen = 3759, tg = 55.14 t/s, tg_3s = 45.18 t/s
10.43.685.465 I slot print_timing: id 0 | task 637 | n_gen = 3922, tg = 55.07 t/s, tg_3s = 53.48 t/s
10.46.719.495 I slot print_timing: id 0 | task 637 | n_gen = 4110, tg = 55.35 t/s, tg_3s = 61.96 t/s
10.49.769.676 I slot print_timing: id 0 | task 637 | n_gen = 4303, tg = 55.66 t/s, tg_3s = 63.27 t/s
10.52.773.426 I slot print_timing: id 0 | task 637 | n_gen = 4465, tg = 55.60 t/s, tg_3s = 53.93 t/s
10.55.787.927 I slot print_timing: id 0 | task 637 | n_gen = 4614, tg = 55.37 t/s, tg_3s = 49.43 t/s
10.58.802.819 I slot print_timing: id 0 | task 637 | n_gen = 4750, tg = 55.01 t/s, tg_3s = 45.11 t/s
11.01.829.267 I slot print_timing: id 0 | task 637 | n_gen = 4931, tg = 55.18 t/s, tg_3s = 59.81 t/s
11.04.860.209 I slot print_timing: id 0 | task 637 | n_gen = 5098, tg = 55.17 t/s, tg_3s = 55.10 t/s
11.08.293.106 I slot print_timing: id 0 | task 637 | n_gen = 5270, tg = 54.99 t/s, tg_3s = 50.10 t/s
11.11.318.003 I slot print_timing: id 0 | task 637 | n_gen = 5396, tg = 54.58 t/s, tg_3s = 41.65 t/s
11.14.371.772 I slot print_timing: id 0 | task 637 | n_gen = 5567, tg = 54.63 t/s, tg_3s = 56.00 t/s
11.17.401.228 I slot print_timing: id 0 | task 637 | n_gen = 5699, tg = 54.31 t/s, tg_3s = 43.57 t/s
11.20.433.590 I slot print_timing: id 0 | task 637 | n_gen = 5868, tg = 54.35 t/s, tg_3s = 55.73 t/s
11.21.715.172 I slot print_timing: id 0 | task 637 | prompt eval time = 95.04 ms / 4 tokens ( 23.76 ms per token, 42.09 tokens per second)
11.21.715.178 I slot print_timing: id 0 | task 637 | eval time = 109236.30 ms / 5972 tokens ( 18.29 ms per token, 54.66 tokens per second)
11.21.715.179 I slot print_timing: id 0 | task 637 | total time = 109331.34 ms / 5976 tokens
11.21.715.179 I slot print_timing: id 0 | task 637 | graphs reused = 3952
11.21.715.267 I slot print_timing: id 0 | task 637 | draft acceptance = 0.37763 ( 4220 accepted / 11175 generated), mean len = 3.43
11.21.715.364 I slot release: id 0 | task 637 | stop processing: n_tokens = 6353, truncated = 0
18.41.181.604 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.867 (> 0.100 thold), f_keep = 0.059
18.41.291.246 I slot launch_slot_: id 0 | task 4353 | processing task, is_child = 0
18.44.649.408 I slot print_timing: id 0 | task 4353 | n_gen = 146, tg = 47.94 t/s, tg_3s = 48.27 t/s
—
Reply to this email directly, view it on GitHub<#1?email_source=notifications&email_token=BFCOGHKSHQACRT73CBGAEGL5RTZIXA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBWG44DOMZRUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVRTG633UMVZF6Y3MNFRWW#discussioncomment-18678731>, or unsubscribe<https://github.com/notifications/unsubscribe-auth/BFCOGHL7GV24JKPK5D6FH435RTZIXAVCNFSNUABJKJSXA33TNF2G64TZHMYTGOBZGAZTEMRQGM5UI2LTMN2XG43JN5XDWMJQHA4TGMRTGWQXMAQ>.
You are receiving this because you authored the thread.Message ID: ***@***.***>
|
Uh oh!
There was an error while loading. Please reload this page.
Running the arc-b580 branch? Post your card, driver, OS, context size and the numbers you get (llama-bench tg128/pp512, or the server's t/s). Results from other setups help most.
All reactions