Repository navigation
Replies: 2 comments 1 reply
|
I would not generally expect With There is an important implementation detail here, though: For example, the mmap path directly returns pointers into the mapped file: llama.cpp/src/llama-model-loader.cpp Lines 1471 to 1490 in bd4f514 On Linux, llama.cpp also uses Lines 470 to 515 in bd4f514 Conversely, the non-mmap path can stage data through pinned host buffers and perform asynchronous uploads to the GPU. So especially with GPU offloading, a performance difference during loading or the first inference may come from the different I/O/upload paths rather than from page faults alone. For steady-state prefill, though, I would expect the difference to mostly disappear once the relevant working set is resident. If mmap is still ~30% slower after several warmup runs, with no meaningful major-page-fault activity, that would be interesting and would suggest that paging alone is probably not the explanation. I would benchmark these separately:
and test both cold and warm page caches. It would also be useful to know the exact model, llama.cpp commit, backend, |
|
Hi @jintakhan, Two things can keep mmap slower after warmup, and both are about whatever part of the model still lives in host memory, not the part in VRAM.
If you're fully offloaded to a discrete GPU, the only thing left on the CPU is the token embedding table (llama-model.cpp: "there is very little benefit to offloading the input layer, so always keep it on the CPU"), and the lookup for a prompt is tiny, so mmap shouldn't be able to cost 30% there. In that case I'd double check that |
Uh oh!
There was an error while loading. Please reload this page.
In almost every instance I've come across, mmap kneecaps prefill performance by at least 30% on both unified memory and non-UMA systems due to the added paging overhead. And from my understanding, while lazy loading helps with model loading times, it can also cause unexpected OOMs. Yet it remains the default load mode for llama.cpp. Is there a situation where mmap actually meaningfully helps with inference performance? Genuinely curious. Thanks.
All reactions