Skip to content
Discussion options

You must be logged in to vote

Thanks, those lines settle most of it, and they also show my estimate last time was ~10 GiB short.

The weights match the file to the MiB: Vulkan0 61222.07 MiB is the non-expert tensors plus all routed experts of UD-IQ4_XS, CPU 27465.95 MiB is the n-gram table (28,800,138,240 bytes), and Vulkan_Host 644.14 MiB is the token embedding.

The name of the CPU buffer is the key: it's CPU, not CPU_Mapped. The 26.8 GiB table has been copied into anonymous RAM, which the kernel can only swap out, not drop. Add the 80.4 GiB of GTT you measured and llama-server alone holds ~107 GiB of 122, which is why the rest of the system ends up in swap.

I think I found why mmap is off. llama.cpp now defaults to -…

Replies: 5 comments 2 replies

Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
1 reply
@159753a52
Comment options

Comment options

You must be logged in to vote
1 reply
@159753a52
Comment options

Answer selected by konst-sh
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants