|
I have downloaded this unsloth ggufs right after the release - before the mtp files were published and successfully run it in unsloth desktop, then the mtp arrived and I have just downloaded the mtp-Qwen3.8-Flash-Next-shared-Q8_0, put in ./snapshots//MTP/ and was able to run the model again just fine. |
Replies: 5 comments 2 replies
|
Some memory arithmetic that may explain why the bigger MTP file tips it over. This assumes the 96 GB is a fixed BIOS carve-out, so the OS and everything CPU-side live in the remaining ~32 GB. From the tensor sizes of UD-IQ4_XS (87.2 GiB across the 3 shards):
So the GPU side is ~60 GiB + the MTP model (~3.8 GiB for a 4.1 GB file) + KV + compute buffers, roughly 70 GiB even at 256K. That fits in 96 with room to spare. The CPU side is the tight one: 27.4 GiB for the table and embedding out of ~32 GB leaves ~4 GB for the OS, the desktop app and its backend. If the table ends up resident instead of being paged from disk, the extra GB or so from the non-shared MTP (presumably it carries its own copies of tensors the Two things worth checking (estimates, I haven't run this on a Strix Halo):
|
|
@159753a52 thanks, I have somehow partially resolved this by switching to 512mb bios vram carve-out and gtt specifcation in grub. Though I feel that i'm still having some memory over-consumption compared to earlier unsloth versions where it was just running up to almost 200k context I'm now having OOM at ~170k + 16gb swap is 100% busy. |
|
@159753a52 /$ cat /sys/class/drm/card*/device/mem_info_gtt_used per_layer_token_embd.weight - there is no such line in the log at that verbosity level btw this is what free-h shows me currently: and these are my gtt grub options: amd_iommu=pt ttm.pages_limit=29360128 ttm.page_pool_size=15728640 - its not too agressive, but from my undestanding the model should still fit without squeezing whole 16 gb to swap |
|
@159753a52 thanks I should note, yesterday in conversation with that qwen 3.8 itself it proposed me to change the load mode to mmap as well, and it seem to resolve memory consumption: -lm mmap --lazy-mode off - with this I have numbers as following at context of 200k /$ free -h Not sure, though if it will impose any performance downsides. Thanks for your help! |
|
The non-shared Q8_0 MTP draft weights inflate base VRAM allocation by an extra ~4-6 GB during KV cache initialization, leaving zero headroom when prompt submission triggers tensor memory reservation. If you cannot downgrade the MTP file through the UI, set We ran into this exact weight-bloat OOM issue and built a quick browser calculator to check exact peak VRAM footprints before loading models: https://huggingface.co/spaces/vivacious-cloud/vivacious-terminal-simulator |
Thanks, those lines settle most of it, and they also show my estimate last time was ~10 GiB short.
The weights match the file to the MiB:
Vulkan0 61222.07 MiBis the non-expert tensors plus all routed experts of UD-IQ4_XS,CPU 27465.95 MiBis the n-gram table (28,800,138,240 bytes), andVulkan_Host 644.14 MiBis the token embedding.The name of the CPU buffer is the key: it's
CPU, notCPU_Mapped. The 26.8 GiB table has been copied into anonymous RAM, which the kernel can only swap out, not drop. Add the 80.4 GiB of GTT you measured and llama-server alone holds ~107 GiB of 122, which is why the rest of the system ends up in swap.I think I found why mmap is off. llama.cpp now defaults to
-…