Repository navigation
How to stop Qwen 3.8 Flash Next from consuming ALL the system ram? #28588
|
I'm on windows 11 running the latest llama.cpp [b10850]. I have 16GB vram and 64GB system ram. I can run Qwen 3.8 Flash Next, Qwen3.8-Flash-Next-UD-Q4_K_XL.gguf, but I can't find any way to limit how much system ram it uses. It always saturates the entire system ram, leaving nothing for any additional apps I might want to run (for instance, if I'm using qwen to code an app and need to run it while we work on it.) I've asked Claude and others to help but they can't work it out. I'd want to keep at least 8-16 GB of system ram free for other things. We've tried a bunch of settings, my latest launch .bat file is below, but there have been a lot of experiments involving mmap and stuff. Nothing has worked. |
Replies: 3 comments
|
Quick hint to take back to Claude — no follow-up support provided: Your RAM saturation is caused by memory duplication and OS pagefile thrashing between Have Claude refactor your launch arguments with these constraints:
|
|
Get more RAM. You need at least 128GB. |
|
Both the above answers were helpful. I needed to use a smaller quant, and I needed to use the new settings. n_cpu_moe 36 seems the best option, fastest tokens per second and still leaves me plenty of system ram. |
Both the above answers were helpful. I needed to use a smaller quant, and I needed to use the new settings.
New settings on their own didn't help, the Q4_K_XL seems to be just too big, one of the AIs suggested it's to do with the size of the lookup table. So I went with a much smaller quant IQ2_M plus new settings, and this seemed to fix the issue. I did some tests, I wouldn't trust they are consistent from run to run, since one example shows different results for two different runs.