Replies: 5 comments 10 replies
|
The context window will be too small when you want to run reasonable local LLM models. You need at least 64-128GB of VRAM or unified memory 🤷♂️ |
|
my setup has rx 9070xt + rtx 2070 super total vram 23~ GB and 32 ram I'm running https://huggingface.co/DavidAU/Qwen3.6-27B-Heretic-Uncensored-FINETUNE-NEO-CODE-Di-IMatrix-MAX-GGUF at Q4KM doing 131k context and it's glorious. |
|
You can easily run a super powerful model like Bartowski's Q6 quant of Gemma 4 31b with that kind of VRAM at your disposal (or Q5_K_M if you don't mind it being a bit dumber and less consistent in exchange for a bit more VRAM headroom). Use a temperature of 1, top_p of 0.95 and top_k of 64 and start with like 32k context, then scale as needed or drop down a tier to Q5 if the maximum you can fit is still inadequate. |
|
It won't be an issue at all. All depends on the model's parameter count. I have an RTX 5060 which has 8 GBs of RAM and I can run qwen2.5-coder:7b pretty smoothly. |
|
Lol you people. 16gb vram 64gb ram here. you can totally run any sort of model. just tinker offload if your model is too big. you can download lmstudio for example where you can change each models parameters to suit you. load on gpu as much as possible and put context as much as needed. havin MoE helps alot. then rest you can dial there untill you LLM is stable. gongrats you are now running 32b llm on 16gbVram! |
Uh oh!
There was an error while loading. Please reload this page.
Like can it run on 32gb vram and 64gb ram?
All reactions