Repository navigation
Recommendation for a good coding local LLM for Mac Mini M4Pro 48GB? #3705
Replies: 5 comments 16 replies
|
I recommend Qwen3.8-27B-oQ4e-mtp or Qwen3.8-Flash-Next oQ4e-mtp |
|
I have the same machine and just tested qwen3.8:27b-mlx. Results were crap on evaluating a very simple 70 line JavaScript Lambda handler focused on image resizing. Even on extra high reasoning it was very inaccurate and took over 2.5 minutes. About to try qwen3-coder:30b but will likely give up on the local llm dream with this machine if it doesn't work well. |
|
I run a 64 GB m5 pro and have for 5 months with omlx surpassing 1Bil input tokens for my daily coding assistant in my work. It is totally doable, particularly with the pi agent and 8 bit turboquant caching. I have also taken a hard look at how these models mess up (usually tool call syntax) and intercept and fix what I can via a pi extension, which definitely took the experience up several notches. I would rather not use qwen, but I must admit 27b (4 bit mlx) is a league above Gemma 4 31B (4bit mlx) for coding and definitely a league and a half above Gemma 4 26b, and with the splash engine (hopefully being ported to omlx?!) it's far more ram friendly than dflash for twice the performance. I tried it as a test on my 36 gb work machine and it will run out to at least a 75k context. I saw you mentioned Sonnet 5, I noticed it is not better than Owen 27b on what I've tried them both on, just faster. I do have to hand hold any of these mentioned models, primarily by giving one (usually a dense model) a detailed plan skill that tells it to write a set of task based instructions and asking me questions about any known-unknowns, potential bugs, etc in my idea. I proofread the instructions, sometimes make tweaks, then open a new pi chat and tell it or an MoE like Gemma 4 26b (must be 4bit QAT or 8bit mlx or oQ8e) to execute it the plan. Usually it does pretty well, but sometimes yes, they really do dumb things. So does Sonnet 5, and even Opus 5 occasionally fwiw if I don't give them a robust enough plan. |
|
Just wanted to share that I am now dropping Ornith 1.5 35B A3B 4bit. It seemed to be doing well and then all of a sudden it started hallucinating big time. And concerning it was hallucinating super early and took me a while to catch that it was saying it was running things that were failing, only that was completely untrue as with me running the same thing in terminal it was totally fine. Not sure if anyone else's experience with Ornith is the same. I'm moving to something else after wasting two days with it on a coding problem. |
|
Caveat: My experience is also relatively limited, but I've played around with these things for the past half year. I'm more in the LLM-assisted "corner" than totally hands-off; I don't trust these parrots enough for that... ;-) On a 64GB machine (M3 Max and M4 Max), I've generally stuck to 6-bit oQe quants; they're a big jump up from 4-bit for longer context work. I've so far been quite happy with KAT-Coder-2.5 ( But my main driver — even though noticeably slower (the prompt processing is more painful part) — has become Qwen-3.8 27B (since prefix caching is essential, this requires patching the chat template to allow in-place system messages with Claude). It does a lot of back-and-forth thinking (even on medium, which is my default), but the results it comes back with are genuinely impressive. My current problem remains memory spikes during long context PP in oMLX; splash seems much better in this regard for now (from what I understand, the 27B should require model size + 16GB for KV cache, but oMLX allocations can exceed that by a big margin — maybe it's temporaries from the prefix caching or something else). |
Uh oh!
There was an error while loading. Please reload this page.
Hi all, I have a Mac Mini M4Pro 48gb and have been searching for a decent coding LLM that I can run locally. So far, that search has not yielded fantastic results. I have landed on Qwen3 32B 4bit but even that failed to deliver good results on some tasks. Gemma4 31B 4bit also wasn't bad but wasn't great either. At least both would run for a while before barfing or hallucinating on whether it was actually doing something or just saying it did something.
I am not searching for perfection given my setup, i'd love to hear what others might recommend.
At the moment I am trying to avoid handholding a model from beginning to end as much as possible. The last project seemed simple - take a Mercator projection map image and draw some region defined boxes (defined by lat/long coordinnates) on top and I literally had to lead it step by step through how to do it, and it could not figure it out on its own, including not being able to read an image and determine key aspects of it for this task's success. Note that even Claude Sonnet 5 failed on this (wasting $ on tokens to produce non-working garbage repeatedly), which makes me wonder that maybe such a task requires a bit more handholding. I did not try more advanced coding models that are not local.
Is handholding inevitable with general project instructions? Seems like some tasks are doable but others are not.
In any case, i'd love to hear your favorite local coding LLM, and a bit on how you use it, like whether it is a coding companion or are you giving it high level direction and letting it go.
Thank you in advance!
All reactions