Repository navigation
Replies: 3 comments 1 reply
|
I just installed oMLX 0.7.0rc1 build and whoa nelly, thats a speedy meatball! 👍 Anyone reading this, I highly recommend upgrading to 0.7.0rc1 or later. This humble pie tastes pretty good! XD Well done and thank you to everyone who made this happen! 🥇 |
|
@JJOne123 mind sharing the settings you are using in oMLX for this model? |
|
Yeah, I ended up using similar settings. Here are my settings if that helps anyone else: |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi
First this is not meant as a bash against omlx. I love omlx and use it daily and have been happy to help file issues etc to help improve the product where I could.
But I decided to try a different product recently (mlx-serve) after reading a post on Reddit about it.
It didn't work oob with my jundot qwen3.8 FN oQ4e image, so I had to download a slightly different model, but which has a similar size, here: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Anyways, on mlx-serve it was telling me that my prefill was getting up to 1.9k at times on a M5 Max 128gb. TG was about 50-80 per second depending on the work. I wasn't able to see any option on mlx-serve to offload the ngram to SSD as I do on omlx to preserve more memory for context etc. But it's size loaded into memory to be fair was quite on par with the jundot model so I presume it also had ngram SSD offloading.
This is a substantially performance difference between the 2 engines. Mlx-serve did feel quicker on my own harness when running it too. Omlx tends to give me around 12-1400pp and about 55tg on 0.7.0-dev4 build.
Ultimately I've reverted my long term work to omlx though as I use other models and I found the jundot model more stable.
But I won't lie, the performance difference is noticeable and it's always nice to have more speed.
Is there any idea how or why mlx-serve is running faster, or is there a way to port some of the performance improvements to omlx?
Maybe it's all down to the model?
I'll try finding a model what works across both engines for a closer comparison.
All reactions