Repository navigation
Comparing MLX and vllm-mlx Performance with Qwen3-8B-4bit on MacBook Pro M5 Max 128GB #637
louisnguyenxt
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
vllm-mlx continuous batching — real-workload benchmark (Qwen3-8B, M5 Max)
Benchmark of
vllm-mlx serve --continuous-batching --use-paged-cacheagainst a plainin-process
mlx-lmbaseline, on a production-style LLM-judge workload (IELTS essayscoring: each request = 2 sequential chat completions, ~7.4k total tokens per
request, ~3.6k-token prompts, JSON output, temperature 0.05/0.1).
Setup
0dd1157), Python 3.12mlx-community/Qwen3-8B-4bitmlx-lmin-process, single stream, speculative decoding (Qwen3-0.6B-4bit draft,num_draft_tokens=2), requests serialized by a lockBoth sides ran the exact same prompts, sampling params, and
max_tokens; one warmuprequest before timing in each mode. The two systems ran in separate phases (never
concurrently). Swap stayed at 0 throughout; server RSS ≈ 5.4 GB peak.
End-to-end throughput (full 2-call requests)
The concurrency tests were performed at 1, 3, 5, 7, 10, and 15. The results are as follows:
The baseline is flat at ~2.1 req/min regardless of offered concurrency (serialized).
vllm-mlx loses ~15% single-stream (no draft-model speculative decoding), breaks even
at ~2 concurrent, and saturates around 2× baseline throughput at c=5–10. Past
c=10 throughput degrades slightly and per-request latency keeps growing.
Token-level metrics (single judge call, streaming)
Same 15 prompts (~3.6k prompt tokens, ~2k completion tokens each):
* TTFT/prefill at c≥5 benefit from the trie prefix cache: the same prompt set was
reused across levels, so prefill was partially cached. Decode and aggregate numbers
are unaffected by this. At c=15 prefill queueing dominates TTFT (p50 ≈ 9.6 s).
Aggregate decode scales from 82 → 215 tok/s (2.6×) between c=1 and c=15, while
per-request decode drops from 87 → 16 tok/s — the classic continuous-batching trade.
Output quality
JSON-mode reliability was perfect in both systems: 90/90 vllm-mlx requests and 45/45
baseline requests parsed and completed with zero errors. The greedy-ish baseline
(temp 0.05) was bit-stable across runs (15/15 identical scores). vllm-mlx produced
identical final scores on 13–14 of 15 essays per level (mean |Δ| ≤ 0.13 on a 0–9
half-band scale), consistent with sampling/batched-numerics differences.
Takeaways
workload at c=5–10, even though the baseline had speculative decoding and
vllm-mlx currently doesn't (draft-model speculative + batching would compound).
latency grows with no throughput gain.
All reactions