M4 + RTX 5080 data to share. Prefill-side cost on Metal, and a q4_0 failure autopsy. Happy to run your diagnostic script. #98
ryanjmichie-git
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi @TheTom I maintain turboquant-101 ([repo link]), a beginner-facing lab for measuring KV-cache compression. Your fork is documented there as the advanced Apple Silicon option. Sharing two findings from paired NIAH runs on an M4 Mac mini (mainline llama.cpp quantized KV) and an RTX 5080 (vLLM turboquant), in case they're useful:
On CUDA the cost of quantized KV lands on decode; on Metal it lands on prefill in my runs. Does that match what you saw building the fork's Metal path?
Our llama.cpp q4_0 failures weren't lost memory. Replaying the originally failing prompts (52 replays) produced the exact needle every time — routed into reasoning_content with an empty reply. Instruction-following degraded; retrieval didn't. Autopsy here
I'd be happy to run your diagnostic script on both my machines (mac mini m4 & 5080) and post results in #20969 (link) or here. The M4 (16GB) and a 5080 might be useful datapoints for you.
All reactions