Repository navigation
Replies: 1 comment
|
Some baseline numbers that show why the window matters on 16 GB. Qwen3.8-27B keeps KV in only 16 of its 64 layers, so it's 64 KiB per token at f16 and ~34 KiB at q8_0:
Question: does the MTP head read from the KVMem window, or does it keep its own small cache? That shifts the budget a bit. (I built a free calculator for these quant/context numbers, if anyone wants to try other combos: [Link removed by moderator]/llm-vram-calculator/qwen3.8-27b/) |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
We designed and implemented KVMem for Qwen3.8-27B: a new way for llama.cpp to manage conversation memory. The basic idea is simple. A 16GB card cannot keep a full 256K KV cache in VRAM, so we do not try. Finished history lives in regular RAM; for each new question the GPU only loads a limited window of the pieces that matter. We checked this at 256K on LongMemEval-S and AgentLongBench: a 32K GPU window is essentially as accurate as keeping the full 256K history (85.6% vs 86.6% accuracy, 60.9% vs 59.5% task success).
For local use, that means a consumer 16GB GPU can keep a long agent chat, tools, and files in one session — without buying a huge card, and without watching token speed collapse as the conversation grows. On an RTX 5060 Ti 16GB, Qwen3.8-27B runs at a full 256K workspace and decode stays around 30–40 tok/s.
Repo: https://github.com/kvmem/kvmem-llama.cpp (
v0.14.0)Paper: https://arxiv.org/abs/2609.04852
It is an OpenAI-compatible llama.cpp server, so existing local clients can point at it. Chat, tools, and vision are supported.
Setup for the numbers below
The conversation can still be 256K tokens long. Only that GPU window sits in VRAM; the rest of the KV stays in RAM and comes back when the current question needs it.
256K tool task (33 requests, ending at 262058 / 262144 tokens):
A shorter image + code task is about 37–41 tok/s decode. GPU vision takes ~0.3s; CPU vision ~11s. Decode includes thinking tokens.
IQ3 (MTP, GPU vision, Q8 KV) is the one we recommend on 16GB. IQ4 (MTP, CPU vision, Q5 KV) if you mostly do text.
How to run it is in the repo: https://github.com/kvmem/kvmem-llama.cpp
Welcome to try it. If you hit anything — build, VRAM, quality, other 16GB cards — discuss it in this thread. I will reply as soon as I see it.
Related: KV streaming
Adaptive KV-cache streaming also runs this same native 256K context on 16GB (they write 262K with K=1000; we write 256K as 256×1024=262144), by moving the full cache through VRAM as needed. It still looks at every token, so quality is exact; speed still drops as the chat gets longer (their 5070 Ti figures: ~50 tok/s at 8K, ~10 at full context). We keep a smaller window on the GPU instead, which is why the 5060 Ti numbers above stay around 30–40 tok/s at 256K.
All reactions