Repository navigation
vLLM Connector
Let a vLLM server reload prompt memory from disk: after a restart, on one GPU or across many.
| Status | ✅ Works on vLLM. Output identical, byte exact across a restart, on one GPU and on models split over several |
| Verified | 29 September 2026: 30 of 30 models on vLLM 0.30, 4 × A10G, 1, 2 and 4 GPUs per model (list below) · Qwen3.5 9B answers after a load token-identical to Galahad off, 5 of 5 runs |
| Needs | vLLM, the Galahad Python package, and a licence (Install) |
When a vLLM server restarts (a deploy, a crash, a node being replaced), every bit of prompt memory it had built up is gone. The first users after a restart wait for work the server had already done minutes earlier.
This plugs into vLLM's own connector interface and writes that memory to disk. After a restart it is read back, so the server begins warm instead of empty.
Your output does not change. The token ids vLLM generates are identical with the connector attached and with it removed. That was measured, not assumed.
A support assistant restarts at 3 a.m. for a routine deploy. Without this, the first fifty customers next morning each wait while the server re-reads the same handbook. With it, the server already knows it.
⭐ Output identity. vLLM's generated token ids are identical with and without the connector, on an A40 with Qwen3-8B.
⭐ Byte exact across a real restart. Memory saved from the GPU, then a full process restart, then loaded back to the GPU: 100% byte identical.
The full steps are in Install. In short:
export GALAHAD_CACHE_DIR=/var/galahad # the store, on fast local disk
export GALAHAD_KEY_FILE=/var/galahad/store.key # or your key manager
export GALAHAD_LICENCE_FILE=/etc/galahad/licence.lic
vllm serve Qwen/Qwen3-8B \
--kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both","kv_load_failure_policy":"recompute"}'vLLM finds the connector by itself: the package registers a vLLM plugin.
⚠ Short prompts are not saved. Below GALAHAD_MIN_TOKENS (default 256)
re-reading is cheaper than loading, so the connector does not bother.
⚠ Write these settings into your unit file or launch script, not only your shell. Settings that live only in the shell you started the server from are gone at the next restart, and the failure looks unrelated.
Split the model the way you normally would (--tensor-parallel-size 4, or
pipeline parallel). Nothing else changes.
⭐ A prompt counts as remembered only when every GPU saved its part. If one GPU's part is missing, it is a miss on all of them, and the model reads the prompt again.
Measured 29 September 2026: Qwen3 14B split over 2 GPUs, 5 of 5 fresh starts saved on both GPUs and loaded back on both.
With "kv_load_failure_policy":"recompute", a block that cannot be read (a
damaged file, a disk error) is refused, vLLM works that prompt out again, and
the user gets the right answer. Galahad then does not offer that block again, on
any GPU.
⚠ vLLM's default is fail: the request returns an error instead. Set
recompute unless you want failures to be loud.
⭐ 30 of 30 work, 29 September 2026, vLLM 0.30, 4 × NVIDIA A10G (24 GB each). "Works" means three things, every time: the model answered, Galahad saved its memory to disk, and after vLLM's own GPU cache was emptied Galahad loaded it back.
| Family | Models | GPUs |
|---|---|---|
| Qwen3 | 8B · 14B · 32B · Coder 30B-A3B | 1 · 2 · 4 · 4 |
| Qwen3.5 / 3.6 / 3.8 (hybrid) | 3.5 9B · 3.6 35B-A3B · 3.8 27B | 2 · 4 · 4 |
| Gemma 4 | 12B · 26B-A4B · 31B | 2 · 4 · 4 |
| Mistral | Ministral 3 8B · Ministral 3 14B · Mistral Small 3.2 24B · Devstral Small 2 24B | 1 · 2 · 4 · 4 |
| GLM | 4.6V-Flash 9B · 4.7-Flash 30B-A3B | 2 · 4 |
| Llama | 3.1 8B Instruct · 3.3 70B Instruct | 1 · 4 |
| DeepSeek-R1 distills | Qwen 7B · Llama 8B · Qwen 14B · Qwen 32B · Llama 70B · R1-0528 Qwen3 8B | 1 · 1 · 2 · 4 · 4 · 1 |
| Phi-4 | 14B · Reasoning 14B | 2 · 2 |
| GPT-oss | 20B | 2 |
| Nemotron 3 | Nano 30B-A3B | 4 |
| Olmo 3 | 7B · 32B | 1 · 4 |
⚠ Five ran as a different build of the same model, because of the test GPUs, not Galahad:
| Model | Ran as | Why |
|---|---|---|
| Ministral 3 8B / 14B | Mistral's own BF16 build | the listed build is FP8; the A10G cannot run FP8 |
| Devstral Small 2 24B | a BF16 build | the listed build is FP8, and Mistral publishes no BF16 build |
| Llama 3.3 70B · R1 Distill Llama 70B | FP8 builds | the full-precision weights do not fit in 96 GB |
⭐ Hybrid models (Qwen3.5 / 3.6 / 3.8) mix normal attention layers with layers that keep a fixed-size running summary instead of a per-word memory. Galahad saves both together, and restores both or neither. Measured on Qwen3.5 9B: after a load, the answer is token-identical to the answer with Galahad switched off, in 5 of 5 runs.
⚠ A model not on this list most likely works. If Galahad cannot handle a
model, it says so at startup (hybrid KV plan unavailable or NOT CACHING)
and vLLM serves normally, without memory. It never guesses.
If many of your requests begin the same way (a system prompt, a policy document, a tool schema), they can share that beginning instead of each computing it:
export GALAHAD_PAYLOAD_ORDER=chunk # default: layer| What it does | makes every 64 tokens a sharing point, instead of one coarse step |
| Measured | a ~700-token shared preamble: 3 of 3 follow-up requests reused it. A ~359-token preamble, shorter than one coarse step: 2 of 4 reused it |
| What it costs | sharing uses extra storage |
| When to leave it off | requests that do not share a beginning. With nothing shared, the extra storage is the only effect |
⚠ Set it once, before first use. Changing it on a running deployment starts
a fresh cache. The old entries are not lost or misread, simply not looked for.
A value other than layer or chunk is refused with an error, not ignored.
See also Prefix Sharing.
The connector reports its counters to your server's log, on a line
beginning KV Transfer metrics:
grep 'KV Transfer metrics' /path/to/your/server.log | tail -1... hits_disk=3, loads=3, saver_chain_saved=1, chain_hits=3, hit_rate=25.0
| Counter | What it tells you |
|---|---|
chain_hits |
requests that reused a shared beginning |
saver_chain_saved |
requests stored so their beginning can be shared |
saver_chain_fallback |
stored the ordinary way instead (see below) |
hits_disk · hit_rate
|
reuse of any kind |
⚠ chain_hits at zero while hits_disk climbs means reuse is happening the
ordinary way, not through shared beginnings. Usually GALAHAD_PAYLOAD_ORDER is
unset, or your requests genuinely do not share one.
⚠ saver_chain_fallback climbing means the setting is on but a save could
not use it, so those requests store the ordinary way. Nothing is lost; the
sharing simply does not apply to them.
⚠ Read the counters from a line printed after the work. They are per-report figures, not running totals, so an old line still shows old numbers.
Models such as Gemma 4 mix sliding-window and full-attention layers. vLLM keeps them apart with its hybrid KV manager, so a sliding layer holds only its window.
⭐ Galahad works with that manager on, so a sliding layer keeps only its window in GPU memory, and Galahad saves only that window. Measured on Gemma 4 31B, RTX A6000, 26 September 2026:
| with the hybrid manager on | |
|---|---|
| KV room in the GPU memory | 60,484 tokens, 2.3× what the same memory holds with the manager off |
| one saved 9,472-token reading | 1.6 GB: each window only, not the whole sequence (8.5 GB) |
With Galahad attached, vLLM logs hybrid KV manager ON for every Gemma 4 and
Olmo 3 model (29 September 2026).
⚠ If the log says hybrid-attention model + no SupportsHMA, the installed
package is not the current one. Install the current package.
⭐ The work is not done. That is the product. A reused prompt costs no prefill at all: no GPU cycles, no energy, no queue time, no capacity consumed.
| GPU work avoided | the prefill never runs. This holds on every machine |
| capacity freed | those GPU-seconds serve other requests |
| time to first token | the user waits for a load, not a computation |
| throughput | depends on your hardware, see below |
Measured on an H100 with Qwen3-30B-A3B:
| turns/s | |
|---|---|
| vLLM without Galahad | 2.643 |
| ⭐ with Galahad | 3.424 |
| with Galahad, 1 in 3 loads deliberately failing | 2.716 |
⭐ The last row is why it is safe to run. With a third of loads failing on purpose: no deadlock, no errors seen by clients, and still above vLLM without Galahad.
⭐ A slower GPU gains more: 2.91× on an RTX PRO 4500 against 1.32× on an H100 SXM. A fast GPU prefills quickly, so avoiding a prefill is worth less. The H100 figures are the floor, not the ceiling.
⚠ Network storage is about 17 times slower than local: 44 ms versus 780 ms for a 580 MB block. Galahad warns at startup if the store is not on a local disk, but does not refuse: 780 ms still beats recomputing.
⚠ A RAM disk does not help. Moving the store from disk to RAM did not raise throughput. A fast local disk is enough.
A TensorRT-LLM connector is not yet available.
| Symptom | Cause | Fix |
|---|---|---|
no [galahad] lines at startup |
the package is not installed in vLLM's environment |
pip install it into the same environment as vLLM |
caching ENABLED never appears |
this model is not supported yet | the startup line says why; vLLM serves normally |
| nothing saved | prompts shorter than GALAHAD_MIN_TOKENS
|
expected; lower it only if loading beats re-reading on your hardware |
| nothing saved, log says no licence | the licence is missing or for fewer points | see Install |
| no hits after a restart | the store moved |
GALAHAD_CACHE_DIR must point at the same store as before |
HTTP 500 once, KV load failure in the log |
a block on disk is damaged, policy is fail
|
set "kv_load_failure_policy":"recompute"
|
| much slower than expected | the store is on network storage | the startup warning names it; use local disk |
| throughput unchanged | async load switched off | leave GALAHAD_ASYNC_LOAD at its default (on) |
- Install · API Reference
- Prefix Sharing: for fleets with a shared preamble
- SGLang Backend: the same idea for SGLang
- Performance