Skip to content

vLLM Connector

Sietse edited this page Sep 29, 2026 · 2 revisions

vLLM Connector

Let a vLLM server reload prompt memory from disk: after a restart, on one GPU or across many.

Status ✅ Works on vLLM. Output identical, byte exact across a restart, on one GPU and on models split over several
Verified 29 September 2026: 30 of 30 models on vLLM 0.30, 4 × A10G, 1, 2 and 4 GPUs per model (list below) · Qwen3.5 9B answers after a load token-identical to Galahad off, 5 of 5 runs
Needs vLLM, the Galahad Python package, and a licence (Install)

In plain words

When a vLLM server restarts (a deploy, a crash, a node being replaced), every bit of prompt memory it had built up is gone. The first users after a restart wait for work the server had already done minutes earlier.

This plugs into vLLM's own connector interface and writes that memory to disk. After a restart it is read back, so the server begins warm instead of empty.

Your output does not change. The token ids vLLM generates are identical with the connector attached and with it removed. That was measured, not assumed.

The everyday example

A support assistant restarts at 3 a.m. for a routine deploy. Without this, the first fifty customers next morning each wait while the server re-reads the same handbook. With it, the server already knows it.


What is proven

⭐ Output identity. vLLM's generated token ids are identical with and without the connector, on an A40 with Qwen3-8B.

⭐ Byte exact across a real restart. Memory saved from the GPU, then a full process restart, then loaded back to the GPU: 100% byte identical.


Using it

The full steps are in Install. In short:

export GALAHAD_CACHE_DIR=/var/galahad          # the store, on fast local disk
export GALAHAD_KEY_FILE=/var/galahad/store.key # or your key manager
export GALAHAD_LICENCE_FILE=/etc/galahad/licence.lic

vllm serve Qwen/Qwen3-8B \
  --kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both","kv_load_failure_policy":"recompute"}'

vLLM finds the connector by itself: the package registers a vLLM plugin.

⚠ Short prompts are not saved. Below GALAHAD_MIN_TOKENS (default 256) re-reading is cheaper than loading, so the connector does not bother.

⚠ Write these settings into your unit file or launch script, not only your shell. Settings that live only in the shell you started the server from are gone at the next restart, and the failure looks unrelated.

Several GPUs

Split the model the way you normally would (--tensor-parallel-size 4, or pipeline parallel). Nothing else changes.

⭐ A prompt counts as remembered only when every GPU saved its part. If one GPU's part is missing, it is a miss on all of them, and the model reads the prompt again.

Measured 29 September 2026: Qwen3 14B split over 2 GPUs, 5 of 5 fresh starts saved on both GPUs and loaded back on both.

A damaged block on disk

With "kv_load_failure_policy":"recompute", a block that cannot be read (a damaged file, a disk error) is refused, vLLM works that prompt out again, and the user gets the right answer. Galahad then does not offer that block again, on any GPU.

⚠ vLLM's default is fail: the request returns an error instead. Set recompute unless you want failures to be loud.


Tested models

⭐ 30 of 30 work, 29 September 2026, vLLM 0.30, 4 × NVIDIA A10G (24 GB each). "Works" means three things, every time: the model answered, Galahad saved its memory to disk, and after vLLM's own GPU cache was emptied Galahad loaded it back.

Family Models GPUs
Qwen3 8B · 14B · 32B · Coder 30B-A3B 1 · 2 · 4 · 4
Qwen3.5 / 3.6 / 3.8 (hybrid) 3.5 9B · 3.6 35B-A3B · 3.8 27B 2 · 4 · 4
Gemma 4 12B · 26B-A4B · 31B 2 · 4 · 4
Mistral Ministral 3 8B · Ministral 3 14B · Mistral Small 3.2 24B · Devstral Small 2 24B 1 · 2 · 4 · 4
GLM 4.6V-Flash 9B · 4.7-Flash 30B-A3B 2 · 4
Llama 3.1 8B Instruct · 3.3 70B Instruct 1 · 4
DeepSeek-R1 distills Qwen 7B · Llama 8B · Qwen 14B · Qwen 32B · Llama 70B · R1-0528 Qwen3 8B 1 · 1 · 2 · 4 · 4 · 1
Phi-4 14B · Reasoning 14B 2 · 2
GPT-oss 20B 2
Nemotron 3 Nano 30B-A3B 4
Olmo 3 7B · 32B 1 · 4

⚠ Five ran as a different build of the same model, because of the test GPUs, not Galahad:

Model Ran as Why
Ministral 3 8B / 14B Mistral's own BF16 build the listed build is FP8; the A10G cannot run FP8
Devstral Small 2 24B a BF16 build the listed build is FP8, and Mistral publishes no BF16 build
Llama 3.3 70B · R1 Distill Llama 70B FP8 builds the full-precision weights do not fit in 96 GB

⭐ Hybrid models (Qwen3.5 / 3.6 / 3.8) mix normal attention layers with layers that keep a fixed-size running summary instead of a per-word memory. Galahad saves both together, and restores both or neither. Measured on Qwen3.5 9B: after a load, the answer is token-identical to the answer with Galahad switched off, in 5 of 5 runs.

⚠ A model not on this list most likely works. If Galahad cannot handle a model, it says so at startup (hybrid KV plan unavailable or NOT CACHING) and vLLM serves normally, without memory. It never guesses.


Sharing a preamble between requests

If many of your requests begin the same way (a system prompt, a policy document, a tool schema), they can share that beginning instead of each computing it:

export GALAHAD_PAYLOAD_ORDER=chunk     # default: layer
What it does makes every 64 tokens a sharing point, instead of one coarse step
Measured a ~700-token shared preamble: 3 of 3 follow-up requests reused it. A ~359-token preamble, shorter than one coarse step: 2 of 4 reused it
What it costs sharing uses extra storage
When to leave it off requests that do not share a beginning. With nothing shared, the extra storage is the only effect

⚠ Set it once, before first use. Changing it on a running deployment starts a fresh cache. The old entries are not lost or misread, simply not looked for. A value other than layer or chunk is refused with an error, not ignored.

See also Prefix Sharing.

Checking it is working

The connector reports its counters to your server's log, on a line beginning KV Transfer metrics:

grep 'KV Transfer metrics' /path/to/your/server.log | tail -1
... hits_disk=3, loads=3, saver_chain_saved=1, chain_hits=3, hit_rate=25.0
Counter What it tells you
chain_hits requests that reused a shared beginning
saver_chain_saved requests stored so their beginning can be shared
saver_chain_fallback stored the ordinary way instead (see below)
hits_disk · hit_rate reuse of any kind

⚠ chain_hits at zero while hits_disk climbs means reuse is happening the ordinary way, not through shared beginnings. Usually GALAHAD_PAYLOAD_ORDER is unset, or your requests genuinely do not share one.

⚠ saver_chain_fallback climbing means the setting is on but a save could not use it, so those requests store the ordinary way. Nothing is lost; the sharing simply does not apply to them.

⚠ Read the counters from a line printed after the work. They are per-report figures, not running totals, so an old line still shows old numbers.


Models with sliding windows

Models such as Gemma 4 mix sliding-window and full-attention layers. vLLM keeps them apart with its hybrid KV manager, so a sliding layer holds only its window.

⭐ Galahad works with that manager on, so a sliding layer keeps only its window in GPU memory, and Galahad saves only that window. Measured on Gemma 4 31B, RTX A6000, 26 September 2026:

with the hybrid manager on
KV room in the GPU memory 60,484 tokens, 2.3× what the same memory holds with the manager off
one saved 9,472-token reading 1.6 GB: each window only, not the whole sequence (8.5 GB)

With Galahad attached, vLLM logs hybrid KV manager ON for every Gemma 4 and Olmo 3 model (29 September 2026).

⚠ If the log says hybrid-attention model + no SupportsHMA, the installed package is not the current one. Install the current package.


What you actually gain

⭐ The work is not done. That is the product. A reused prompt costs no prefill at all: no GPU cycles, no energy, no queue time, no capacity consumed.

GPU work avoided the prefill never runs. This holds on every machine
capacity freed those GPU-seconds serve other requests
time to first token the user waits for a load, not a computation
throughput depends on your hardware, see below

Measured on an H100 with Qwen3-30B-A3B:

turns/s
vLLM without Galahad 2.643
⭐ with Galahad 3.424
with Galahad, 1 in 3 loads deliberately failing 2.716

⭐ The last row is why it is safe to run. With a third of loads failing on purpose: no deadlock, no errors seen by clients, and still above vLLM without Galahad.

⭐ A slower GPU gains more: 2.91× on an RTX PRO 4500 against 1.32× on an H100 SXM. A fast GPU prefills quickly, so avoiding a prefill is worth less. The H100 figures are the floor, not the ceiling.


Storage

⚠ Network storage is about 17 times slower than local: 44 ms versus 780 ms for a 580 MB block. Galahad warns at startup if the store is not on a local disk, but does not refuse: 780 ms still beats recomputing.

⚠ A RAM disk does not help. Moving the store from disk to RAM did not raise throughput. A fast local disk is enough.


TensorRT-LLM

A TensorRT-LLM connector is not yet available.


Troubleshooting

Symptom Cause Fix
no [galahad] lines at startup the package is not installed in vLLM's environment pip install it into the same environment as vLLM
caching ENABLED never appears this model is not supported yet the startup line says why; vLLM serves normally
nothing saved prompts shorter than GALAHAD_MIN_TOKENS expected; lower it only if loading beats re-reading on your hardware
nothing saved, log says no licence the licence is missing or for fewer points see Install
no hits after a restart the store moved GALAHAD_CACHE_DIR must point at the same store as before
HTTP 500 once, KV load failure in the log a block on disk is damaged, policy is fail set "kv_load_failure_policy":"recompute"
much slower than expected the store is on network storage the startup warning names it; use local disk
throughput unchanged async load switched off leave GALAHAD_ASYNC_LOAD at its default (on)

Related

Clone this wiki locally