Repository navigation
Performance
What Galahad costs, what it was measured on, and which numbers you should not carry to your own hardware.
| Status | ✅ Measured. Every figure names its hardware |
| Supported platform | Linux x86-64. Windows, macOS and arm64 are not supported |
Every number on this page was measured on a specific machine. Where a number does not transfer to your hardware, this page says so.
The speedup grows as the GPU gets slower.
| GPU | Speedup |
|---|---|
| RTX PRO 4500 | 2.91× |
| H100 SXM | 1.32× |
⭐ The H100 numbers are the floor, not the ceiling. A fast GPU prefills quickly, so avoiding a prefill is worth less. On older or cheaper GPUs you gain more than the headline figures suggest.
⚠ Storage is a second, separate factor and is not predicted by the GPU.
Nothing to switch on. It works on its own.
- A long text is too big for the model to read at once, so it is read in parts.
- Galahad saves the model's reading of each part.
- When a question comes, the part that holds the answer is given back to the model, which answers from it. Nothing is read again.
- The GPU only ever holds one part, so GPU memory stays the same, whether the text is 1 million or 50 million tokens. Only the disk use grows.
One NVIDIA A40, Gemma 4 12B on llama.cpp, September 2026. A fact was placed at different depths in the text, then asked for:
| Text length | GPU memory (peak) |
|---|---|
| 1 million tokens | 33.8 GB |
| 15 million tokens | 34.1 GB |
| 30 million tokens | 34.1 GB |
| 50 million tokens | 34.1 GB |
Depth barely changes the time. The deepest fact at 1 million tokens came back in 7.7 s, the deepest at 50 million in 9.6 s: 52 times deeper for 1.25 times the time.
Plan your disk for the part of your text you want to keep: every saved part stays on disk until you remove it.
Nothing to switch on. You read the long text in parts and let Galahad save each part; the window moves over disk on its own.
With vLLM or SGLang: serve the model with the Galahad connector (see Install), set a store directory on a local disk, and send the text in parts the model's context can hold. Each part's reading is saved as it is read, and a later question restores the part it needs. No per-request flag.
export GALAHAD_CACHE_DIR=/var/galahad # a local disk with room for the text
# then serve as on the Install page; feed the text in context-sized partsInside your own program (llama.cpp / libgalahad.so): run llama-server with
slot save and restore pointed at a local directory, deposit each part, and save
its slot. A later question restores the matching slot and answers from it.
# one part's reading streams to disk, GPU memory stays flat as the window moves
llama-server -m your-model.gguf -c 16384 --slot-save-path /var/galahad
# -c is one part's size; the whole text can be far larger than -cSliding-window models (for example Gemma): add --swa-full so a restored
part is complete. See
Models with sliding windows.
Sizing.
| You set | What it means |
|---|---|
| one part | how many tokens the model reads at once (its context); fits in GPU memory |
| the store directory | on a local disk, large enough for every part you keep |
| GPU memory | stays flat at one part's size, no matter how long the whole text is |
⚠ Processed is not the same as retained is not the same as retrieved. The model processes every part once as it reads it; it retains in GPU memory only the one part it holds now; it retrieves from disk the part a question needs. The whole text is not attended to at once. The window moves.
| local disk, 580 MB block | 44 ms |
| network volume, same block | 780 ms, about 17 times slower |
⚠ Galahad warns at startup if the store is not on a local disk, and does not refuse. Some deployments have no local disk, and 780 ms still beats recomputing.
⚠ A RAM disk does not help. Moving the store from disk to RAM did not raise throughput. A fast local disk is enough; measure before you spend money on storage.
| Records | Flush |
|---|---|
| 10k | 1.4 ms |
| 100k | 18.8 ms |
| 1M | 327.6 ms |
⚠ Skipping merlin_shutdown does not lose saved memory. It loses the usage
counters that decide what is loaded first after a restart.
⚠ Timing on a busy machine measures the machine. Pin your benchmarks
(taskset), take medians, and check whether your container's CPU quota matches
what the operating system reports.
What to measure:
-
merlin_min_useful_prefix(): below it, reuse costs more than it saves, and the answer depends on your hardware. - Local versus network storage, with your real prompt sizes.
- Your prefill time. Every speedup ratio has it as the denominator; a fast GPU shrinks the win.
- Whether your model uses sliding window attention: see Models with sliding windows.
| Platform | Linux x86-64 only |
| GPU telemetry | real on NVIDIA GPUs only; on other GPUs the telemetry figures are placeholders |
- vLLM Connector: throughput, storage and sliding windows
- Prefix Sharing: where the largest measured savings are
- Install · API Reference