Skip to content

Performance

Sietse edited this page Oct 4, 2026 · 5 revisions

Performance and Limits

What Galahad costs, what it was measured on, and which numbers you should not carry to your own hardware.

Status ✅ Measured. Every figure names its hardware
Supported platform Linux x86-64. Windows, macOS and arm64 are not supported

In plain words

Every number on this page was measured on a specific machine. Where a number does not transfer to your hardware, this page says so.


A slower GPU gains more

The speedup grows as the GPU gets slower.

GPU Speedup
RTX PRO 4500 2.91×
H100 SXM 1.32×

⭐ The H100 numbers are the floor, not the ceiling. A fast GPU prefills quickly, so avoiding a prefill is worth less. On older or cheaper GPUs you gain more than the headline figures suggest.

⚠ Storage is a second, separate factor and is not predicted by the GPU.


Very long texts

Nothing to switch on. It works on its own.

How it works

  1. A long text is too big for the model to read at once, so it is read in parts.
  2. Galahad saves the model's reading of each part.
  3. When a question comes, the part that holds the answer is given back to the model, which answers from it. Nothing is read again.
  4. The GPU only ever holds one part, so GPU memory stays the same, whether the text is 1 million or 50 million tokens. Only the disk use grows.

Measured

One NVIDIA A40, Gemma 4 12B on llama.cpp, September 2026. A fact was placed at different depths in the text, then asked for:

Text length GPU memory (peak)
1 million tokens 33.8 GB
15 million tokens 34.1 GB
30 million tokens 34.1 GB
50 million tokens 34.1 GB

Depth barely changes the time. The deepest fact at 1 million tokens came back in 7.7 s, the deepest at 50 million in 9.6 s: 52 times deeper for 1.25 times the time.

Plan your disk for the part of your text you want to keep: every saved part stays on disk until you remove it.

How to use it

Nothing to switch on. You read the long text in parts and let Galahad save each part; the window moves over disk on its own.

With vLLM or SGLang: serve the model with the Galahad connector (see Install), set a store directory on a local disk, and send the text in parts the model's context can hold. Each part's reading is saved as it is read, and a later question restores the part it needs. No per-request flag.

export GALAHAD_CACHE_DIR=/var/galahad      # a local disk with room for the text
# then serve as on the Install page; feed the text in context-sized parts

Inside your own program (llama.cpp / libgalahad.so): run llama-server with slot save and restore pointed at a local directory, deposit each part, and save its slot. A later question restores the matching slot and answers from it.

# one part's reading streams to disk, GPU memory stays flat as the window moves
llama-server -m your-model.gguf -c 16384 --slot-save-path /var/galahad
#   -c is one part's size; the whole text can be far larger than -c

Sliding-window models (for example Gemma): add --swa-full so a restored part is complete. See Models with sliding windows.

Sizing.

You set What it means
one part how many tokens the model reads at once (its context); fits in GPU memory
the store directory on a local disk, large enough for every part you keep
GPU memory stays flat at one part's size, no matter how long the whole text is

⚠ Processed is not the same as retained is not the same as retrieved. The model processes every part once as it reads it; it retains in GPU memory only the one part it holds now; it retrieves from disk the part a question needs. The whole text is not attended to at once. The window moves.


Storage

local disk, 580 MB block 44 ms
network volume, same block 780 ms, about 17 times slower

⚠ Galahad warns at startup if the store is not on a local disk, and does not refuse. Some deployments have no local disk, and 780 ms still beats recomputing.

⚠ A RAM disk does not help. Moving the store from disk to RAM did not raise throughput. A fast local disk is enough; measure before you spend money on storage.


Shutdown

Records Flush
10k 1.4 ms
100k 18.8 ms
1M 327.6 ms

⚠ Skipping merlin_shutdown does not lose saved memory. It loses the usage counters that decide what is loaded first after a restart.


Measuring on your own machine

⚠ Timing on a busy machine measures the machine. Pin your benchmarks (taskset), take medians, and check whether your container's CPU quota matches what the operating system reports.

What to measure:

  1. merlin_min_useful_prefix(): below it, reuse costs more than it saves, and the answer depends on your hardware.
  2. Local versus network storage, with your real prompt sizes.
  3. Your prefill time. Every speedup ratio has it as the denominator; a fast GPU shrinks the win.
  4. Whether your model uses sliding window attention: see Models with sliding windows.

Known limits

Platform Linux x86-64 only
GPU telemetry real on NVIDIA GPUs only; on other GPUs the telemetry figures are placeholders

Related

Clone this wiki locally