Repository navigation
Prefix Sharing
Many agents, one shared preamble, computed once.
| Status | ✅ Works: through the C API, and through the vLLM connector with one setting |
| Verified | A40: 50 of 50 agents bit identical to a full read · 23.1× fewer prefill tokens at 50 agents |
| Minimum chunk | 64 tokens. A smaller chunk is refused, not rounded up |
Ten agents doing the same job usually start from the same instructions: the same handbook, the same rules, the same system prompt. Normally every one of them reads all of it before doing anything.
This lets them share that reading. The common beginning is computed once; each agent only pays for the part that is different: its own question.
The saving grows with the size of the fleet.
Fifty support agents, all working from the same 2,000 word policy document. One of them reads it. The other forty nine start from that same understanding and go straight to their own customer's question.
How it works
- Many requests often start with the same text: rules, a handbook, a system prompt.
- Galahad saves that shared start once.
- Every next request reuses it and only computes its own, different part.
How to use it
-
With vLLM: set
GALAHAD_PAYLOAD_ORDER=chunkonce, before first use. -
With the C API: call
merlin_router_set_calibrationonce at start-up, then save and look up in pieces of at least 64 tokens. - Check the counters to see how much was reused.
On an A40, all layers on the GPU, three repetitions:
| Agents | Prefill tokens | Saving |
|---|---|---|
| 4 | 8,392 → 2,248 | 3.7× |
| 10 | 20,980 → 2,548 | 8.2× |
| 50 | 104,900 → 4,548 | 23.1× |
| 200 | 419,600 → 12,048 | 34.8× |
Wall clock at 50 agents: 39.09 s → 7.27 s (5.4×). Mean time to first token 706 ms → 57 ms.
⚠ Token saving and wall clock are different numbers. 23.1× fewer tokens gave 5.4× wall clock, because the GPU was already fast at prefill. Quote the one that matches your question.
⭐ The first read is shared out over the fleet. At 4 agents it is most of the cost; at 200 it hardly counts.
⚠ A poor case: 256 shared tokens in front of a 2,048 token task. If your agents share only a short preamble and then diverge heavily, this feature is not for you.
⭐ Memory built from shared chunks is bit identical to reading the whole prompt fresh. Measured: 0 of 1,368,040 state bytes differ. Each of the 50 agents was compared against a full read, in both internal state and output: 50 of 50 identical.
Step 1, once at start-up: tell Galahad your hardware's real costs, so Galahad only reuses when it is faster.
merlin_router_set_calibration(prefill_ns_per_token, /* your measured prefill time per token */
read_ns_per_byte, /* your measured storage read time per byte */
kv_bytes_per_token, /* your model's memory size per token */
0); /* 0 = built-in overhead */⚠ Without this call, nothing is reused. Every prefix lookup is declined.
Step 2, per request:
/* Is reuse worth it on this storage? (in tokens) */
int32_t min_tokens = merlin_min_useful_prefix(MERLIN_TRANSPORT_LOCAL);
/* Find the longest saved beginning */
merlin_prefix_result r = {0};
r.struct_size = sizeof r;
merlin_lookup_prefix(tokens, n_tokens, tenant_id, 64, MERLIN_TRANSPORT_LOCAL, &r);
/* r.tokens_cached tokens are ready; prefill only the rest */
/* Save this request's chunks so later requests can share them */
merlin_deposit_chain(tokens, n_tokens, tenant_id, 64, chunk_bytes, chunk_lens, n_chunks);See API Reference for every argument, and for
merlin_hash_prefix_chain.
⭐ Ask merlin_min_useful_prefix first. Below that length, reuse costs more
than it saves, and the number depends on your hardware.
merlin_lookup_prefix(tokens, n, tenant_id, 32, MERLIN_TRANSPORT_LOCAL, &r); /* MERLIN_ERR_INVALID_ARG */⚠ A chunk below 64 tokens (MERLIN_MIN_CHUNK_TOKENS) is refused, not rounded
up, so you always get the chunk size you asked for.
⚠ Only shared beginnings are reused. Shared middles are not.
The same words at a different position give different internal state. Two documents that share a paragraph in the middle share nothing reusable. Put the shared text first.
✅ Verified 17 September 2026 on a running vLLM 0.29 server with a 31B model. Set one variable and requests that share a beginning reuse it:
export GALAHAD_PAYLOAD_ORDER=chunk # default: layer| What was measured on a live server | |
|---|---|
| a ~700-token shared preamble, three follow-up questions | 3 of 3 reused it |
| a ~359-token preamble, shorter than one coarse step | 2 of 4 reused it, at 64-token resolution |
| answers | correct, 0.4–1.4 s per request |
⭐ With this setting, every 64 tokens is a sharing point. Without it, reuse happens in coarser steps, so a short shared preamble may not be reused at all.
| Default | layer |
Storage cost of chunk |
sharing uses extra storage |
| When it pays | many requests sharing a long beginning. With nothing shared, the extra storage is the only effect |
⚠ Set it once, before first use, and leave it. Changing it starts a fresh
cache: entries saved the other way are not found (they are never misread). Any
value other than layer or chunk is refused at start-up.
| Symptom | Cause | Fix |
|---|---|---|
MERLIN_ERR_INVALID_ARG on a prefix call |
chunk_size below 64 |
raise it to 64 or more |
| every prefix lookup is declined | not calibrated | call merlin_router_set_calibration once at start-up |
merlin_min_useful_prefix returns INT32_MAX
|
not calibrated, or reuse is slower than recomputing on this storage | calibrate; check your storage speed |
tokens_cached is 0 and declined_reason is MERLIN_GRAFT_DECLINED_BY_ROUTER
|
the prefix is saved but too short to be worth reusing | expected; the request prefills |
| prefix lookups always miss | nothing stored that prefix as a chain | through the C API, call merlin_deposit_chain; through the connector, see the setting
|
| saving much smaller than the table | short shared preamble, long divergent tasks | see the poor case above |
| reuse seems not worth it | below merlin_min_useful_prefix
|
ask the function rather than guessing |
everything misses after changing GALAHAD_PAYLOAD_ORDER |
expected: it starts a fresh cache | the cache refills; change it once, before first use |
- API Reference: prefix sharing
- Agent Connectors: the fleets this is built for
- vLLM Connector
- Install