Skip to content

Install

Sietse edited this page Oct 6, 2026 · 5 revisions

Install

Three ways to run Galahad, from nothing to a working memory, and how to prove each one works before you rely on it.

Status ✅ Works. Tested as a customer would, on fresh machines
Verified 1 October 2026, fresh NVIDIA L40S (Ubuntu 22.04): licence flow, and a zero-config encrypted store round-trip (deposit → restart → decrypt byte-exact) with no key file — the licence supplies the key. Also 29 September 2026, 4 × NVIDIA A10G (Ubuntu 24.04), vLLM 0.30
Pick one Bare (your own C code) · with vLLM · with SGLang

What you get

Everything installs with pip install galahad-kv — one package, no separate download. Inside it:

libgalahad.so all of Galahad, in one file, shipped inside the wheel. Self-contained: it needs only libcrypto.so.3 (OpenSSL 3), libzstd.so.1 and the C library, which every current Linux has
the vLLM + SGLang connectors the production data plane, loaded automatically
galahad command one CLI: galahad help, galahad doctor, and the full licence flow (free / buy / quote / status)
the licence tools bundled in the package and driven by the galahad command
the C headers merlin.h (storage / Taliesin) and galahad_blaise.h (Blaise), on disk inside the package for C/C++ hosts

The C functions use the merlin_ prefix. See Names.

What your machine needs

OS Linux x86-64, glibc 2.28 or newer (Ubuntu 20.04+, Debian 11+, RHEL 8+)
OpenSSL 3.0 or newer (libcrypto.so.3). Standard on Ubuntu 22.04+ and RHEL 9+
GPU an NVIDIA GPU the CUDA driver sees — a licence is bound to one. AMD and CPU-only hosts are not supported yet
Python 3.10–3.14 (the wheels cover these; there is no 3.9 wheel)

The bundled command-line tools (galahad licence etc.) carry their own extra libraries (such as libcurl) inside the package, so they also run on a minimal Linux. The engine libgalahad.so needs libcrypto.so.3, libzstd.so.1 and the C library, which every current Linux has.

# does the library load here at all? Nothing printed = yes.
ldd libgalahad.so | grep "not found"

Verify it really came from Corbenic (recommended)

pip install galahad-kv downloads from PyPI over HTTPS, and pip verifies each wheel's hash from the index automatically.

For a locked, reproducible install (recommended for production), pin the exact hashes so pip refuses anything that does not match:

# capture the hash of the wheel pip would install, once, in a trusted environment
pip download galahad-kv --no-deps -d ./galahad-dl
sha256sum ./galahad-dl/galahad_kv-*.whl        # record this value

# then require it everywhere
pip install galahad-kv --require-hashes -r requirements.txt   # with the pinned sha256

The published project is galahad-kv, owner Corbenic AI. Install only that name from PyPI.


Step 0: the licence (every install needs one)

Galahad does not write to disk without a licence. One GPU is free.

All licence commands are subcommands of the single galahad command that pip install galahad-kv puts on your PATH. A licence needs an NVIDIA GPU (AMD and CPU-only hosts are not supported yet).

# 1. how many points is this machine? (points = measured GPU speed)
galahad points

# 2a. one GPU: the free licence (non-commercial terms, 12 months, renewable)
galahad free --org <your-secret-key> --email you@company.com \
             --accept-non-commercial

# 2b. two or more GPUs: buy one (opens a checkout, then fetch it)
galahad buy   --org <your-secret-key> --gpus 4
galahad fetch --reference <the reference the checkout gave you>

# 3. what a refusal code means, if you get one
galahad explain GLH-L07

This writes two files to /etc/galahad/: licence.lic and org.key. Keep them together: Galahad reads org.key from the same folder as the licence, and stores nothing without it. Somewhere else: pass --out PATH and set GALAHAD_LICENCE_FILE=PATH wherever Galahad runs; move both files.

⚠ --org is your store identifier — keep it identical for as long as you want to read what you stored; a new --org means a new, empty store. Note --org is sent to the activation service (with your email, node-id and GPU info) to issue the signed licence, so it is an identifier, not your encryption secret.

Your store key, and what is sent

The key that encrypts your stored memory is held on your machine and is never transmitted to us. It is derived locally:

  • Passphrase (recommended, zero-knowledge): set GALAHAD_STORE_PASSPHRASE to a strong secret (≥ 12 bytes). The store key is scrypt(passphrase, salt = org), computed on your machine; we never receive or hold it.
  • Otherwise the local org.key (32 random bytes written to /etc/galahad/). It stays on disk and is not transmitted.

The licence does not carry a store key, so a copy of your licence alone cannot open your store — the passphrase or org.key is required. There is no recovery: lose both and the stored memory cannot be read by anyone, including us.

⚠ How you know the licence is working: run galahad doctor — it reports whether the library loads, a licence is installed and enforcing, and the store is writable. Without a valid licence Galahad serves normally but stores nothing. A new --org means a new, empty store.

⭐ Price: see corbenic.ai/pricing. A licence covers up to its points (+2%); a faster machine is refused with points-exceeded.


Option 1: Bare (your own code)

No serving engine. Your program links libgalahad.so and calls it. The library and headers ship inside the pip package — copy them out to a stable location:

pip install galahad-kv
# find what pip installed and copy the .so + headers out
PKG=$(python -c "import galahad_vllm, os; print(os.path.dirname(galahad_vllm.__file__))")
LIB=$(python -c "import merlin, os; print(os.path.join(os.path.dirname(merlin.__file__),'lib'))")
sudo mkdir -p /opt/galahad/lib /opt/galahad/include /var/galahad
sudo cp "$LIB"/libgalahad.so /opt/galahad/lib/
sudo ln -sf libgalahad.so /opt/galahad/lib/libgalahad.so.1
# headers are in the package; or keep using them in place

⚠ On every upgrade, re-run pip install -U galahad-kv and re-copy libgalahad.so (and refresh the .so.1 symlink). The loader opens the .1 name, so a stale copy there keeps running the previous version.

The five minute check

#include "merlin.h"
#include <stdio.h>
#include <string.h>

int main(void) {
    merlin_config cfg;
    memset(&cfg, 0, sizeof cfg);
    cfg.struct_size       = sizeof cfg;              /* required */
    cfg.ledger_path       = "/var/galahad/ledger.bin";
    cfg.block_dir         = "/var/galahad/blocks";
    cfg.tenancy           = MERLIN_TENANCY_SINGLE;   /* required, no default */
    cfg.model_fingerprint = 0xA11CE;                 /* required, no default */

    /* The store is encrypted. The KEY COMES FROM YOUR LICENCE -- nothing to
       configure here. Install it first (Step 0); merlin_init() finds it. */
    merlin_status st = merlin_init(&cfg);
    printf("merlin_init -> %s\n", merlin_status_string(st));
    merlin_shutdown();
    return st == MERLIN_OK ? 0 : 1;
}
cc -I /opt/galahad/include check.c -o check \
   -L /opt/galahad/lib -l:libgalahad.so -Wl,-rpath,/opt/galahad/lib
./check

Expected: merlin_init -> ok (with a licence installed at /etc/galahad/). Without a licence you get MERLIN_ERR_NO_KEY_PROVIDER — install one (Step 0).

⚠ No key file, and none needed. The licence supplies the storage key automatically, so there is nothing to generate, place beside the data, or rotate by hand. If instead you run a key manager (KMS), register it through merlin_set_key_provider() before merlin_init(). See Encryption at Rest.

The two fields with no default

tenancy: MERLIN_TENANCY_SINGLE for one organisation, MERLIN_TENANCY_MULTI when separate customers share the machine. Left at zero, Galahad refuses to start: a wrong guess would silently miss, or serve one customer's memory to another.

model_fingerprint: one number for model + tokenizer + quantisation. Memory is only shared between identical fingerprints: reusing across a mismatch gives fluent, wrong output instead of an error.

struct_size: always sizeof cfg, never a number. It lets a newer library accept your older struct.

Next: API Reference.


Option 2: With vLLM

No code. One setting on vllm serve.

Install

pip install vllm                         # if you do not have it yet
pip install galahad-kv                    # the package; libgalahad.so is inside it

Check:

galahad doctor                           # library loads? licence? store writable?

Run

export GALAHAD_CACHE_DIR=/var/galahad          # the store (fast local disk)
# No GALAHAD_KEY_FILE: the licence supplies the storage key automatically.
export GALAHAD_LICENCE_FILE=/etc/galahad/licence.lic   # only if not in /etc/galahad

vllm serve Qwen/Qwen3-8B \
  --kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both"}'

That is the whole setup. vLLM finds Galahad by itself (the package registers a vLLM plugin); there is no module path to set. Encryption at rest is on with no key file — the licence is the key. (A customer-managed KMS is the alternative; see Encryption at Rest.)

Several GPUs

Nothing extra. Split the model as you normally would:

vllm serve Qwen/Qwen3-32B --tensor-parallel-size 4 \
  --kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both"}'

A prompt counts as remembered only when it was saved on every GPU.

Recommended: recompute instead of failing

--kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both",
                       "kv_load_failure_policy":"recompute"}'

vLLM's default for a memory that cannot be read (a damaged file on disk) is to fail that request. With recompute, vLLM works it out again and the user never notices. Galahad then does not offer that block again.

How you know it works

Look for these lines when vLLM starts:

[galahad] caching ENABLED: 36 layers, uniform geometry
[galahad] multi-GPU: the scheduler reads 4 GPU stores (tp0, tp1, tp2, tp3)   <- several GPUs only

Ask the same long question twice. The second time vLLM's log shows the prompt was served from Galahad, and the answer is the same.

Next: vLLM Connector: what it does, and the tested models.


Option 3: With SGLang

Galahad sits under SGLang's own cache as a durable disk tier.

⚠ Host RAM matters for SGLang. SGLang's --enable-hierarchical-cache builds a large CPU-RAM mirror of the GPU KV cache (sized by --hicache-ratio), and Galahad is the disk tier under it. On a small-RAM box SGLang refuses to start with Not enough host memory available. Requesting X GB but only have Y GB free — this is SGLang sizing its host cache, before Galahad is even reached, not a Galahad fault. Rule of thumb: system RAM ≥ model + SGLang's static allocation + (GPU-KV × hicache-ratio) + a few GB free (e.g. a 24 GB GPU wants a 32 GB+ host). A 16 GB host cannot run the hierarchical cache; use a bigger host, a smaller --hicache-ratio, or run without --enable-hierarchical-cache (then there is no disk tier). vLLM does not have this requirement.

Install

pip install "sglang[all]==0.5.20"        # pinned: tested version
pip install galahad-kv

⚠ Pin SGLang 0.5.20 (and vLLM 0.30 on the vLLM path). Newer SGLang/vLLM change internal APIs the connector binds to; an unpinned pip install can pull a version that fails to register the backend.

Run

export GALAHAD_CACHE_DIR=/var/galahad
# No GALAHAD_KEY_FILE: the licence supplies the storage key automatically.
export GALAHAD_LICENCE_FILE=/etc/galahad/licence.lic   # only if not in /etc/galahad

python -m sglang.launch_server --model-path Qwen/Qwen3-8B \
  --page-size 64 \
  --enable-hierarchical-cache \
  --hicache-io-backend direct --hicache-mem-layout page_first_direct \
  --hicache-storage-prefetch-policy wait_complete \
  --hicache-storage-backend dynamic \
  --hicache-storage-backend-extra-config \
    '{"backend_name":"galahad","module_path":"galahad_vllm.sglang_backend","class_name":"GalahadHiCacheStorage"}'

⚠ Keep --hicache-storage-prefetch-policy wait_complete. With SGLang's default (timeout), SGLang can start computing before the memory has arrived from disk, and then computes the whole prompt anyway.

How you know it works

At start:

Creating dynamic storage backend 'galahad' (galahad_vllm.sglang_backend.GalahadHiCacheStorage)
[galahad] disk budget: cap ... GB, floor 5.0 GB free (/var/galahad/blocks)

When a prompt is served from Galahad (here after SGLang's own cache was emptied), SGLang logs the tokens it loaded and computes only the rest:

HiCache prefetch success ... loaded=1792
Prefill batch, #new-seq: 1, #new-token: 32, #cached-token: 1792, ...

Next: SGLang Backend.


Settings you may want

Setting Default What it does
GALAHAD_CACHE_DIR /tmp/galahad where the store lives. ⚠ /tmp is lost on reboot and shared by all users: set this
GALAHAD_LICENCE_FILE /etc/galahad/licence.lic the licence — it supplies the storage key, so no key file is needed
GALAHAD_DISK_BUDGET_GB half of the free disk space at start the most Galahad may write. 0 turns the limit off
GALAHAD_DISK_FLOOR_GB 5 always leave this much free. 0 turns it off
GALAHAD_TENANCY single multi when separate customers share the server
GALAHAD_MIN_TOKENS 256 shorter prompts are not saved (re-reading them is cheaper)
GALAHAD_DISABLED not set 1 turns Galahad off. The server keeps serving, without memory
GALAHAD_LIBRARY_PATH the copy inside the package use this libgalahad.so instead. If it cannot be loaded, Galahad stops with an error rather than use another copy

Every setting name starts with GALAHAD_.

⚠ Put the store on a fast local disk. A network volume turns a 44 ms read into 780 ms; Galahad prints a warning at startup when it detects one.


What lands on disk

Everything Galahad writes lives in one directory: GALAHAD_CACHE_DIR with vLLM or SGLang, or the ledger_path and block_dir you set in C. It holds an index and the memory itself, encrypted (AES-256-GCM).

disk space give that directory the space you want Galahad to use; GALAHAD_DISK_BUDGET_GB caps it
backup back up the whole directory, together with your key and org.key. Without them, the backup cannot be read
moving it move the whole directory, and point GALAHAD_CACHE_DIR at the new place

Nothing else is written, and nothing is sent anywhere.


Troubleshooting

Symptom Cause Fix
cannot open shared object file the loader cannot find the library -Wl,-rpath, or install into a path the loader searches
libcrypto.so.3: cannot open OpenSSL 1.1 only install OpenSSL 3 (Ubuntu 22.04+ has it)
GLIBC_2.xx not found the OS is older than glibc 2.28 use Ubuntu 20.04 or newer
nothing written to disk, log says why no licence, or org.key is not beside it check GALAHAD_LICENCE_FILE and that org.key sits in the same folder
points-exceeded the machine is faster than the licence a licence with more points: galahad buy
ERR_NO_KEY_PROVIDER / GLH-E17 the store is encrypted and no key was given install a licence (galahad free / galahad install) — it supplies the key; or register a KMS with merlin_set_key_provider()
tenancy mode not configured cfg.tenancy left at zero set SINGLE or MULTI
ABI/struct_size mismatch header and library disagree use the header shipped with the library
vLLM: No such file or directory: 'ninja' vLLM builds kernels and needs ninja on PATH run from the environment vLLM is installed in (source venv/bin/activate)
vLLM: hybrid KV plan unavailable a model layout Galahad does not support yet vLLM serves normally, without memory. Tell us the model
a model runs but Galahad saves nothing prompt shorter than GALAHAD_MIN_TOKENS expected: short prompts are cheaper to re-read

Related

Clone this wiki locally