Skip to content
Sietse edited this page Sep 30, 2026 · 5 revisions

Galahad

Long-term memory for AI. It makes running a model cheaper, faster and smarter.


What Galahad is

An AI model has no memory of its own. Every time it reads something, it works through it from the beginning. The moment it answers, all of that work is thrown away. Tomorrow the same document arrives and the model starts again from nothing.

Galahad gives the model a memory that lasts. What the model works out is kept, and given back to it later, exactly as it was. Not a summary of it. Not something similar. The same thing, byte for byte.

That memory has two halves, because a model needs to remember two different kinds of thing:

Taliesin remembers thinking: the work the model already did on text it has read
Blaise remembers documents: the exact text of what you gave it

⭐ And it works with the model you already have. You do not change the model, you do not train anything, and you do not send your data anywhere.


What it gives you

A memory that lasts makes a model cheaper to run, faster to answer, and better at answering. Every number below is measured, and the page that measured it is linked.

Cheaper and faster, and they are the same thing

The expensive part of running a model is making it read. Reading a long document costs GPU time every single time, unless the reading is remembered.

Measured
50 agents sharing one document (A40) 23.1× fewer tokens to read, 104,900 down to 4,548
the same job on the clock 39.09 s down to 7.27 s
time to the first word 706 ms down to 57 ms
on an RTX PRO 4500 2.91× faster
on an H100 1.32× faster
encryption on top of it costs under 1%
what you are billed for counted per tenant, not estimated. Cost Reporting

⭐ The GPU work does not get faster. It does not happen.

At 50 agents the work falls from 104,900 tokens to 4,548. That work was not sped up. It was never started. And that is why cheaper and faster are one effect rather than two:

You buy fewer GPU seconds reading is the expensive part, and this is reading that never happens
The GPU has room again the cycles that were re-reading yesterday's document are free
So it generates more, faster a GPU with room serves more requests and produces tokens sooner. The free capacity turns straight back into speed

See Performance and Prefix Sharing.

Smarter

⭐ A model with memory gives better answers. Three things change:

It remembers across restarts what it read yesterday comes back exactly, not worked out again
It answers consistently the same question gets the same answer, every time, because the memory is byte-identical rather than worked out again
It can read the source again Blaise gives back the exact text of your documents, not a summary

What you actually get

One file libgalahad.so, self-contained, with 104 functions you can call from C (91 for Taliesin and storage, 13 for Blaise). See API Reference
It runs on its own no serving engine needed. The library is the product, and it works with whatever you already run
Or with one setting using vLLM or SGLang? Add one line of config and it works, with no code at all
Your data stays yours it runs on your machine; nothing is sent anywhere
Every model ⭐ 30 of 30 current models tested with vLLM on 29 September 2026: Qwen3 / 3.5 / 3.6 / 3.8, Gemma 4, Mistral, GLM, Llama, DeepSeek-R1, Phi-4, GPT-oss, Nemotron, Olmo. List
One GPU or many split a model over 2, 4 or more GPUs as you normally would; nothing else changes
One GPU is free from two GPUs on you pay by measured GPU speed. Install

The names you will see

Galahad the full system: what you install and run
Taliesin the memory for thinking (the KV cache)
Blaise the memory for documents

The C functions use the merlin_ prefix, and Blaise's use galahad_blaise_. Every setting starts with GALAHAD_. Names has the detail.


Where to start

You are Start here
new to this, or curious about AI Names, then Install
running your own code, no serving engine Install, then API Reference. This is the bare library, and it needs nothing else
running vLLM Install, then vLLM Connector
running SGLang Install, then SGLang Backend
running Kubernetes Install, then Running on Kubernetes
an engineer with the library in hand Install, then API Reference

⚠ Every page starts simple and gets technical after. You do not need to read C to understand what a page is about.


Pages

Start here

Page Status For
Names ⭐ The vocabulary Galahad, Taliesin, Blaise: what each one is and which question it answers. Read it first if a name here is unfamiliar
Install ✅ Works The licence, then three ways to run it: bare, with vLLM, with SGLang. Tested as a customer would on a fresh machine
API Reference ✅ Works All 104 functions, grouped by what you are trying to do

Connecting it to your engine

Page Status For
vLLM Connector ✅ Works Prompt memory on disk for vLLM, on one GPU or many. 30 tested models
SGLang Backend ✅ Works A durable disk tier under RadixAttention
Running on Kubernetes ✅ Works Probes, metrics, and the storage setting that matters most

Capabilities

Page Status For
Agent Connectors ✅ Works Adding audit and memory to an existing agent: LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, Letta
Branching and Snapshots ✅ Works Exploring many possibilities from one memory without copying it, and returning to an exact moment, bit for bit
Prefix Sharing ✅ Works Many agents sharing one preamble: 23.1× fewer prefill tokens at 50 agents. And through the vLLM connector with one setting
Replay ✅ Works Regression testing an agent that never gives the same answer twice
Memory Inspector ✅ Works A read-only window into what was remembered, and when

Running it safely

Page Status For
Encryption at Rest ✅ Works AES-256-GCM over everything on disk, with your own key manager holding the key
Multi Tenant Isolation ✅ Works, capacity limits are configured Keeping customers apart, and seeing who crowds whom
Collision Safety ✅ Works Two independent checks before anything is reused, and the function name that decides whether you get it
Payload Validation ✅ Works Refusing a corrupted block before the model ever sees it

Understanding the numbers

Page Status For
Performance and Limits ✅ Measured What it costs, texts up to 50 million tokens, and which figures do not transfer to yours
Cost Reporting ⚠ Ships inert by design What reuse saved in seconds and money, and why the money line stays blank until you supply a price

More pages arrive as each capability is verified.


What goes in this wiki

✅ Yes: what a function does · its signature · what it returns · what it guarantees · what it refuses · the limits · worked examples · the error messages you will actually see.

❌ No: how it works inside. That is not part of the delivery and is not documented here.

⚠ Every number in this wiki is measured on real hardware.


Status legend used on every page

✅ Works verified on real hardware
⚙ Works once configured the feature is there, and inert until you set it
⚠ Works, with a stated limit read the limit before relying on it

Reading this wiki

  • Dates are written out in full, for example 17 September 2026.
  • Arrows (→) in call examples and log lines stand for the two character ASCII arrow a program prints.

Clone this wiki locally