-
Notifications
You must be signed in to change notification settings - Fork 1
Home
Long-term memory for AI. It makes running a model cheaper, faster and smarter.
An AI model has no memory of its own. Every time it reads something, it works through it from the beginning. The moment it answers, all of that work is thrown away. Tomorrow the same document arrives and the model starts again from nothing.
Galahad gives the model a memory that lasts. What the model works out is kept, and given back to it later, exactly as it was. Not a summary of it. Not something similar. The same thing, byte for byte.
That memory has two halves, because a model needs to remember two different kinds of thing:
| Taliesin | remembers thinking: the work the model already did on text it has read |
| Blaise | remembers documents: the exact text of what you gave it |
⭐ And it works with the model you already have. You do not change the model, you do not train anything, and you do not send your data anywhere.
A memory that lasts makes a model cheaper to run, faster to answer, and better at answering. Every number below is measured, and the page that measured it is linked.
The expensive part of running a model is making it read. Reading a long document costs GPU time every single time, unless the reading is remembered.
| Measured | |
|---|---|
| 50 agents sharing one document (A40) | 23.1× fewer tokens to read, 104,900 down to 4,548 |
| the same job on the clock | 39.09 s down to 7.27 s |
| time to the first word | 706 ms down to 57 ms |
| on an RTX PRO 4500 | 2.91× faster |
| on an H100 | 1.32× faster |
| encryption on top of it | costs under 1% |
| what you are billed for | counted per tenant, not estimated. Cost Reporting |
⭐ The GPU work does not get faster. It does not happen.
At 50 agents the work falls from 104,900 tokens to 4,548. That work was not sped up. It was never started. And that is why cheaper and faster are one effect rather than two:
| You buy fewer GPU seconds | reading is the expensive part, and this is reading that never happens |
| The GPU has room again | the cycles that were re-reading yesterday's document are free |
| So it generates more, faster | a GPU with room serves more requests and produces tokens sooner. The free capacity turns straight back into speed |
See Performance and Prefix Sharing.
⭐ A model with memory gives better answers. Three things change:
| It remembers across restarts | what it read yesterday comes back exactly, not worked out again |
| It answers consistently | the same question gets the same answer, every time, because the memory is byte-identical rather than worked out again |
| It can read the source again | Blaise gives back the exact text of your documents, not a summary |
| One file |
libgalahad.so, self-contained, with 104 functions you can call from C (91 for Taliesin and storage, 13 for Blaise). See API Reference
|
| It runs on its own | no serving engine needed. The library is the product, and it works with whatever you already run |
| Or with one setting | using vLLM or SGLang? Add one line of config and it works, with no code at all |
| Your data stays yours | it runs on your machine; nothing is sent anywhere |
| Every model | ⭐ 30 of 30 current models tested with vLLM on 29 September 2026: Qwen3 / 3.5 / 3.6 / 3.8, Gemma 4, Mistral, GLM, Llama, DeepSeek-R1, Phi-4, GPT-oss, Nemotron, Olmo. List |
| One GPU or many | split a model over 2, 4 or more GPUs as you normally would; nothing else changes |
| One GPU is free | from two GPUs on you pay by measured GPU speed. Install |
| Galahad | the full system: what you install and run |
| Taliesin | the memory for thinking (the KV cache) |
| Blaise | the memory for documents |
The C functions use the merlin_ prefix, and Blaise's use galahad_blaise_.
Every setting starts with GALAHAD_. Names has the detail.
| You are | Start here |
|---|---|
| new to this, or curious about AI | Names, then Install |
| running your own code, no serving engine | Install, then API Reference. This is the bare library, and it needs nothing else |
| running vLLM | Install, then vLLM Connector |
| running SGLang | Install, then SGLang Backend |
| running Kubernetes | Install, then Running on Kubernetes |
| an engineer with the library in hand | Install, then API Reference |
⚠ Every page starts simple and gets technical after. You do not need to read C to understand what a page is about.
| Page | Status | For |
|---|---|---|
| Names | ⭐ The vocabulary | Galahad, Taliesin, Blaise: what each one is and which question it answers. Read it first if a name here is unfamiliar |
| Install | ✅ Works | The licence, then three ways to run it: bare, with vLLM, with SGLang. Tested as a customer would on a fresh machine |
| API Reference | ✅ Works | All 104 functions, grouped by what you are trying to do |
| Page | Status | For |
|---|---|---|
| vLLM Connector | ✅ Works | Prompt memory on disk for vLLM, on one GPU or many. 30 tested models |
| SGLang Backend | ✅ Works | A durable disk tier under RadixAttention |
| Running on Kubernetes | ✅ Works | Probes, metrics, and the storage setting that matters most |
| Page | Status | For |
|---|---|---|
| Agent Connectors | ✅ Works | Adding audit and memory to an existing agent: LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, Letta |
| Branching and Snapshots | ✅ Works | Exploring many possibilities from one memory without copying it, and returning to an exact moment, bit for bit |
| Prefix Sharing | ✅ Works | Many agents sharing one preamble: 23.1× fewer prefill tokens at 50 agents. And through the vLLM connector with one setting |
| Replay | ✅ Works | Regression testing an agent that never gives the same answer twice |
| Memory Inspector | ✅ Works | A read-only window into what was remembered, and when |
| Page | Status | For |
|---|---|---|
| Encryption at Rest | ✅ Works | AES-256-GCM over everything on disk, with your own key manager holding the key |
| Multi Tenant Isolation | ✅ Works, capacity limits are configured | Keeping customers apart, and seeing who crowds whom |
| Collision Safety | ✅ Works | Two independent checks before anything is reused, and the function name that decides whether you get it |
| Payload Validation | ✅ Works | Refusing a corrupted block before the model ever sees it |
| Page | Status | For |
|---|---|---|
| Performance and Limits | ✅ Measured | What it costs, texts up to 50 million tokens, and which figures do not transfer to yours |
| Cost Reporting | ⚠ Ships inert by design | What reuse saved in seconds and money, and why the money line stays blank until you supply a price |
More pages arrive as each capability is verified.
✅ Yes: what a function does · its signature · what it returns · what it guarantees · what it refuses · the limits · worked examples · the error messages you will actually see.
❌ No: how it works inside. That is not part of the delivery and is not documented here.
⚠ Every number in this wiki is measured on real hardware.
| ✅ Works | verified on real hardware |
| ⚙ Works once configured | the feature is there, and inert until you set it |
| ⚠ Works, with a stated limit | read the limit before relying on it |
- Dates are written out in full, for example 17 September 2026.
- Arrows (→) in call examples and log lines stand for the two character ASCII arrow a program prints.