Repository navigation
Kubernetes
Probes, metrics, an off switch, and the storage mistake that silently makes every read about 17 times slower.
| Status | ✅ Works |
| What you set up | a local disk volume for the store, the Galahad settings, three health probes, and a metrics scrape |
| Needs | a vLLM or SGLang image with the Galahad package installed, and a licence (Install) |
On Kubernetes, things restart constantly: deploys, autoscaling, spot reclaims, node drains. Several times a day, by design.
That makes durable memory worth much more here than on a single server. On Kubernetes the platform restarts things for you, and Galahad means a restarted pod does not start empty.
A rollout replaces every pod at lunchtime. Without this, every pod starts empty and the next hour is slow for everyone. With it, the new pods read the same memory the old ones built.
Kubernetes' default storage is usually network backed, and it is about 17 times slower.
| local disk, 580 MB block | 44 ms |
| network volume, same block | 780 ms |
⚠ The pod stays healthy. The cache works and everything looks correct, but
every read costs seventeen times what it should. Galahad prints a warning at
startup, and the merlin_storage_class metric shows the class (see
Metrics); check for both.
Two more storage choices that lose the cache:
| Trap | What happens |
|---|---|
emptyDir |
erases the cache on every restart |
| pod-lifetime NVMe | fast, looks local, and is wiped on every restart |
⭐ Give the store a local disk that outlives the pod, for example a Kubernetes local PersistentVolume on the node's NVMe:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: galahad-local-nvme
provisioner: kubernetes.io/no-provisioner
volumeBindingMode: WaitForFirstConsumer
reclaimPolicy: Retain
---
apiVersion: v1
kind: PersistentVolume
metadata:
name: galahad-local-nvme-node1
spec:
capacity:
storage: 340Gi # must not be larger than the disk
accessModes: [ReadWriteOnce]
persistentVolumeReclaimPolicy: Retain
storageClassName: galahad-local-nvme
local:
path: /mnt/disks/galahad-nvme
nodeAffinity:
required:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/hostname
operator: In
values: ["<your-node-name>"]⚠ Size the claim to the disk. A claim larger than the local disk never
binds: the pod stays Pending, while the same size on network storage works at
once. That makes local disk look broken when it is not.
Mount the volume and point Galahad at it:
env:
- name: GALAHAD_CACHE_DIR
value: /data/galahad
- name: GALAHAD_TENANCY
value: single # or multi, when pods serve different customers
- name: GALAHAD_LICENCE_FILE
value: /etc/galahad/licence.lic
volumeMounts:
- name: galahad-store
mountPath: /data/galahadGALAHAD_TENANCY accepts only single or multi. Any other value is refused
with an error, so a typo cannot silently turn the cache off.
With vLLM, the Galahad connector serves three health endpoints on port 8081
(set GALAHAD_HEALTH_PORT to change it):
| Path | Use it for | What it checks |
|---|---|---|
/health/started |
startupProbe |
startup has finished, including reading the store's index |
/health/ready |
readinessProbe |
the server can serve. Stays ready with a degraded cache |
/health/live |
livenessProbe |
the server process is alive. Does not check the cache |
ports:
- { name: health, containerPort: 8081 }
- { name: metrics, containerPort: 9090 }
startupProbe:
httpGet: { path: /health/started, port: health }
periodSeconds: 5
failureThreshold: 60 # a large store takes longer to read at startup
readinessProbe:
httpGet: { path: /health/ready, port: health }
periodSeconds: 10
livenessProbe:
httpGet: { path: /health/live, port: health }
periodSeconds: 20
failureThreshold: 3⭐ Liveness does not include the cache. If the cache is unavailable, the pod is degraded, not dead, so it keeps serving instead of being restarted.
⚠ Readiness never reports "healthy" when it cannot tell.
⚠ Give the startup probe enough time. A pod that holds a large store needs longer to read its index at startup. If the startup probe gives up too early, the pod is killed while it is still warming, and it looks like a crash loop.
env:
- name: GALAHAD_DISABLED
value: "1"Galahad starts switched off and the server serves normally, without memory. A
bad storage class or an unwritable store directory cannot put your pod into a
crash loop, and you can turn Galahad off without rebuilding an image. An empty
value or 0 means Galahad is on.
merlin_shutdown() writes pending records to disk:
| Records | Flush |
|---|---|
| 10k | 1.4 ms |
| 100k | 18.8 ms |
| 1M | 327.6 ms |
The Kubernetes default of terminationGracePeriodSeconds: 30 is plenty: the
flush is about 1.6% of that at a million records.
⚠ Skipping shutdown does not lose saved memory. It loses the usage counters that decide what is loaded first after a restart.
With vLLM, the Galahad connector serves Prometheus metrics at /metrics on port
9090 (set GALAHAD_METRICS_PORT to change it), with no prometheus_client
dependency, so nothing extra goes into your image.
Metric names start with merlin_. The ones to watch:
| Metric | What it tells you |
|---|---|
merlin_backend_attached |
1 Galahad is attached. 0 it is installed but doing nothing |
merlin_storage_class |
the storage class of the block folder, as a class label. Anything but local is much slower |
merlin_store_durable |
1 the store survived on this node. 0 it was wiped (pod-lifetime storage). -1 unknown, for example a new volume or a pod on a new node |
merlin_hits_total · merlin_misses_total
|
memory found and not found |
merlin_write_failures_total |
saves that failed |
⭐ Scrape merlin_store_durable in particular: it tells you whether the
store you think is durable actually is.
⚠ Alert on merlin_store_durable == 0, not on != 1. A pod on a new
volume correctly reports -1, so an alert on != 1 fires on normal rollouts.
To turn the health and metrics endpoints off, set GALAHAD_OBSERVABILITY=0.
| Ray Serve support | not available. The vLLM and SGLang paths are the supported ones |
| Autoscaling, including scale to zero | not available |
| Multi-node rollout figures (hit rate, time to first token, recovery time) | ⚠ not published: no multi-node cluster measurement yet |
⚠ The single-node results on Performance do not carry over to a cluster. Rollout behaviour depends on your storage, your node count and your model size. Measure it on your own cluster, or ask for a measurement run there.
| Symptom | Cause | Fix |
|---|---|---|
| everything works but is slow | network backed storage | use a local disk volume; check the startup warning and merlin_storage_class
|
| cache empty after every restart |
emptyDir, or pod-lifetime NVMe |
both are wiped on restart; use a local PersistentVolume |
volume claim stays Pending
|
the claim is larger than the local disk | make the claim smaller than the disk |
| pod crash loops at startup | bad storage class or unwritable store directory | set GALAHAD_DISABLED=1 to get serving, then fix storage |
| pod killed while warming up | startup probe gives up too early | raise the startup probe's failureThreshold
|
| pod killed during a deploy while flushing | grace period too short for your record count | raise terminationGracePeriodSeconds: 327.6 ms at 1M records |
| liveness flapping | you added the cache to the liveness probe | do not; degraded is not dead |
- Install · vLLM Connector · SGLang Backend
- Multi Tenant Isolation: if pods serve different customers
- Performance