Skip to content

Kubernetes

Sietse edited this page Sep 29, 2026 · 2 revisions

Running on Kubernetes

Probes, metrics, an off switch, and the storage mistake that silently makes every read about 17 times slower.

Status ✅ Works
What you set up a local disk volume for the store, the Galahad settings, three health probes, and a metrics scrape
Needs a vLLM or SGLang image with the Galahad package installed, and a licence (Install)

In plain words

On Kubernetes, things restart constantly: deploys, autoscaling, spot reclaims, node drains. Several times a day, by design.

That makes durable memory worth much more here than on a single server. On Kubernetes the platform restarts things for you, and Galahad means a restarted pod does not start empty.

The everyday example

A rollout replaces every pod at lunchtime. Without this, every pod starts empty and the next hour is slow for everyone. With it, the new pods read the same memory the old ones built.


The storage default that costs you

Kubernetes' default storage is usually network backed, and it is about 17 times slower.

local disk, 580 MB block 44 ms
network volume, same block 780 ms

⚠ The pod stays healthy. The cache works and everything looks correct, but every read costs seventeen times what it should. Galahad prints a warning at startup, and the merlin_storage_class metric shows the class (see Metrics); check for both.

Two more storage choices that lose the cache:

Trap What happens
emptyDir erases the cache on every restart
pod-lifetime NVMe fast, looks local, and is wiped on every restart

⭐ Give the store a local disk that outlives the pod, for example a Kubernetes local PersistentVolume on the node's NVMe:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: galahad-local-nvme
provisioner: kubernetes.io/no-provisioner
volumeBindingMode: WaitForFirstConsumer
reclaimPolicy: Retain
---
apiVersion: v1
kind: PersistentVolume
metadata:
  name: galahad-local-nvme-node1
spec:
  capacity:
    storage: 340Gi                 # must not be larger than the disk
  accessModes: [ReadWriteOnce]
  persistentVolumeReclaimPolicy: Retain
  storageClassName: galahad-local-nvme
  local:
    path: /mnt/disks/galahad-nvme
  nodeAffinity:
    required:
      nodeSelectorTerms:
        - matchExpressions:
            - key: kubernetes.io/hostname
              operator: In
              values: ["<your-node-name>"]

⚠ Size the claim to the disk. A claim larger than the local disk never binds: the pod stays Pending, while the same size on network storage works at once. That makes local disk look broken when it is not.

Mount the volume and point Galahad at it:

env:
  - name: GALAHAD_CACHE_DIR
    value: /data/galahad
  - name: GALAHAD_TENANCY
    value: single                  # or multi, when pods serve different customers
  - name: GALAHAD_LICENCE_FILE
    value: /etc/galahad/licence.lic
volumeMounts:
  - name: galahad-store
    mountPath: /data/galahad

GALAHAD_TENANCY accepts only single or multi. Any other value is refused with an error, so a typo cannot silently turn the cache off.


The probes

With vLLM, the Galahad connector serves three health endpoints on port 8081 (set GALAHAD_HEALTH_PORT to change it):

Path Use it for What it checks
/health/started startupProbe startup has finished, including reading the store's index
/health/ready readinessProbe the server can serve. Stays ready with a degraded cache
/health/live livenessProbe the server process is alive. Does not check the cache
ports:
  - { name: health,  containerPort: 8081 }
  - { name: metrics, containerPort: 9090 }
startupProbe:
  httpGet: { path: /health/started, port: health }
  periodSeconds: 5
  failureThreshold: 60           # a large store takes longer to read at startup
readinessProbe:
  httpGet: { path: /health/ready, port: health }
  periodSeconds: 10
livenessProbe:
  httpGet: { path: /health/live, port: health }
  periodSeconds: 20
  failureThreshold: 3

⭐ Liveness does not include the cache. If the cache is unavailable, the pod is degraded, not dead, so it keeps serving instead of being restarted.

⚠ Readiness never reports "healthy" when it cannot tell.

⚠ Give the startup probe enough time. A pod that holds a large store needs longer to read its index at startup. If the startup probe gives up too early, the pod is killed while it is still warming, and it looks like a crash loop.


GALAHAD_DISABLED=1: the off switch

env:
  - name: GALAHAD_DISABLED
    value: "1"

Galahad starts switched off and the server serves normally, without memory. A bad storage class or an unwritable store directory cannot put your pod into a crash loop, and you can turn Galahad off without rebuilding an image. An empty value or 0 means Galahad is on.


Shutdown timing, measured

merlin_shutdown() writes pending records to disk:

Records Flush
10k 1.4 ms
100k 18.8 ms
1M 327.6 ms

The Kubernetes default of terminationGracePeriodSeconds: 30 is plenty: the flush is about 1.6% of that at a million records.

⚠ Skipping shutdown does not lose saved memory. It loses the usage counters that decide what is loaded first after a restart.


Metrics

With vLLM, the Galahad connector serves Prometheus metrics at /metrics on port 9090 (set GALAHAD_METRICS_PORT to change it), with no prometheus_client dependency, so nothing extra goes into your image.

Metric names start with merlin_. The ones to watch:

Metric What it tells you
merlin_backend_attached 1 Galahad is attached. 0 it is installed but doing nothing
merlin_storage_class the storage class of the block folder, as a class label. Anything but local is much slower
merlin_store_durable 1 the store survived on this node. 0 it was wiped (pod-lifetime storage). -1 unknown, for example a new volume or a pod on a new node
merlin_hits_total · merlin_misses_total memory found and not found
merlin_write_failures_total saves that failed

⭐ Scrape merlin_store_durable in particular: it tells you whether the store you think is durable actually is.

⚠ Alert on merlin_store_durable == 0, not on != 1. A pod on a new volume correctly reports -1, so an alert on != 1 fires on normal rollouts.

To turn the health and metrics endpoints off, set GALAHAD_OBSERVABILITY=0.


Not available in this release

Ray Serve support not available. The vLLM and SGLang paths are the supported ones
Autoscaling, including scale to zero not available
Multi-node rollout figures (hit rate, time to first token, recovery time) ⚠ not published: no multi-node cluster measurement yet

⚠ The single-node results on Performance do not carry over to a cluster. Rollout behaviour depends on your storage, your node count and your model size. Measure it on your own cluster, or ask for a measurement run there.


Troubleshooting

Symptom Cause Fix
everything works but is slow network backed storage use a local disk volume; check the startup warning and merlin_storage_class
cache empty after every restart emptyDir, or pod-lifetime NVMe both are wiped on restart; use a local PersistentVolume
volume claim stays Pending the claim is larger than the local disk make the claim smaller than the disk
pod crash loops at startup bad storage class or unwritable store directory set GALAHAD_DISABLED=1 to get serving, then fix storage
pod killed while warming up startup probe gives up too early raise the startup probe's failureThreshold
pod killed during a deploy while flushing grace period too short for your record count raise terminationGracePeriodSeconds: 327.6 ms at 1M records
liveness flapping you added the cache to the liveness probe do not; degraded is not dead

Related

Clone this wiki locally