feat: prefix-cache-aware routing for distributed mode by localai-bot · Pull Request #10071 · mudler/LocalAI

localai-bot · 2026-05-29T20:09:02Z

Summary

Adds prefix-cache-aware routing to distributed mode (PostgreSQL + NATS) across all backends supported (text inferencing only). Today each request is routed to a loaded replica purely by load (in_flight ASC, last_used ASC, available_vram DESC), with no awareness of which replica already holds the relevant KV/prefix cache. This biases routing toward the cache-warm replica so backends reuse the prefix cache instead of recomputing it, as a load-guarded hint that never routes worse than today's round-robin.

Approach is frontend-side and engine-agnostic (no backend protocol changes), validated against the designs of SGLang, vLLM production-stack, Ray Serve, llm-d, AIBrix, and NVIDIA Dynamo.

How it works

A generic, isolated prefix tree (pkg/radixtree) maps cumulative prompt-prefix hashes to nodes (TTL, recency-weighted scoring, longest-prefix match).
core/services/nodes/prefixcache turns the rendered prompt into a deterministic xxhash chain (head-first chunking, so a multi-turn conversation or a volatile system-prompt tail still matches its stable head), keeps per-model trees, and makes a filter-then-score decision: narrow to load-eligible replicas (absolute + relative balance guards), then prefer the longest-prefix match above a min-match ratio, else cold-place onto the lowest-cache-weight node.
The decision feeds a preferredNodeID into the existing atomic FindAndLockNodeWithModel (SELECT ... FOR UPDATE), so the pick+lock+increment stays atomic and a stale/saturated/unhealthy preferred node transparently falls back to today's ordering.
Observations sync across frontends over NATS (prefixcache.observe/prefixcache.invalidate), mirroring the existing OpCache live-sync pattern. No JetStream.
The reconciler scales a model up on forced-disturb pressure (a strong cache match the load guard forced us off), bounded by MaxReplicas/capacity and the existing UnsatisfiableUntil backoff.
Per-model config (route_policy, min_prefix_match, balance_abs_threshold, balance_rel_threshold) via /api/nodes/scheduling + the React scheduling form; global TTL/enable via CLI flags.

Safety: round-robin is the floor

Affinity is only ever an additive, load-guarded hint. Four independent staleness layers (global TTL, live-registry validation in the lock, load guard, active eviction on unload/stale). When disabled, the entire feature is a byte-for-byte no-op.

Defaults / disabling

On by default in distributed mode (load-guarded, so safe). Worth a release-notes callout.
Fully disabled with --distributed-prefix-cache=false (env LOCALAI_DISTRIBUTED_PREFIX_CACHE=false): hook never set, provider/pressure nil, routing reverts to round-robin.

Follow-ups (tracked, not in this PR)

Deferred routing enhancements from the prior-art survey are tracked under epic #10063 (reported KV-event mode, multi-tier overlap, scorer pipeline, load-shaping, PD disaggregation, VTC fairness, minor tuning).

Test Plan

go test ./pkg/radixtree/ ./core/services/nodes/... ./core/services/messaging/ ./pkg/distributedhdr/ ./core/http/endpoints/localai/... (nodes + endpoints use Postgres testcontainers; green locally)
go test -race ./pkg/radixtree/ ./core/services/nodes/prefixcache/ (race-clean)
golangci-lint run --new-from-merge-base=master (0 findings)
Manual: with --distributed, confirm a multi-turn conversation stays on one replica and a saturating burst still spreads (round-robin floor)
Manual: --distributed-prefix-cache=false reverts to round-robin

Built with strict TDD across 10 phases; the round-robin-floor invariant, forced-disturb signal, and disabled no-op are each covered by tests.

🤖 Generated with Claude Code

localai-bot · 2026-05-29T21:48:48Z

Post-review hardening (12 commits)

A fresh-eyes multi-angle review of the full branch surfaced cross-cutting issues that the per-phase reviews did not. All but one are now fixed on this branch; the last was evaluated and intentionally dropped (rationale below).

Correctness

Autoscale over-scaling: Pressure.Count was non-draining, so one forced-disturb scaled up every tick within the window. Added Pressure.Reset to consume the signal after a pressure scale-up.
Autoscale overshoot: the busy-burst and pressure blocks could both fire in one tick on a stale current. Now at most one scale-up per reconcile tick (and no scale-down in a tick that scaled up).
Forced-disturb false positive: Decide searched the whole tree ignoring candidacy, so an offline/unloaded/excluded warm node produced a bogus saturation signal. Decide now drops a hot match whose node is not a current candidate.
Invalidation gap: invalidation was wired at only 3 router sites, missing scale-down, the probe reaper, the health-monitor reap, and the remote unloader. Added a SetReplicaRemovedHook chokepoint on the registry (RemoveNodeModel/RemoveAllNodeModelReplicas) wired to prefixSync.Invalidate, so every removal path invalidates by construction.
Partial-update footgun: a scheduling POST that omitted prefix-cache fields wiped them. The four prefix fields are now pointers with PATCH-style merge (omitted = preserved); existing fields keep their prior semantics.

Efficiency / dedup

GetModelScheduling was read 3x per request; now fetched once and threaded.
radixtree.Weight was O(N) per candidate (O(K·N) under lock); added single-walk WeightsFor.
ExtractChain allocated a hasher per block; now reuses one hasher (output byte-identical, verified).
Single source of truth: ModelConfig.ModelID() (was duplicated 2x) and prefixcache.ValidateThresholds (shared by the endpoint and Config.Validate); removed an unreachable branch in Select.

Evaluated and dropped: folding the LoadedReplicaStats read into the SELECT ... FOR UPDATE transaction. The current design reads stats unlocked (cheap, non-serializing) and locks a single replica row, so concurrent same-model requests proceed in parallel. Folding would require locking all of a model's replica rows to run selection in-transaction, serializing them - a throughput regression on hot models. Keeping the current two-read design.

Deferred follow-up: the forced-disturb autoscale signal is per-frontend (in-process counter); cluster-wide aggregation is tracked in #10083 (under the routing roadmap epic #10063).

Each fix was developed by a sub-agent owning a disjoint set of files (TDD), verified independently, and the integrated branch passes golangci-lint --new-from-merge-base=master with 0 issues and the full nodes/prefixcache/endpoints/radixtree/config suites green.

localai-bot · 2026-05-29T22:52:06Z

Second-pass review hardening (5 more commits)

A second fresh-eyes review pass (run after the first round of fixes) found issues the first pass missed - including two that the first round's invalidation fix had not fully resolved. All fixed here:

Invalidation chokepoint was incomplete. The registry hook fired only from RemoveNodeModel/RemoveAllNodeModelReplicas, but Register (same-ID re-register), MarkOffline, MarkDraining, and Deregister delete node_models rows in bulk and bypassed it - leaving stale prefix-index entries (worst case: a worker re-registering with the same ID stays a false hot-match until TTL). All four bulk-delete paths now fire the hook; the chokepoint is genuinely complete.
Invalidate empty-tree leak + broadcast amplification (introduced when the hook started firing for every model). Index.Invalidate no longer interns an empty radix tree for models that never used prefix-cache, and Sync no longer broadcasts a NATS invalidate when there was nothing to invalidate.
Router double-invalidate. With the registry hook live, the router's 3 explicit Invalidate calls were redundant (double tree-prune + double broadcast); removed.
Pressure consumed on a failed scale-up. scaleUp now reports success; the forced-disturb Pressure.Reset (and the MinReplicas ClearUnsatisfiable) only fire when a replica was actually scheduled, so a failed scale-up no longer wipes the signal.
Efficiency: Decide derives candidacy from the WeightsFor result (dropped a redundant map) and skips the sort for a single candidate; LoadedReplicaStats now SELECTs only the two columns the router reads.

Each fix was a disjoint-file sub-agent (TDD), verified independently; integrated branch passes golangci-lint --new-from-merge-base=master with 0 issues and the nodes (Postgres testcontainer) + prefixcache suites green.

Follow-ups noted (not dropped): explicit min_prefix_match=0 semantics (needs nullable columns), full PATCH/PUT endpoint coherence, longest-prefix-match restricted to candidate nodes, and a Pressure interface seam for the cross-frontend aggregation already tracked in #10083.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…empty-key guard Address review findings on the generic prefix tree: - Extract a shared pruneWalk helper parameterized by a shouldClear predicate and use it from Evict, Remove, and the MaxEntries path. Previously evictOldestLocked cleared a victim's value but never removed the now value-less node or its childless ancestors, so internal nodes accumulated under sustained churn at the cap. The MaxEntries path now prunes the victim and its empty ancestors. - DRY: pruneWalk replaces the duplicated logic in the former pruneLocked and Remove's inner closure. - Switch Tree.mu to sync.RWMutex; LongestMatch, Weight and Len take the read lock (RLock) while Insert, Evict and Remove keep the write lock. Confirmed race-clean under go test -race. - Document the strict greater-than TTL boundary on Options.TTL and expired: age exactly equal to TTL is still live. - Guard Insert against an empty key (no-op): the root never holds a value. Adds Ginkgo specs covering MaxEntries eviction, ancestor reclamation, the no-growth-past-cap invariant, the TTL boundary, and empty-key behavior for both Insert and LongestMatch. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…sort Address code-quality review findings in the prefixcache package. Correctness: ExtractChain now chunks from absolute offset 0 with fixed [0,W),[W,2W),... boundaries and caps the chain to the FIRST MaxDepth head blocks. The previous tail-keeping logic shifted the byte offset by a non-window amount once a conversation grew past MaxDepth*WindowBytes, changing every hash each turn and silently breaking cross-turn longest-prefix matching. The reusable KV/prefix cache lives at the head of the prompt, so anchoring at offset 0 makes the chain a true prefix-chain: P and P+suffix share their full leading overlap. Add a regression spec proving cross-turn stability past the cap. Performance: Index.Decide precomputes each candidate's Weight once (decorate-sort-undecorate) instead of calling the O(tree size) Weight inside the O(n log n) sort comparator. Behavior is unchanged. Lint: encode prev with binary.LittleEndian.PutUint64 instead of a manual byte loop, clearing the modernize rangeint finding. Also add a concurrent Decide/Observe/Invalidate spec to exercise Index's documented concurrency safety under go test -race. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…ain build Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…uting Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Add RoutePolicy and per-model balance/prefix-match override columns to ModelSchedulingConfig and include them in the SetModelScheduling upsert DoUpdates list so updates are not dropped on conflict. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Add a RoutePreference type and a new pref parameter so the atomic pick+lock+increment can be biased toward a preferred node without weakening atomicity. A nil preference reproduces the previous ORDER BY behavior exactly. Update the ModelRouter interface, both router.go call sites (pass nil for now; Phase 5 builds the real preference), the test doubles, and the distributed e2e caller. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Sync.Observe now returns whether the local index treated the assignment as new or extended, and Sync gains an Evict method that delegates to the wrapped index. Together these let SmartRouter hold a single prefixcache.Provider that broadcasts via NATS. Adds a compile-time Provider assertion and an Evict-delegates behavioral test. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

….Route Add a PrefixProvider + PrefixConfig to SmartRouterOptions/SmartRouter (nil keeps routing byte-for-byte the round-robin floor). On each request Route now calls buildPreference: it reads the prompt prefix chain from ctx (distributedhdr.PrefixChain), resolves the per-model policy/thresholds over the global config, loads candidate replica in-flight via a new registry read LoadedReplicaStats (deduped to one entry per node using the MIN in-flight across that node's replicas), asks the provider to Decide, and runs prefixcache.Select. The chosen node is passed as the RoutePreference to FindAndLockNodeWithModel on all three pick paths (cache hit, locked re-pick, cold scheduleAndLoad), and the served node is recorded via Observe only when the resolved policy is prefix_cache so round-robin models never pollute the tree. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

UnloadModel and both staleness fall-through paths in Route (after a failed gRPC probe and RemoveNodeModel) now call prefixProvider.Invalidate(model, nodeID), guarded by a nil-provider check so the round-robin floor is unchanged. At runtime the provider is the *prefixcache.Sync, so invalidations also broadcast to peer frontends. Adds a test that a previously hot prefix no longer Decides to a node after UnloadModel. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Add a concurrency-safe per-model rolling counter that tracks how many times a request had a usable hot prefix match but the load guard forced it off the warm node. Entries outside the window are dropped lazily on Count so the backing slice stays bounded. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Wire the rolling forced-disturb counter into the SmartRouter and the ReplicaReconciler. Router: in buildPreference, after Decide + Select, record a forced-disturb when a usable hot prefix match existed (d.HotNodeID != "" and d.MatchRatio >= cfg.MinPrefixMatch) but Select chose a different node (or nothing) because the load guard ruled the warm node out. This is the scale-worthy signal: the cache-warm replica is saturated. It deliberately does not fire for all-unique workloads (no hot match), avoiding false-positive scale-ups. Pressure is optional on SmartRouterOptions; nil keeps the path a no-op. Reconciler: read the same Pressure instance in reconcileModel as an extra scale-up reason, reusing the existing MaxReplicas + ClusterCapacityForModel guards and the UnsatisfiableUntil cooldown that gates the whole method. Pressure never overrides MaxReplicas and never force-evicts; a no-capacity model does not spin. Window and threshold come from prefixcache.Config (PressureWindow default 1m, PressureScaleThreshold default 1) and are configurable via ReplicaReconcilerOptions. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

The c.Name fallback to c.Model was duplicated in core/backend/options.go (feeding model.WithModelID) and hand-copied into core/backend/llm.go (the prefix-chain salt). These MUST agree or the prefix-cache salt diverges silently from the id the model loader tracks. Consolidate both into a new config.ModelConfig.ModelID() helper and call it from both sites. Behavior is identical. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

ExtractChain allocated a fresh xxhash.New() Digest per block (up to MaxDepth per call) and grew the chain slice without preallocation. Reuse a single Digest via Reset() before each block and preallocate the chain to min(nBlocks, MaxDepth). xxhash seed 0 is stateless, so Reset()+Write produces the byte-identical value to a fresh New()+Write. Output hashes are unchanged, preserving the cross-process determinism that peers rely on over NATS. Verified by capturing ExtractChain output for the existing test inputs before and after the refactor: identical. Existing extractor tests pass unchanged. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…; weigh cold candidates in one walk Index.Decide called radixtree.LongestMatch over the whole tree, so the deepest match could be a node that is offline, unloaded, or simply not in the passed candidate set. Honoring that as HotNodeID produced a false forced-disturb signal upstream (buildPreference records pressure when chosen != HotNodeID), making it look like a warm replica was load saturated when it was actually absent. Build the candidate set once and only set HotNodeID/MatchRatio when the matched node is an actual candidate; otherwise fall back to cold placement. A future refinement could ask the tree for the longest match restricted to the candidate nodes (shallower-but-valid) instead of dropping it. Also replace the per-candidate tree.Weight call in the cold-order sort with a single tree.WeightsFor walk, turning O(K*N) under the read lock into O(N + K). Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…back buildPreference always passes ColdOrder as a permutation of the full candidate set, so the cold-order loop hits every eligible candidate. The trailing best/bestIF scan was dead. Replace it with a plain "return """ and document that ColdOrder is guaranteed to cover all candidates, so "" means none were eligible. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

GetModelScheduling was read three times per request - in resolveSelectorCandidates, buildPreference, and nodeMatchesScheduling - three DB round-trips for one row that is immutable for the life of the request, and not a consistent snapshot. Fetch it once near the top of Route and thread the *ModelSchedulingConfig (may be nil) into all three helpers. scheduleNewModel keeps its own fetch since it runs outside the Route snapshot. Behavior is identical for nil sched. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Pressure.Count is non-draining (it prunes only by age), so a single burst of forced-disturbs stays within the rolling window for the whole window and keeps Count >= threshold on every reconciler tick. The reconciler will use Reset to clear a model's events after acting on the signal so a fresh scale-up requires fresh forced-disturbs to accumulate, rather than one burst driving the model toward MaxReplicas. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…sure Two autoscale bugs: 1. Over-scaling: the pressure scale-up block read Pressure.Count but never consumed it. With a non-draining counter a single forced-disturb burst kept Count >= threshold across the whole window, firing scaleUp on every tick and pushing the model toward MaxReplicas off one transient burst. After a successful pressure-triggered scale-up the reconciler now calls Pressure.Reset to consume the signal. 2. Double scale-up in one tick: the all-replicas-busy block and the pressure block could both fire in the same reconcileModel pass, each calling scaleUp(+1) against the same `current` read once at the top, so a model that was both busy and over threshold scaled +2 and could overshoot MaxReplicas by one. A scaledUp flag now enforces at most one scaleUp(+1) per tick: the pressure block is skipped if the busy block already scaled, and scale-down is skipped in any tick that scaled up. MinReplicas enforcement, UnsatisfiableUntil backoff, and capacity guards are unchanged. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…ation Add SetReplicaRemovedHook to NodeRegistry and fire it from both RemoveNodeModel and RemoveAllNodeModelReplicas after a successful delete. This is the single chokepoint every replica-removal path funnels through (router eviction, reconciler scale-down, probe reaper, health-monitor node-down reap, RemoteUnloaderAdapter), so the prefix-cache index can be invalidated by construction rather than wiring each call site individually. The hook is stored in an atomic.Pointer so the startup wiring (setter) and the request/reconcile-time fire are race-free; it is nil-safe when unset. GORM Delete reports no error for a no-op delete, so the hook also fires when nothing was removed; the consumer's Invalidate(model, node) is idempotent so this is harmless. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

… registry hook Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Extract ValidateThresholds into prefixcache/config.go so the per-model override validation (nodes.go endpoint) and Config.Validate share one implementation of the numeric bounds (min_prefix_match in [0,1], balance_abs_threshold >= 0, balance_rel_threshold == 0-or->= 1) instead of hard-coding them in two places. The route_policy allow-list stays explicit (not ParsePolicy, which maps typos to Default). Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

A scheduling POST that omitted route_policy/thresholds (e.g. a min_replicas-only update) full-replaced every column and silently reset the model's previously-configured prefix-cache settings to empty/zero. Make the four prefix-cache request fields pointers so omitted is distinguishable from explicit zero, and merge PATCH-style in SetSchedulingEndpoint: a provided pointer wins, an omitted one preserves the existing config value (zero default when none). Non-prefix fields keep their full-replace PUT semantics. Validation now runs on the resolved values via prefixcache.ValidateThresholds. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…p empty broadcasts A registry chokepoint fires Sync.Invalidate(model, nodeID) for every replica removal of every model, including round-robin models that never used the prefix cache. Index.Invalidate previously called tree(model), which lazily created and permanently retained an empty radix tree for any model that ever lost a replica, growing the trees map without bound. Sync.Invalidate also published a NATS PrefixCacheInvalidateEvent on every call, amplifying no-op removals across the cluster. Index.Invalidate now looks the tree up read-only via existingTree and returns without allocating when none exists. The Provider interface is unchanged; Sync gates the broadcast through an optional invalidateExisting(bool) capability type-asserted from the wrapped Index, falling back to the prior always-broadcast behavior for other Provider implementations. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…rivial sort WeightsFor already returns a map keyed by every requested candidate, so the separate candidates set built to validate the hot match was redundant: a node is a candidate iff it is a key in the weights map. Drop the extra map and gate the hot-match check on weights membership. Also skip the sort when there is at most one candidate, since the input order is already the cold order. Behavior is unchanged. Deferred follow-up: skipping the WeightsFor walk entirely when a hot match wins would need lazy cross-file changes and is out of scope here. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…im LoadedReplicaStats columns Bulk node-scoped node_models deletes (Register re-register cleanup, MarkOffline, MarkDraining, Deregister) removed rows directly without firing the replica-removed hook, so the prefix-cache index kept pointing at nodes whose models were gone. Capture the DISTINCT model names before each bulk delete and fire fireReplicaRemoved once per model after a successful delete, restoring the single-chokepoint invariant for all removal paths. The pre-query is skipped when no hook is set so the no-hook path stays cheap. Also narrow LoadedReplicaStats to SELECT only node_id and in_flight (the only fields the router consumer reads), dropping the JOIN-side available_vram fetch and unused columns while keeping the []ReplicaCandidate return type unchanged. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

scaleUp was fire-and-forget (void) yet its callers unconditionally consumed the pressure signal (Pressure.Reset) and the MinReplicas hysteresis (ClearUnsatisfiable) right after calling it. If scaleUp added nothing (ScheduleAndLoadModel errored, or no node could be loaded) the saturated warm replica got no new replica AND its accumulated forced-disturb history was wiped, forcing the signal to re-accumulate over a full PressureWindow before the next attempt. Make scaleUp return whether at least one replica was actually scheduled, and gate the side effects on it: - pressure block (2b): set scaledUp and call Pressure.Reset only on success; on failure preserve the signal so the next tick retries off the same accumulated pressure. - busy-burst block (2): set scaledUp from the return value so a failed attempt does not suppress the pressure path or scale-down. - MinReplicas block: call ClearUnsatisfiable only on success so a failed attempt does not reset the unsatisfiable counter. All existing invariants (MaxReplicas, capacity gating, UnsatisfiableUntil cooldown, at-most-one-scale-up-per-tick) are preserved. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

The NodeRegistry removal chokepoint (RemoveNodeModel / RemoveAllNodeModelReplicas) now fires SetReplicaRemovedHook, which invalidates the prefix-cache index. The router was also calling prefixProvider.Invalidate explicitly right after each registry removal on the two stale-replica health-probe fall-throughs in Route and in UnloadModel, so every router-side eviction invalidated twice (double tree-prune + double NATS broadcast). Remove the three redundant explicit Invalidate calls and their empty nil-guards. Each removed call sat immediately after a registry removal that fires the hook, so invalidation is preserved via the chokepoint. Decide/Observe usage is untouched. Re-point the unit test (fake registry fires no hook) to assert the removal chokepoint is exercised on unload instead of the router's direct invalidation. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…rontend coherence Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…ould panic) Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

MarkOffline, MarkDraining, and the Register re-register cleanup ran the nodeModelNames SELECT and the bulk node_models DELETE as two separate statements on r.db with no transaction. A SetNodeModel landing between the two was deleted but its replica-removed hook never fired, leaving the prefix-cache index pointing at a removed replica until TTL or candidacy self-heal. Wrap the capture and the delete in a single db.Transaction in each path (mirroring how Deregister already does it). The captured model names are collected into a slice declared outside the closure; the replica-removed hook fires for each only after the transaction commits, so a rollback never invalidates the index for a removal that did not persist. The set of fired hooks now equals exactly the set of node_models rows actually deleted, with no interleaving gap. The status flip in MarkOffline/MarkDraining (setStatus) is a separate, pre-existing operation and routing already filters non-healthy nodes, so it stays outside the transaction; return contracts are unchanged. Deregister was already correct and is untouched. The cheap-path skip (no hook -> skip the SELECT) is preserved. Adds a spec asserting MarkOffline fires hooks for exactly the rows it deletes and leaves no node_models row behind (consistent snapshot). Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…servations Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Insert recorded the value (node id) only on the final node of the key chain, leaving every intermediate prefix node valueless. LongestMatch returns the deepest node that hasValue, so two chains that share a leading block but diverge in the tail never matched: only exact-repeat queries hit. That broke the prefix-cache routing core use cases (shared system prompt, multi-turn extension, volatile tail), all of which rely on prefix matching rather than exact-repeat. Set value/hasValue/lastSeen at every node along the chain so each prefix-block node remembers the node id that served that prefix (SGLang/vLLM-style). The deepest match wins, and the last writer owns a shared prefix node (a recency heuristic: the most recent chain through a block is the one most likely still warm). size now counts valued nodes, which is the intended meaning. Updated radixtree tests to the new semantics: deepest-prefix test uses non-overlapping chains, a new test asserts last-writer-owns-shared-node, Evict/Remove/MaxEntries expectations recomputed for per-prefix-node counting, and a shared-prefix LongestMatch red test added. Added a prefixcache Decide test proving a prefix-only query routes to the warm node. No prefixcache .go logic changed. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Add a DB-backed e2e spec that drives SmartRouter against a real NodeRegistry (Postgres testcontainer) and the real prefixcache.Index radix-tree provider, using a fake gRPC backend factory so no real inference runs. Covers the five behaviors validated by hand: 1. Cold miss + observe: an unseen prefix chain cold-places and is recorded. 2. Hot-match affinity: the same chain returns to its warm node X. 3. Shared-prefix match: a divergent chain sharing X's leading prefix still routes to X (the radix-tree regression we fixed). 4. Negative control: an unrelated chain is a cold miss, not a false hot match on X. 5. Failover + invalidation: removing X's replica fires the registry chokepoint hook to invalidate the prefix entry, and the chain fails over to surviving node Y and re-homes there. Replaces the need for manual docker-compose re-runs. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Track prefix-cache affinity per loaded replica (a backend process with its own KV cache) instead of per node, so multiple replicas of the same model on one node each keep distinct affinity and a hot prefix routes back to the exact replica that served it. - radixtree: add RemoveFunc(pred) and reimplement Remove on top of it. - prefixcache: introduce ReplicaKey{NodeID, Replica}; Index/Candidate/ PrefixDecision/Select/Provider now key on ReplicaKey. Add InvalidateNode to drop every replica of a node; Invalidate drops one replica. Select returns (ReplicaKey, bool) and gains a deterministic least-in-flight eligible fallback (tiebreak NodeID then Replica). - messaging: carry Replica on PrefixCacheObserveEvent and PrefixCacheInvalidateEvent (Replica < 0 means all replicas of the node). - Sync delegates + broadcasts with replica; InvalidateNode broadcasts Replica=-1; ApplyInvalidate routes negative replica to InvalidateNode. This is part 1 of 2; the registry/router/wiring consumers are updated separately. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Wire the SmartRouter, NodeRegistry, and distributed startup to the replica-keyed prefixcache API. Affinity is now tracked per replica (each replica is a separate process with its own KV cache), so a prefix served by (node,0) no longer leaks onto the same-node sibling (node,1). - RoutePreference gains PreferredReplica; FindAndLockNodeWithModel locks the EXACT (node_id, replica_index) row, falling through to the default ORDER BY when that replica is not loaded. - SetReplicaRemovedHook now carries replicaIndex; RemoveNodeModel fires the specific replica, RemoveAllNodeModelReplicas and the four bulk node-scoped deletes fire replica<0 (all replicas of the node). - buildPreference builds one Candidate per loaded replica and locks the exact replica the policy chose; observePrefix records the served ReplicaKey at every call site. - distributed.go routes the hook to InvalidateNode (replica<0) or Invalidate(key). - Tests updated to the replica-keyed API plus new coverage: a hot prefix on (node,0) prefers replica 0 over the same-node sibling (router unit + e2e), and FindAndLock locks the exact preferred replica. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

…plate models Prefix-cache-aware routing built its prompt-prefix chain from the rendered prompt string `s` in ModelInference. For models with TemplateConfig.UseTokenizerTemplate the frontend never renders a prompt - the backend tokenizes the structured messages itself - so `s` is empty, the chain is empty, and routing silently falls back to round-robin. That covers the bulk of modern chat models (qwen3, llama3, ...), so the feature effectively never engaged for them. Fall back to messagesPrefixSource(messages): a deterministic, prefix-stable head-first serialization of the conversation (role + content per turn). Two requests sharing a leading system prompt and early turns share a leading byte prefix, which ExtractChain maps to a shared chain prefix - landing both on the same cache-warm replica. The rendered `s` is still preferred when present (higher fidelity for non-template models). Found via the multi-replica-per-node e2e: zero "prefix-cache routing decision" logs despite per-request Route calls, traced to the empty-chain guard. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

Add a routing-and-caching roadmap section to the distributed-mode guide, linking the epic (#10063) and the follow-up issues (#10064-#10070) surfaced from a survey of SGLang, vLLM production-stack, Ray Serve, llm-d, AIBrix, and NVIDIA Dynamo. Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

localai-bot mentioned this pull request May 29, 2026

Distributed routing: cluster-wide forced-disturb pressure (per-frontend today) #10083

Open

mudler added 27 commits May 30, 2026 10:30

feat(radixtree): generic prefix tree skeleton with longest-match

c52a44d

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(radixtree): Insert with path recency refresh and entry cap

f1e2e78

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(radixtree): TTL idle-expiry and Evict sweep with branch pruning

8ff0935

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(radixtree): recency-weighted per-value Weight

9131c55

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(radixtree): Remove all entries for a value

1d3977f

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

test(radixtree): race-free concurrency smoke test

288e05c

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): RoutePolicy enum with parse/resolve

20c53cd

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): Config with defaults and validation

f8373ac

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): deterministic xxhash prefix-chain extractor

3892d10

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): pure filter-then-score replica selection

4f7be8b

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): Provider interface and radix-tree-backed Index

e0c8bf5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

style(prefixcache): gofmt policy enum comment alignment

b4b2595

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(messaging): prefixcache observe/invalidate subjects and payloads

1fcd09a

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(prefixcache): NATS sync publish/apply for observe and invalidate

e59a67a

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(distributedhdr): ctx carrier for prefix-hash chain

732d877

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(distributedhdr): PrefixChainHook indirection for backend-side ch…

5aed51c

…ain build Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

feat(backend): stash prompt prefix chain on ctx before distributed ro…

1f368e8

…uting Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

fix(backend): mirror modelID fallback for prefix-chain salt parity

2493a5f

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

mudler added 26 commits May 30, 2026 10:30

feat(distributed): invalidate prefix-cache on any replica removal via…

8925c8a

… registry hook Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

fix(prefixcache): broadcast invalidations unconditionally for cross-f…

cd4d6ad

…rontend coherence Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

fix(prefixcache): reject TTL<=0 in Config.Validate (eviction ticker w…

9841386

…ould panic) Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

chore(nodes): debug logging for prefix-cache routing decisions and ob…

5affd48

…servations Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

mudler force-pushed the feat/prefix-cache-routing branch from 039c96d to 058b1b6 Compare May 30, 2026 13:20

mudler added the enhancement New feature or request label May 30, 2026

mudler merged commit a44bdb2 into master May 30, 2026
59 checks passed

mudler deleted the feat/prefix-cache-routing branch May 30, 2026 21:24

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

feat: prefix-cache-aware routing for distributed mode#10071

feat: prefix-cache-aware routing for distributed mode#10071
mudler merged 61 commits into
masterfrom
feat/prefix-cache-routing

localai-bot commented May 29, 2026 •

edited by mudler

Loading

Uh oh!

localai-bot commented May 29, 2026

Uh oh!

localai-bot commented May 29, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Uh oh!

Conversation

localai-bot commented May 29, 2026 • edited by mudler Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

How it works

Safety: round-robin is the floor

Defaults / disabling

Follow-ups (tracked, not in this PR)

Test Plan

Uh oh!

localai-bot commented May 29, 2026

Post-review hardening (12 commits)

Uh oh!

localai-bot commented May 29, 2026

Second-pass review hardening (5 more commits)

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

localai-bot commented May 29, 2026 •

edited by mudler

Loading