Added
-
Anonymous community telemetry, on by default. The router sends one summary a day: which models ran, request and token counts, latency percentiles, and error counts by category. Never prompts, never completions, never raw error text, never your hostname. Opt out with
FLEET_NODE_TELEMETRY=false, or from the dashboard's new Community Telemetry section. A one-time notice is printed the first timeherd-nodestarts, before anything is ever sent, and the opt-out is honoured on that first run — a node that has opted out writes no files at all, not even the identifier. Every field that can be sent is published at ollamaherd.com/telemetry; that page is the contract, and the payload builder mirrors it under tests that fail the build if the two drift.Identity is a random
install_idUUID stored in~/.fleet-manager/install_id— delete it and you are a new install, delete it and opt out and you are gone. It is never derived from anything about the machine; tests assert it is neither the hostname in any form nor a hash of it, so it cannot be turned into a stable fingerprint later.Reporting is per herd, not per machine. The router is the only component that knows a fleet is one fleet, so it sends one payload containing a
devices[]row per node (chip, memory, cores, agent and Ollama version, that node's request share). Fleet totals are derived from those rows server-side rather than sent as scalars, so a total can never disagree with its parts. Each device carries adevice_idderived from itsnode_idbut hashed and salted with the herd'sinstall_id: stable enough to aggregate across days, and impossible to reverse to a hostname or correlate across herds. -
Name your herd for the public leaderboard — a second, separate opt-in on top of telemetry. Set it in the dashboard or via
FLEET_NODE_HERD_NICKNAME. Telemetry alone is never public; a nickname is the only field that is, and the UI says so before you set one. Names are limited to 30 characters of letters, numbers, spaces,-,_,., validated where they are typed rather than only in the browser. -
X-Fleet-Affinitymakes session routing visible. Every scored endpoint now reportsmatchedwhen a request went back to the node already holding that conversation, ornewwhen it did not — and omits the header entirely on routes that never score, because "did not apply" and "missed" are different facts.It reports the routing decision, not a cache hit, and deliberately so: Ollama folds llama.cpp's
cache_nback intoprompt_n(ollama/ollama#16428), soprompt_eval_countcannot yield a hit ratio and inferring one from it produced a false "zero prefix-cache reuse" report here once already. Time-to-first-token is the honest proof: a matched turn shows a large drop. -
usage.prompt_tokens_details.cached_tokensis reported on backends that actually measure it (MLX), and omitted — never zeroed — on backends that cannot (Ollama). Zero means "measured, nothing reused"; absent means "cannot measure". Conflating them is a bug vLLM shipped and SGLang still has. -
The Ollama and mlx-lm versions are collected in the heartbeat and surfaced to the router, so a fleet can answer "which runtimes are actually out there?" before a version-gated model or a changed default lands.
Changed
-
Session affinity now decays with queue depth. A pinned node contributes the full bonus when idle and progressively less as its queue grows, so a warm but saturated node stops winning while an idle peer sits free. Every production router does this — NVIDIA Dynamo, Ray Serve, SGLang — and a flat bonus is the bug vLLM's production-stack shipped
loadawarerouting to fix. Signal 3 already prices congestion, so this shrinks the bonus rather than adding a second penalty, and because decay can only shrink, affinity still never outweighs a thermal warning. -
MLX prompt cache raised from 4 to 10, matching
mlx_lm.server's own default. The previous value had no recorded rationale and sat below upstream, silently halving how many conversations could keep their KV cache warm — which is the actual limit on how many sessions affinity can honour.
Fixed
-
Telemetry now sends a missed day on startup instead of sleeping past it. The scheduler slept until the next 00:05 UTC before its first send, so start time of day decided whether an install reported at all: a router started at 00:10 UTC waited 23h55m, and one restarted daily after 00:05 never sent. Both are indistinguishable from "nobody uses this" on the receiving end. Found because our own fleet ran 12 hours of clean uptime with zero automatic sends.
-
The published telemetry opt-out worked in name only. Moving the sender from the node to the router silently repointed which environment variables it read (
ServerSettingscarries a different prefix), soFLEET_NODE_TELEMETRY=false— the opt-out documented on the website — stopped disabling anything, with no error anywhere, because "off" and "unset" are indistinguishable in a boolean default. Both that andFLEET_NODE_HERD_NICKNAMEare now read correctly, with tests pinning the names as a published contract. -
Dashboard toggles now survive a restart.
POST /dashboard/api/settingsonly mutated settings in memory, so every Feature Toggle silently reverted on the next start. They are now written to~/.fleet-manager/envwith a line-based writer that preserves the comments and hand edits in that file, and the response reportsnot_persisted/restart_requiredinstead of implying an effect it cannot deliver. Merely annoying forauto_pull; it would have been a broken promise for a telemetry opt-out. -
The gotomy.ai platform panel is removed from Settings — that service is shut down.
Removed
- The account-based platform telemetry opt-ins are untouched, but the dashboard's platform-connection UI is gone along with the service it targeted.