You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Installing the agent rules is now serverpilot mcp policy instead of a script in the repository. The rule text is rendered from one source inside the package rather than hand-copied into six places. --check only means something inside a source checkout, and says so plainly when run from an installed copy.
Nothing needs to be installed on a GPU server any more. The published serverpilot-collect command is gone. Observation had already become one fixed read-only SSH probe, and no configuration could make ServerPilot invoke a collector on the remote host, so the command shipped but could never be called. Hosts that already have it can uninstall it, and upgrading the control plane no longer means matching versions on every server.
The database no longer only grows. Beyond collection history, audit records, replay keys, resolved alerts and process sightings accumulated for the life of the install. Each kind now has a retention window (telemetry and replay keys one day, audit records 30 days, resolved alerts and process sightings seven), applied by the hourly pass that already ran, and freed space is returned to the disk once enough of the file is unused. Audit history is therefore kept for 30 days.
Servers no longer report a connection failure at the slightest disturbance, and no longer fail six at a time. One timeout was serving as both the SSH connect limit and the limit for the whole observation. A successful observation measured 2.3-2.9 seconds on average and up to 12.8 at the tail, against a budget of 8, so any network wobble pushed every host over the same line at once. Connecting and running the remote command now have separate budgets, collection reuses a persistent SSH connection (nothing is installed on the servers), and each server is probed on its own schedule. Measured on the same machine: one observation went from 2.3-2.9 seconds to 0.5-0.9, and the worst case from 12.8 seconds to 1.0.
A server that recovers is shown as online immediately, instead of a minute later. Three consecutive failures used to stop probing that server for 60 seconds, during which it kept reporting a connection failure even after it was fixed.
One server's trouble no longer becomes every server's trouble. All servers shared one collection pass and one deadline, so a stall anywhere failed all of them together. Each server now has its own clock. A server whose settings cannot be read is likewise only its own problem, instead of silently stopping the whole collection pass.
A failed connection now says what actually failed - authentication, a changed host key, a refused connection, an unreachable network, a name that does not resolve, a remote command that ran too long, a plugin failure - rather than one undifferentiated "connection failed".
One way to observe a server, instead of three.linux-nvidia, linux-host and server-script-v1 are replaced by linux. The probe already worked out for itself whether a machine has NVIDIA cards, so the choice only gave people a way to register a server that connects but reports no GPUs. Registered servers migrate automatically, and plugins for shared clusters are unaffected.
The log can be read. The daemon log now carries a UTC timestamp, the process id and the server it is about, and rotates. It used to be one ever-growing file with no timestamps at all.
Only one daemon can run on a machine. Two installations could claim the same port and restart each other indefinitely, and a retired startup item was only unloaded rather than removed, so it came back at the next login to claim the port again.
A server can no longer be "disabled" or "paused". Nothing could put a server into either state — the methods that wrote them were gone — yet the app and MCP still carried the vocabulary. monitor.status and gpu.state no longer include DISABLED or DRAINING, and the three database columns that never changed are dropped. Nothing on screen looks different, because neither state was reachable.
One claim, one record. Claiming used to write a request row first and issue a lease from it, so one act had two identities. That row existed for a queue, and across every request row this install ever wrote, not one was ever queued. The task name, purpose, constraints and duration now live on the lease itself, and the request table goes, taking GET/POST /api/v1/requests, POST /api/v1/requests/{id}/cancel, the serverpilot request command group, and the GET /metrics endpoint nobody read with it. The audit history is kept.
no_capacity now says what stood in the way. It used to be one sentence. It now carries a count per reason — held by another lease, held by occupancy, telemetry too old, and so on.
The Usage page loses a tile that was always zero. "Requested" counted queued requests, and nothing was ever queued.
The Windows desktop app is gone. It lagged the macOS app (no reassignment, no clearing an idle hold) and never managed the daemon's lifecycle; the CLI and MCP keep working on Windows.