Repository navigation
Releases: JinPLu/ServerPilot
Release list
v3.0.0
- Installing the agent rules is now
serverpilot mcp policyinstead of a script in the repository. The rule text is rendered from one source inside the package rather than hand-copied into six places.--checkonly means something inside a source checkout, and says so plainly when run from an installed copy. - Nothing needs to be installed on a GPU server any more. The published
serverpilot-collectcommand is gone. Observation had already become one fixed read-only SSH probe, and no configuration could make ServerPilot invoke a collector on the remote host, so the command shipped but could never be called. Hosts that already have it can uninstall it, and upgrading the control plane no longer means matching versions on every server. - The database no longer only grows. Beyond collection history, audit records, replay keys, resolved alerts and process sightings accumulated for the life of the install. Each kind now has a retention window (telemetry and replay keys one day, audit records 30 days, resolved alerts and process sightings seven), applied by the hourly pass that already ran, and freed space is returned to the disk once enough of the file is unused. Audit history is therefore kept for 30 days.
- Servers no longer report a connection failure at the slightest disturbance, and no longer fail six at a time. One timeout was serving as both the SSH connect limit and the limit for the whole observation. A successful observation measured 2.3-2.9 seconds on average and up to 12.8 at the tail, against a budget of 8, so any network wobble pushed every host over the same line at once. Connecting and running the remote command now have separate budgets, collection reuses a persistent SSH connection (nothing is installed on the servers), and each server is probed on its own schedule. Measured on the same machine: one observation went from 2.3-2.9 seconds to 0.5-0.9, and the worst case from 12.8 seconds to 1.0.
- A server that recovers is shown as online immediately, instead of a minute later. Three consecutive failures used to stop probing that server for 60 seconds, during which it kept reporting a connection failure even after it was fixed.
- One server's trouble no longer becomes every server's trouble. All servers shared one collection pass and one deadline, so a stall anywhere failed all of them together. Each server now has its own clock. A server whose settings cannot be read is likewise only its own problem, instead of silently stopping the whole collection pass.
- A failed connection now says what actually failed - authentication, a changed host key, a refused connection, an unreachable network, a name that does not resolve, a remote command that ran too long, a plugin failure - rather than one undifferentiated "connection failed".
- One way to observe a server, instead of three.
linux-nvidia,linux-hostandserver-script-v1are replaced bylinux. The probe already worked out for itself whether a machine has NVIDIA cards, so the choice only gave people a way to register a server that connects but reports no GPUs. Registered servers migrate automatically, and plugins for shared clusters are unaffected. - The log can be read. The daemon log now carries a UTC timestamp, the process id and the server it is about, and rotates. It used to be one ever-growing file with no timestamps at all.
- Only one daemon can run on a machine. Two installations could claim the same port and restart each other indefinitely, and a retired startup item was only unloaded rather than removed, so it came back at the next login to claim the port again.
- A server can no longer be "disabled" or "paused". Nothing could put a server into either state — the methods that wrote them were gone — yet the app and MCP still carried the vocabulary.
monitor.statusandgpu.stateno longer includeDISABLEDorDRAINING, and the three database columns that never changed are dropped. Nothing on screen looks different, because neither state was reachable. - One claim, one record. Claiming used to write a request row first and issue a lease from it, so one act had two identities. That row existed for a queue, and across every request row this install ever wrote, not one was ever queued. The task name, purpose, constraints and duration now live on the lease itself, and the request table goes, taking
GET/POST /api/v1/requests,POST /api/v1/requests/{id}/cancel, theserverpilot requestcommand group, and theGET /metricsendpoint nobody read with it. The audit history is kept. no_capacitynow says what stood in the way. It used to be one sentence. It now carries a count per reason — held by another lease, held by occupancy, telemetry too old, and so on.- The Usage page loses a tile that was always zero. "Requested" counted queued requests, and nothing was ever queued.
- The Windows desktop app is gone. It lagged the macOS app (no reassignment, no clearing an idle hold) and never managed the daemon's lifecycle; the CLI and MCP keep working on Windows.
v2.4.0
- A card that is running a job is no longer read as empty because one collection failed to list its processes. On a containerized server nvidia-smi sometimes cannot see its own compute processes, and that single reading used to be enough.
- A staged job's gap between two batches of work no longer costs it its cards. gpu_status(lease_id=…) now doubles as the holder's heartbeat, so agents need change nothing. While something in the claim is still running, cards it does not use are returned one by one as before; when nothing is running anywhere, its cards are returned only once the holder has also gone quiet.
- A claim that declared a duration (the CLI and App paths) used to be exempt from observed-idle reclaim entirely. It is now treated like any other claim: cards that stay idle while the holder stays silent past the reclaim window are returned. The declared duration still decides when the lease expires.
- A card does not become claimable the instant its process ends: it waits until a minute of unbroken, healthy collections have all left the process out. A server outage does not count toward that minute — it restarts it, rather than waiting the process out on its behalf.
- Clearing an idle hold by hand no longer releases a lease whose holder was working moments ago: while a process or heartbeat has been seen recently the request is refused and says how long to wait, or to let the holder release it. ServerPilot's own occupancy holds, and leases on a server that stopped answering, are still reclaimable at once.
- The macOS app now reads releasability from the control plane instead of inferring it from GPU states, and no longer offers a button that would be refused.
- A card gets a ten-minute cooldown after its workload lease is released before occupancy can claim it again.
- The occupancy stop confirmation now says it is a whole-server switch, how many cards it stops together, and that claiming GPUs already frees the cards an agent needs.
v2.3.0
The numbers you see are the numbers you can claim, and a card's state reads at a glance.
- CPU and memory now show what the machine actually gives this endpoint. A containerized server reports the whole node's cores and memory, so the interface read "128 cores, 974 GB, 121 free" for an endpoint that only owns 60 cores and 480 GB — sizing parallelism off that number oversubscribes the job. The GUI, MCP, and the allocator now read one resolved figure, labelled as a container quota where it comes from a cgroup budget; CPU and memory requests are admitted against that same budget instead of the node's core count.
- A network blip no longer makes a healthy server show as unreachable. A single failed SSH probe used to turn a server red immediately, even one that had reported in seconds earlier; it now takes a sustained silence to call it unreachable. A disabled or draining server is also no longer shown as a connection failure.
- An agent can now tell "held but idle" apart from "actually running a job." The MCP status values existed in code but were never explained in words, so an agent could see a card was held but not whether anything was computing on it.
v2.2.0
Cards can be claimed, released, and deleted, and the interface stops stalling.
- A server that stopped answering no longer blocks an apply. The cards it held itself still counted as reclaimable capacity, so an apply picked that host, spent its whole budget on an SSH timeout and failed, while other machines in the group had free cards all along.
- A server you can no longer reach can now be deleted, and the cards it holds released. Releasing required a successful collection and deleting required the leases released first — both asking a machine that no longer exists for proof, so the cards stayed locked in the ledger forever. The release is recorded as "server unreachable; operator settled the ledger". A host that still answers is refused exactly as before.
- The interface and the CLI no longer stall on a cycle. Every collection cycle restarted a batch of subprocesses to ask each plugin what it was, enough to stop the whole service for a second or two. The server list went from 1.7s to 0.05s.
- A multi-card apply is no longer killed by a 20-second timeout. The wait is now the server's real budget rather than a guess from the card count, and applies pinned to different hosts or groups no longer block each other.
- Registering a server now connects to it on the spot and tells you what happened. Registration used to be a database write that answered "added", so a host that can never connect was first reported as a new server and then became an unexplained red row. It now says plainly: connected and how many cards were found, or not connected and why.
- An agent can now register a server into a group, and pick the right observation profile.
gpu_add_server/gpu_update_serverdid not acceptserver_group_id, so a server registered through them could never be selected; the default profile also named the most special case, which connects to an ordinary machine and reports no cards. - The state page no longer claims occupancy is running when it cannot see the card. It used to replay the last recorded "running"; it now says that a card it cannot see cannot confirm what is on it.
- Clearing an "idle" hold now tells you when those cards last ran a compute process. A job between two batches looks exactly like one that has ended, and clearing the wrong lease wedges its cards. Only processes started after you took the cards count; the judgement stays yours.
- A lease's telemetry reflects only the load during your own hold. The ten-minute average used to start before the claim, so sizing a batch from it meant sizing it against someone else's work.
- Deleted servers and released tasks no longer leave warnings behind. 21 of the 23 live warnings pointed at resources that no longer existed, burying the ones that mattered.
- Re-registering a machine on a new port no longer lets two registrations hold the same card. The registration that can still see the card keeps it; the old one steps aside.
- The two server-management tools can finally be called the way they are documented. They demanded three parameters the instructions never mention, so following the documentation could only fail. The CLI also receives the structured detail behind a failure now, such as the list of groups to choose from.
v2.1.0
Upgrading once now finishes the job: macOS carries a single backend, a reinstall moves the running control plane onto it, and a release is one pushed tag.
- The macOS app no longer carries a second backend; installing the CLI is installing the backend. The app used to bundle a whole Python runtime, and opening it pinned the background service to that copy.
uv tool install --force .then replaced the package on disk while the running process stayed on the app's older code, so a backend change survived any number of reinstalls. The app is now only the interface. If you never installed the CLI separately, the app names the command to run rather than quietly failing to start. Building the app also drops from minutes to seconds. - A reinstall now moves the running control plane onto the new version. Deciding whether a process was stale relied on a hand-maintained capability list, so the same version number could sit in front of an older process indefinitely with nothing to show for it. A version or capability mismatch now restarts it.
- MCP says so when it reaches a control plane that has not been restarted. A new tool hitting an older API used to return an ordinary HTTP error that gave no hint an upgrade was pending. It now names both versions, the missing capabilities, and the restart.
- Releasing is one pushed tag. Pushing the tag creates the GitHub Release, and the Windows archive is uploaded as before. If the version has been bumped and the changelog entry sealed but the tag is missing, CI now fails and names the next step —
2.0.0sat in exactly that state: commits pushed, CI green, no tag and no release, with nothing to catch it. serverpilot doctorlists the versions in play. Control plane, local CLI, MCP entry point, and each server's collector, and it names the next step — restart the control plane, or reinstall the collector on a specific machine — when they disagree. That was previously a per-machine guess.- The server collection entry point reports its own version. An older script that omits the field keeps working, so a partially upgraded fleet does not start refusing allocations.
- A local plugin on the wrong contract no longer disappears without a word. A plugin still on the old protocol was skipped during discovery and looked simply absent; it now surfaces as an explicit error in
doctor. - History charts stop drawing the same card twice. A server whose GPUs changed — a rebuilt container, for instance — leaves the old cards in the database, and history returned them too, so an eight-card machine drew sixteen lines and every card appeared twice in the legend. History now returns only the cards currently on the machine, and a line with no data in the selected window no longer takes a legend slot.
- The server detail sheet is no longer sparse and over-tall. The group card stacked every field into one column and left most of the width empty; it now fills the width with the same layout the host card uses, with long paths and notes on their own rows.
- Two easily confused card counts were renamed. "At most 4 cards" and "per-lease ceiling 8 cards" used to sit side by side and read as a contradiction. They are now the number available on this machine right now (which moves) and the ceiling one apply can reach (which does not); when they agree only one is shown, and an unknown ceiling is left blank rather than invented.
- An upgrade checklist collects the steps that cannot be automated. Confirming the control-plane process actually changed version, reconnecting MCP in the client so the tool list refreshes, reinstalling the agent policy, and upgrading the collection entry point and occupancy helper on each server.
- When the local control plane is up,
serverpilot daemon statusreports it live instead of following a system proxy orHTTP_PROXYinto a false dead state. CLI, MCP, and the Windows desktop bridge no longer send 127.0.0.1 through an environment or OS HTTP proxy. On macOS, Python still reads the System Configuration proxy when the process has no proxy variables; those probes used to go through it, so curl could succeed while the CLI said the service was down. gpu_status(server_id=…)no longer includes groups you did not ask about, and no longer reports those groups as a 0-card ceiling. The rule that a group without per-card inventory must still appear was also matching groups that were only out of scope for this call; an empty member set then aggregated the one-apply limit to 0, so an agent or the desktop would treat a cluster with free cards as empty. A delegated cluster still appears when the query is not narrowed — it has a member, just not per-card rows. A direct group that is genuinely full still reportslargest_allocatable_block0.
ServerPilot 2.0.0
Source is now the daily product: a local daemon, the desktop apps, and five MCP tools. The browser UI and the scheduler/planning surfaces are gone; the claim loop you were using is not. And there is one kind of cluster — a plugin-adapted cluster is no longer a parallel concept, so the same experiment can finally be compared across both.
- A plugin-adapted cluster (Slurm, for example) appears in
gpu_statusas an ordinary group, with its name, shared workspace, GPU model, and SSH, in the same shape as a bare-metal group. There is no parallelscheduler_serversbucket. - Every cluster carries the same apply constraints: whether a lease ends on release or is killed at a time limit, how long it may run, how many cards one apply can take, how much CPU and memory each card comes with, how long an apply may block, and whether it queues. These used to live only inside plugin source — a three-hour training run placed on a cluster with a one-hour limit died at minute 60.
- The largest block one apply can take and what is left in the pool are two different numbers. A partition with 27 free cards spread across nodes may not open a single 8-card job when the job must land on one node. When the one-apply ceiling is unknown, the number is withheld rather than filled in from the remaining total.
- A routine claim selects a group first, then the scheduler best-fits one host inside it. A plugin-adapted group can be claimed by group as well.
- Routine MCP is exactly five tools:
gpu_status,gpu_apply,gpu_release,gpu_add_server,gpu_update_server. There is no browser UI. - The desktop app no longer shows a plugin cluster as a CPU node, its detail sheet stopped repeating the table row, and the history panel draws one shared GPU legend.
Plugin authors: the plugin contract is schema_version 3 and info must declare limits. A plugin still on v2 is not discovered. The bundled slurm-immediate is on v3.
ServerPilot is not published to PyPI. Install the tag you intend to run:
uv tool install --force "git+https://github.com/JinPLu/ServerPilot.git@v2.0.0"Full notes: CHANGELOG.en.md · 中文
ServerPilot 1.9.1
1.9.1 - 2026-08-27
ServerPilot 1.9.1 makes the new MCP entry panel work on an ordinary installation.
- The Settings page reported no MCP entry at all after
uv tool install, which is the ordinary macOS layout. The lookup followed the interpreter symlink before looking beside it, which walks out of the directory that holdsserverpilot-mcpand into the base interpreter's own. The daemon runs from launchd with no PATH of its own, so nothing else could find it either. - Backing up the database works on Windows. Each SQLite connection is now closed before the finished copy is published, instead of being left open by a context manager that only ends the transaction — Windows refuses to replace a file that still has a handle. The copy is flushed through a writable handle: Windows cannot fsync a file opened only for reading, so after the connections were closed the backup still failed there while working everywhere else.
- A Slurm submission that really went through is no longer reported as a failure.
sbatchdoes not drain its stdin on the paths where it exits early, which killed the process feeding it the script and turned that signal into the pipeline's exit status. Its own verdict was lost: a quota refusal or a bad partition arrived as a meaningless 141, and a job that had actually been accepted was left running with nothing here pointing at it. - The
mcpServersblock you copy out of the App shows the executable path as it is, instead of escaping every slash into\/. - A GPU row that marks itself unavailable while its own status text still claims the card is free is rejected again. Keying that check on the state code alone let the contradiction through whenever the state was something other than allocatable.
1.9.0 - 2026-08-27
ServerPilot 1.9.0 lets an agent reach it from a Windows machine and read everything it says, and stops the allocator breaking up your fleet one card at a time.
- Connecting an agent is one command:
serverpilot mcp install --client codex|claude|cursorregisters through each client's own mechanism, and merges into~/.cursor/mcp.jsonwithout disturbing servers already there.serverpilot mcp configprints the registration instead of writing it, and the README and agent guide now carry the standardmcpServersblock. - The desktop app Settings page now shows this installation's MCP entry as an absolute path and a pasteable
mcpServersblock, so you can copy them into Codex, Claude, or Cursor without reconstructing the command. If the executable is missing, the same panel says so and tells you how to install it. - The Windows archive ships
serverpilot-mcp.exenext to the app. The download previously contained only the GUI, so there was no MCP entry point on the machine at all and the documented registration commands only ever worked from a source install. - The MCP instructions, the three tool descriptions, and every value ServerPilot sends back are English. An allocatable card reports
available; a busy one reports why in a stable code such asrunning,held_idle, orbusy_unmanaged. The desktop app stays in Chinese — the two are separate surfaces now, so changing one no longer disturbs the other. gpu_applytakes the server with the fewest free GPUs that can still serve the whole request, and one lease always lands on a single machine. This is the difference between a fleet that keeps working and one that fills with holes: cards used to be taken in alphabetical server order, so a single-card request landed on the first eight-GPU machine, and a few of those left nothing able to start an eight-GPU run. An eight-GPU request could also come back as 5+3 across two machines — useless for a single-node job, while holding all eight away from someone who could use them.no_capacitycomes back as data rather than an error string, so an agent can tell a full fleet from a broken connection instead of retrying an answer. Releasing a lease that is already released confirms it instead of failing.- Shared scheduler clusters register through a local plugin instead of as bare metal. The plugin registers only the cards in your own jobs;
gpu_statusreports cluster headroom asscheduler_serversbefore you request anything, real GPU identities appear after, and idle reclaim cancels the job. A Slurm reference plugin,slurm-immediate, ships with the package. - A cluster that refused you no longer looks like a cluster that is merely full. A quota refusal, an unreachable scheduler, or a broken plugin reaches you as a failure with its reason, instead of
no_capacityfor an agent to wait out. - Two ways a cluster job could be left running are fixed: ServerPilot no longer asks a plugin to allocate again after the cards are already assigned, and on a lease spanning several machines every job is now recorded, so releasing leaves nothing behind. Requesting or releasing cluster resources also no longer blocks everything else — other local work previously waited up to a minute.
- A cluster job ServerPilot could not cancel while reclaiming an idle lease is now recorded in the audit trail. It used to vanish silently, leaving the job running against your quota with nothing here pointing at it. Releasing a lease yourself still refuses outright rather than reporting a success it did not achieve.
- When the control plane is unreachable,
serverpilot keepalive inspectandserverpilot keepalive stopreport and stop the workers still holding cards, andserverpilot daemon reclaimtakes the port back. None of them needs the daemon running. The port-ownership error now names the process holding it and its command line instead of only saying the port is foreign. - The MCP handshake reports ServerPilot's own version rather than the MCP SDK's, and each tool declares its effect, so a client can tell a read from a lease mutation instead of gating all three the same way.
- The standalone Windows app can create its database again. The packaged build was dropping the migration scripts, which left a source deployment as the only working option.
- ServerPilot installs from PyPI. A release refuses to publish when the tag and the package version disagree, or when the wheel is missing its migration scripts or arrives with a bundled plugin that is not executable — a plugin without that bit is invisible to discovery rather than failing loudly.
serverpilot --versionexists.- The security documentation matches the implementation: the control plane has no authentication, ServerPilot manages its own occupancy processes and plugin-side allocations but never your workloads, and fail-closed admission trusts the SSH user and the remote collector. Keepalive workers keep holding GPUs after the control plane stops, until it returns and reconciles or someone stops them on the server.
- A slow or hanging local daemon call no longer stalls every other MCP tool on the same connection, and each tool now advertises a real input schema instead of free-form arguments.
ServerPilot 1.8.0
ServerPilot 1.8.0 publishes telemetry only where the occupancy belongs to the caller: a free card reports capacity, and your own lease reports per-GPU utilisation and the card lagging behind.
gpu_statusnow answers three questions in three groups: an allocatable card reports capacity only (model, VRAM, available), a busy card reports who holds it, and telemetry appears only on cards the caller holds. Free cards used to carry telemetry too — and every bit of load observable on a free card comes from ServerPilot's own keepalive hold (80% of VRAM, released only when the card is actually allocated), so a machine with eight free cards read as eight-tenths full and an agent that checked availability against it concluded there was nothing to claim.- New
gpu_status(lease_id=…)returns your own lease: per-GPU rolling ten-minute averages and the latest sample, plus a lease summary — average utilisation, the smallest free VRAM across the lease, and on multi-GPU leases the utilisation spread and the card lagging behind. These are the numbers behind "is my job using these cards well, can I raise the batch size, is one card holding the rest back"; a card you held previously fell into the compactbusy_gpuslist with nothing but its task name. gpu_statusno longer takesinclude_busy: busy cards always come back inbusy_gpuswith their task, and your own cards come fromlease_id. Allocatable cards report one status, "available", instead of exposing keepalive's internal variants. Agpu_statusresponse for an eight-GPU machine went from 5,957 to 1,749 bytes.
ServerPilot 1.7.0
ServerPilot 1.7.0 turns the servers page back into a table you can compare down a column, with a pressure bar on all four metrics, and makes idle reclaim per GPU.
- The servers page is a table again, one 44pt row per machine: GPU utilisation, VRAM, CPU load and memory are all drawn the same way — a percentage and a bar, equal width, equal weight — so the eye can compare them straight down a column. CPU and memory previously carried a number with no bar, while the sort control still offered to sort by CPU load, which left the resulting order with nothing visible behind it.
- Narrow windows fold columns from the right instead of switching to a different layout: the full SSH command is never truncated at any width, and none of the four bars ever folds.
GPU modeldrops at 1280 andproject / taskbelow that; both stay in the row's tooltip and in the detail sheet. - Column headers are the sort controls: click one to sort by it, click again to reverse, and the active column darkens and carries an arrow. The headers now have accessibility names too, so screen readers no longer meet a row of unnamed buttons.
- Rows no longer print the absence of a task, and a host with no GPUs shows its core count and total memory where a GPU model would go — that is what that machine actually is.
- Idle reclaim is now per GPU: a claim that takes eight cards and uses one returns the other seven individually as each idle window elapses, while the working card keeps its claim. Previously a single running process protected every other GPU in the same claim.
- CPU cores, total memory, peak temperature, absolute VRAM and the full remote workspace path move into a new "host" card in the detail sheet. The workspace path could only ever render as
…Data/tmp/ljpin a row, which carries no information; all of it also stays in the row's tooltip. - The usage and settings pages now speak the same card language as the servers page: a group of facts sits in one white card, rows are separated by hairlines instead of each carrying its own fill and border. Usage detail gains a "resource total" card, and settings gains a "data state" card (connection, snapshot freshness and revision, server / GPU / lease counts, and whether resource changes can run).
- The server detail sheet drops its translucent material for the same plane as every other page, and the per-GPU grid's minimum column width now fits its whole contents, so mid-word truncations like
4 / 8…,32 / …andtask: …are gone. - Turning on the system Increase Contrast setting now actually changes the interface: cards gain an outline, hairlines deepen, status colours re-solve to 7:1, and bar tracks darken — applied immediately, with no restart.
- Settings cards align to the left margin instead of centring in a wide window, and the connection fact no longer claims a live local service while the read-only test fixture is in use.
- The usage page and the server detail sheet used to collapse into a single element, leaving screen reader users with one summary sentence and no access to any button or value inside; both are now readable item by item.
- Status colour is rebuilt as two tiers: a deep mark tier (
#00832F/#B05A00/#E40021) for dots, bars and status words, and a luminous area tier (#E7F8EB/#FFF1E5/#FFE7E8) that carries the brightness. Darkening the whole palette so a small dot could clear contrast on its own had left the green muddy and the amber mustard. - The interface drops from three background planes to two (white content over
#E9ECF1); the old three differed by only 1.06-1.09 each, which reads as a rendering fault rather than as depth. - The type ramp gains a 26pt display step, lifting the largest-to-smallest ratio from 1.70 to 2.60 so something can finally lead, and every numeral is now tabular so columns stop twitching on refresh.
- Every font size across the interface now comes from a six-step Apple semantic ramp (previously 19 sizes including half-pixel steps), so the interface follows the system text size.
- Settings drops a third "Settings" heading that repeated the sidebar and page title, and the filter control's duplicate label no longer stacks vertically in wide windows.
ServerPilot 1.6.0
ServerPilot 1.6.0 cuts about seventy percent of the context an agent spends reading GPU status, and returns idle-but-claimed GPUs to the pool on their own.
- Agents spend far less context reading GPU status: connection details and the remote working directory are now returned once per server instead of repeated on every GPU, and
gpu_statusalso lists busy cards with the task holding each one — so deciding where to place work takes a single call instead of two. On one 8-GPU server that decision path dropped from roughly 21,800 to 6,700 characters. gpu_statusaccepts aserver_idargument to narrow the response to one server.gpu_releasenow echoes the released lease id and its settled state, so an agent holding several leases can confirm them one by one instead of assuming one release finished everything.- The desktop App refreshes more cheaply: it no longer fetches the generic-resource and external-scheduler projections it never displays, cutting the measured state payload from roughly 77,000 to 59,700 characters (-22.4%) per refresh. Everything the interface actually renders — servers, GPUs, leases, resource usage — is byte-for-byte unchanged.
- Three desktop details now match the design contract: the server column header uses the
server.rackicon, the setting reads "data collection interval" (distinct from "refresh", which only re-reads local state), and the usage page's empty state reads "no current resource allocation". - Status colors now meet the accessibility floor: the normal green and caution amber darken to
#339653and#AA7C00(lightness only, hues unchanged), lifting status dots and pressure bars from as low as 1.53 to at least 3.01 against both the content surface and the page background — the WCAG threshold for non-text graphics. The error red already passed and is unchanged. - Idle GPUs come back on their own: when a lease's GPUs show no compute process across observations the collector can actually see, ServerPilot raises a warning first and then releases the lease back into the allocatable pool. An agent that forgets to release, or a job that finished without cleanup, no longer locks the cards indefinitely. The idle clock resets whenever telemetry goes stale, so a collector outage never reclaims a job that is merely unobserved.
- GPU status no longer collapses into "task in use": it now distinguishes a running task, a claim with no observed task, an unmanaged process, and an attribution conflict — so a card that is claimed but idle is visible at a glance.
- Upgrade note: an external agent's global rules are a static copy on disk. After upgrading, re-run
python3 scripts/install_agent_policy.py all --install(Cursor:--printthen paste), otherwise the rules describe the old response shape.