Repository navigation
v2.2.0
Cards can be claimed, released, and deleted, and the interface stops stalling.
- A server that stopped answering no longer blocks an apply. The cards it held itself still counted as reclaimable capacity, so an apply picked that host, spent its whole budget on an SSH timeout and failed, while other machines in the group had free cards all along.
- A server you can no longer reach can now be deleted, and the cards it holds released. Releasing required a successful collection and deleting required the leases released first — both asking a machine that no longer exists for proof, so the cards stayed locked in the ledger forever. The release is recorded as "server unreachable; operator settled the ledger". A host that still answers is refused exactly as before.
- The interface and the CLI no longer stall on a cycle. Every collection cycle restarted a batch of subprocesses to ask each plugin what it was, enough to stop the whole service for a second or two. The server list went from 1.7s to 0.05s.
- A multi-card apply is no longer killed by a 20-second timeout. The wait is now the server's real budget rather than a guess from the card count, and applies pinned to different hosts or groups no longer block each other.
- Registering a server now connects to it on the spot and tells you what happened. Registration used to be a database write that answered "added", so a host that can never connect was first reported as a new server and then became an unexplained red row. It now says plainly: connected and how many cards were found, or not connected and why.
- An agent can now register a server into a group, and pick the right observation profile.
gpu_add_server/gpu_update_serverdid not acceptserver_group_id, so a server registered through them could never be selected; the default profile also named the most special case, which connects to an ordinary machine and reports no cards. - The state page no longer claims occupancy is running when it cannot see the card. It used to replay the last recorded "running"; it now says that a card it cannot see cannot confirm what is on it.
- Clearing an "idle" hold now tells you when those cards last ran a compute process. A job between two batches looks exactly like one that has ended, and clearing the wrong lease wedges its cards. Only processes started after you took the cards count; the judgement stays yours.
- A lease's telemetry reflects only the load during your own hold. The ten-minute average used to start before the claim, so sizing a batch from it meant sizing it against someone else's work.
- Deleted servers and released tasks no longer leave warnings behind. 21 of the 23 live warnings pointed at resources that no longer existed, burying the ones that mattered.
- Re-registering a machine on a new port no longer lets two registrations hold the same card. The registration that can still see the card keeps it; the old one steps aside.
- The two server-management tools can finally be called the way they are documented. They demanded three parameters the instructions never mention, so following the documentation could only fail. The CLI also receives the structured detail behind a failure now, such as the list of groups to choose from.