gpuledger v0.4.0
What the GPUs cost: counters of the seconds each GPU spends in each state and of each
job's GPU-seconds, held or idle; gpuledger report reading them back from Prometheus;
the waste in the dashboard and an alert on it. Every number checked through a real
Prometheus on real Nomad allocations in the matrix; the driver side, as before, only
against a fake nvidia-smi (GL-10, v0.5.0).
Upgrading from 0.3.0
- Nothing to change: the counters are new series, and
reportis a new command. serve --history(which the job specs set) keeps the counters across restarts; without
it they start from zero at each start, which Prometheus reads as a counter reset.- The dashboard has three new panels and the rules a new alert; re-import them to get
them.
Added
- The cost in the dashboard and the rules: GPU-hours per job over the range, the idle
share per job, GPUs per state from the counters;GPULedgerJobMostlyIdle(info) when
a job with at least a GPU-hour reserved over a day left more than half of it idle.
Unit-tested withpromtool test rules, a mutated threshold checked to fail; every
panel's query run against the real Prometheus in the matrix (GL-32). gpuledger report --prometheus URL --since 7d: GPU-hours per job — reserved, held,
idle and the idle share, most idle first — and per node and state, from the counters
in Prometheus;--jsonunder the JSON contract; a bearer token by variable name
(GL-31).- Counters:
gpuledger_gpu_state_seconds_total{state}per GPU (GL-29) and
gpuledger_job_gpu_seconds_total{namespace,nomad_job,use}per job,useheldor
idle(GL-30) — GPU-hours and waste over any window withincrease(). Kept in the
--historyfile across restarts; gaps and partial reads accrue nothing. The Nomad
matrix checks them through a real Prometheus on real reservations.