Skip to content

gpuledger v0.4.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 07:52
· 6 commits to main since this release
3135bbe

What the GPUs cost: counters of the seconds each GPU spends in each state and of each
job's GPU-seconds, held or idle; gpuledger report reading them back from Prometheus;
the waste in the dashboard and an alert on it. Every number checked through a real
Prometheus on real Nomad allocations in the matrix; the driver side, as before, only
against a fake nvidia-smi (GL-10, v0.5.0).

Upgrading from 0.3.0

  • Nothing to change: the counters are new series, and report is a new command.
  • serve --history (which the job specs set) keeps the counters across restarts; without
    it they start from zero at each start, which Prometheus reads as a counter reset.
  • The dashboard has three new panels and the rules a new alert; re-import them to get
    them.

Added

  • The cost in the dashboard and the rules: GPU-hours per job over the range, the idle
    share per job, GPUs per state from the counters; GPULedgerJobMostlyIdle (info) when
    a job with at least a GPU-hour reserved over a day left more than half of it idle.
    Unit-tested with promtool test rules, a mutated threshold checked to fail; every
    panel's query run against the real Prometheus in the matrix (GL-32).
  • gpuledger report --prometheus URL --since 7d: GPU-hours per job — reserved, held,
    idle and the idle share, most idle first — and per node and state, from the counters
    in Prometheus; --json under the JSON contract; a bearer token by variable name
    (GL-31).
  • Counters: gpuledger_gpu_state_seconds_total{state} per GPU (GL-29) and
    gpuledger_job_gpu_seconds_total{namespace,nomad_job,use} per job, use held or
    idle (GL-30) — GPU-hours and waste over any window with increase(). Kept in the
    --history file across restarts; gaps and partial reads accrue nothing. The Nomad
    matrix checks them through a real Prometheus on real reservations.