gpuledger v0.2.0
The whole cluster, and watched: fleet across nodes, since-when per GPU, per-card
thresholds from NVIDIA's own numbers, Podman, findings as metrics with alert rules and a
dashboard, and a JSON contract. Every Nomad-dependent fact re-observed on every stable
Nomad from 1.0.18 to 2.0.7; the driver side is still observed only against a fake
nvidia-smi — GL-10, now in v0.3.0.
Upgrading from 0.1.0
- The tenant metrics'
joblabel is nownomad_job(Prometheus had been renaming it
exported_job): queries ongpuledger_tenant_*{job=…}must change. --encoder-maxdefaults to 0, the card's own cap — none on Quadro and datacenter
cards. To keep 0.1.0's behaviour, pass--encoder-max 8.hotfollows the driver's thermal margin and slowdown flags where the driver reports
them;--temp-maxis the fallback.- The system job adds
--historyon a sticky ephemeral disk, and takesdatacentersas
a variable.
Added
- The JSON contract:
"schema": 1onls --json,check --json,/ledger,
/findings,fleet --jsonand the history file; the rule for what bumps it in the
README; golden files ofls,checkandfleet lsintestdata/goldenthat fail
the tests on any change until regenerated on purpose.fleetshows each node's
schema, and marks a node on another one in the table (GL-20). gpuledger_findings{node,code,level}— the count of each code at the last refresh,
with a zero for every code — andgpuledger_worst_level(0 OK … 3 ERROR) (GL-17).deploy/prometheus/gpuledger.rules.yml: source down, unreserved and unmanaged
tenants, hot, thermal slowdown, reserved-idle over six hours; unit-tested with
promtool test rulesin CI (GL-18).deploy/grafana/gpuledger.json: states, findings, reserved-idle durations,
utilisation, memory, encoder sessions, temperature and the driver's margin, memory by
tenant, unreserved Nomad tenants. CI imports it into the latest Grafana; the Nomad
matrix runs every panel's query and every rule against a real Prometheus (GL-19).- A test that every metric and label the rules and the dashboard name exists in
/metrics. - Podman, next to Docker:
--podman(default:/run/podman/podman.sockwhen it exists,
off, or an endpoint),libpod-<id>cgroups (conmon excluded), and — for Nomad's
podman driver, which labels nothing withoutextra_labels— the allocation id from the
container's name<task>-<alloc id>, trusted only when Nomad returns it;allocFromName
in the ledger JSON. The Nomad matrix runs a realnomad-driver-podmantask on every
version (GL-16). internal/cards: NVENC engines, generation and session cap per card, with the source.- Optional nvidia-smi queries, each on its own so a missing field never breaks the main
one:temperature.gpu.tlimitand the thermal slowdown flags (clocks_event_reasons.*,
thenclocks_throttle_reasons.*). In/ledgerasthermalMarginCand
thermalSlowdown, in/metricsasgpuledger_gpu_thermal_margin_celsiusand
gpuledger_gpu_thermal_slowdown. --history FILE:serverecords each GPU's state and since when, atomically, each
refresh;lsandcheckread it.reserved-idleandidlesay for how long,ls
has astatecolumn,/metricshasgpuledger_gpu_state{state}and
gpuledger_gpu_state_since_timestamp_seconds. A silence over three intervals, or a
partial read, restarts the clock. The system job keeps the file on a sticky disk (GL-14).- The ledger JSON has
stateandstateSinceper GPU;fleetshares the one
classification (ledger.Classify). gpuledger fleet lsandfleet check: every node's/ledger, from--targetsor
Consul's health API (--consul, default$CONSUL_HTTP_ADDR; the token from the variable
named by--consul-token-env). Each GPU is counted as held, reserved-idle, unaccounted
or free, per node and per job;fleet checkevaluates every node with one policy. A
node or a Consul that cannot be read is asource-unavailableERROR (GL-13).- The Nomad matrix runs
fleetagainst the real node, directly and through a Consul dev
agent with the system job's/healthzcheck.
Changed
encoder-saturateduses each card's published cap: none on Quadro RTX 4000, L4, T4
and A10 ("Unrestricted" in NVIDIA's support matrix), 12 on GeForce.--encoder-max
now defaults to 0 (the card's cap);Napplies to every card,-1turns it off. The
old default, 8, flagged a limit the farm's cards do not have (GL-15).hottrusts the driver first: an active thermal slowdown, or 5 °C or less to the
card's own slowdown temperature.--temp-maxdecides only when the driver reports
neither.
Fixed
- The tenant metrics'
joblabel collided with thejobPrometheus attaches to every
target, and was renamedexported_jobon ingestion: the Nomad job never reached
Prometheus under its name. It is nownomad_job. Found by running the dashboard
against a real Prometheus.