Releases: HiWay-Media/gpuledger
Release list
gpuledger v0.5.0
A page to look at: serve shows its node at /, fleet serve shows the cluster from
one process, both server-rendered, refreshing without script, behind GET-only
endpoints and a strict Content-Security-Policy; the fleet view as a Nomad job, run for
real in the matrix — which found that a Consul service keeps a job off a node without
Consul from Nomad 1.3. The driver side, as before, only against a fake nvidia-smi
(GL-10, v0.6.0).
Upgrading from 0.4.0
- Every HTTP endpoint answers GET and HEAD only; anything else is now a 405.
- The job specs take
-var consul=falsefor a cluster without Consul (default on:
nothing changes where Consul runs). fleet serveanddeploy/nomad/gpuledger-fleet.nomad.hclare new; they listen on 9878.
Added
gpuledger fleet serve: polls the fleet every--intervaland serves its page (each
node linked to its own),/fleetJSON,/metricswith the cluster's totals
(gpuledger_fleet_*) and/healthz, behind the same GET-only, CSP headers; listens
on 9878 by default (GL-34).deploy/nomad/gpuledger-fleet.nomad.hcl: fleet serve as one service instance,
discovery as a variable, checksum required. The Nomad matrix runs it withnomad job runon every version, the binary of the commit served locally (GL-35).- The node's page:
serveanswers/with the GPUs, their state and since when, who
reserved and who holds each, and the findings — server-rendered HTML, light and dark,
refreshing every--intervalwithout a script (GL-33). - Safe to expose: every endpoint GET and HEAD only (405 otherwise), a
Content-Security-Policy that allows the page's own stylesheet and nothing else,
nosniff,no-referrer,X-Frame-Options: DENY,no-store; names escaped by
context, tested with a container named like a script (GL-36).
Fixed
- On a cluster without Consul, from Nomad 1.3, the system job was never placed: a Consul
serviceblock adds the constraint${attr.consul.version} >= 1.8.0. All three job
specs take-var consul=false, which leaves the registration out. Found by running
the fleet job in the Nomad matrix. - The site's navbar wrapped onto a second line and out of the header once the README
grew to seven sections. It is one line now whatever the README grows to: the sections
scroll sideways when they do not fit, with a fade only on the side that hides some,
GitHub pinned outside the scroll, and only the logo and GitHub below 900 px.
gpuledger v0.4.0
What the GPUs cost: counters of the seconds each GPU spends in each state and of each
job's GPU-seconds, held or idle; gpuledger report reading them back from Prometheus;
the waste in the dashboard and an alert on it. Every number checked through a real
Prometheus on real Nomad allocations in the matrix; the driver side, as before, only
against a fake nvidia-smi (GL-10, v0.5.0).
Upgrading from 0.3.0
- Nothing to change: the counters are new series, and
reportis a new command. serve --history(which the job specs set) keeps the counters across restarts; without
it they start from zero at each start, which Prometheus reads as a counter reset.- The dashboard has three new panels and the rules a new alert; re-import them to get
them.
Added
- The cost in the dashboard and the rules: GPU-hours per job over the range, the idle
share per job, GPUs per state from the counters;GPULedgerJobMostlyIdle(info) when
a job with at least a GPU-hour reserved over a day left more than half of it idle.
Unit-tested withpromtool test rules, a mutated threshold checked to fail; every
panel's query run against the real Prometheus in the matrix (GL-32). gpuledger report --prometheus URL --since 7d: GPU-hours per job — reserved, held,
idle and the idle share, most idle first — and per node and state, from the counters
in Prometheus;--jsonunder the JSON contract; a bearer token by variable name
(GL-31).- Counters:
gpuledger_gpu_state_seconds_total{state}per GPU (GL-29) and
gpuledger_job_gpu_seconds_total{namespace,nomad_job,use}per job,useheldor
idle(GL-30) — GPU-hours and waste over any window withincrease(). Kept in the
--historyfile across restarts; gaps and partial reads accrue nothing. The Nomad
matrix checks them through a real Prometheus on real reservations.
gpuledger v0.3.0
A hardened cluster: Nomad and Consul over mutual TLS, a workload identity in place of a
static token, fleet through Nomad's own service discovery, and release binaries with
build provenance. Every item observed in the Nomad matrix on each version that has the
feature; the driver side, as before, only against a fake nvidia-smi (GL-10, v0.4.0).
Upgrading from 0.2.0
- The system job now requires
-var checksum=sha256:…of the binary, from the
release's checksums file (README: Run it); Nomad refuses a download that does not match. - Both job specs pass
--nomad-node-id ${node.unique.id}; nothing to do unless you run
gpuledger by other means with a workload identity before Nomad 1.11. - This is the first release with build provenance:
gh attestation verify(README: Install).
Added
--nomad-node-id: the node's Nomad id given,/v1/agent/selfis not called. Both
job specs pass${node.unique.id}(GL-25).deploy/nomad/gpuledger.wi.nomad.hcl: the system job with no static token, Nomad
1.5+ —identity { env = true }and the policy file bound to the job. The Nomad
matrix runs gpuledger with a task's workload identity on every version from 1.5, and
observed that/v1/agent/selfrefuses one before 1.11 (GL-25).- Verifiable releases: build provenance attestations on every binary and the checksums
file, verified by the release workflow before it publishes;gh attestation verifyin
the README's install steps; a dry run of those steps, with a tampered binary that must
fail, on every change to the release workflow (GL-28).
Changed
- The system job requires
-var checksum=sha256:…and passes it to the artifact, so
Nomad refuses a binary that does not match the release's checksums file. The Nomad
matrix checks the job validates with it and is refused without it (GL-28). - Consul over TLS for
fleet:--consul-ca-cert,--consul-ca-path,
--consul-client-cert,--consul-client-key,--consul-tls-server-name, defaulting
to the Consul CLI's variables, andCONSUL_HTTP_SSLfor a bare address. The matrix
runs a Consul dev agent withverify_incoming(GL-27). fleet --nomad-service NAME(and--nomad-namespace): the endpoints from Nomad's
own service discovery, Nomad 1.3+, with the same address, token and TLS as the rest —
for clusters without Consul. Observed on every matrix version from 1.3 (GL-26).- Nomad over mutual TLS:
--nomad-ca-cert,--nomad-ca-path,--nomad-client-cert,
--nomad-client-key,--nomad-tls-server-name, defaulting to the Nomad CLI's
NOMAD_CACERT,NOMAD_CAPATH,NOMAD_CLIENT_CERT,NOMAD_CLIENT_KEY,
NOMAD_TLS_SERVER_NAME. The Nomad matrix runs an agent withverify_https_clienton
every version (GL-24).
gpuledger v0.2.0
The whole cluster, and watched: fleet across nodes, since-when per GPU, per-card
thresholds from NVIDIA's own numbers, Podman, findings as metrics with alert rules and a
dashboard, and a JSON contract. Every Nomad-dependent fact re-observed on every stable
Nomad from 1.0.18 to 2.0.7; the driver side is still observed only against a fake
nvidia-smi — GL-10, now in v0.3.0.
Upgrading from 0.1.0
- The tenant metrics'
joblabel is nownomad_job(Prometheus had been renaming it
exported_job): queries ongpuledger_tenant_*{job=…}must change. --encoder-maxdefaults to 0, the card's own cap — none on Quadro and datacenter
cards. To keep 0.1.0's behaviour, pass--encoder-max 8.hotfollows the driver's thermal margin and slowdown flags where the driver reports
them;--temp-maxis the fallback.- The system job adds
--historyon a sticky ephemeral disk, and takesdatacentersas
a variable.
Added
- The JSON contract:
"schema": 1onls --json,check --json,/ledger,
/findings,fleet --jsonand the history file; the rule for what bumps it in the
README; golden files ofls,checkandfleet lsintestdata/goldenthat fail
the tests on any change until regenerated on purpose.fleetshows each node's
schema, and marks a node on another one in the table (GL-20). gpuledger_findings{node,code,level}— the count of each code at the last refresh,
with a zero for every code — andgpuledger_worst_level(0 OK … 3 ERROR) (GL-17).deploy/prometheus/gpuledger.rules.yml: source down, unreserved and unmanaged
tenants, hot, thermal slowdown, reserved-idle over six hours; unit-tested with
promtool test rulesin CI (GL-18).deploy/grafana/gpuledger.json: states, findings, reserved-idle durations,
utilisation, memory, encoder sessions, temperature and the driver's margin, memory by
tenant, unreserved Nomad tenants. CI imports it into the latest Grafana; the Nomad
matrix runs every panel's query and every rule against a real Prometheus (GL-19).- A test that every metric and label the rules and the dashboard name exists in
/metrics. - Podman, next to Docker:
--podman(default:/run/podman/podman.sockwhen it exists,
off, or an endpoint),libpod-<id>cgroups (conmon excluded), and — for Nomad's
podman driver, which labels nothing withoutextra_labels— the allocation id from the
container's name<task>-<alloc id>, trusted only when Nomad returns it;allocFromName
in the ledger JSON. The Nomad matrix runs a realnomad-driver-podmantask on every
version (GL-16). internal/cards: NVENC engines, generation and session cap per card, with the source.- Optional nvidia-smi queries, each on its own so a missing field never breaks the main
one:temperature.gpu.tlimitand the thermal slowdown flags (clocks_event_reasons.*,
thenclocks_throttle_reasons.*). In/ledgerasthermalMarginCand
thermalSlowdown, in/metricsasgpuledger_gpu_thermal_margin_celsiusand
gpuledger_gpu_thermal_slowdown. --history FILE:serverecords each GPU's state and since when, atomically, each
refresh;lsandcheckread it.reserved-idleandidlesay for how long,ls
has astatecolumn,/metricshasgpuledger_gpu_state{state}and
gpuledger_gpu_state_since_timestamp_seconds. A silence over three intervals, or a
partial read, restarts the clock. The system job keeps the file on a sticky disk (GL-14).- The ledger JSON has
stateandstateSinceper GPU;fleetshares the one
classification (ledger.Classify). gpuledger fleet lsandfleet check: every node's/ledger, from--targetsor
Consul's health API (--consul, default$CONSUL_HTTP_ADDR; the token from the variable
named by--consul-token-env). Each GPU is counted as held, reserved-idle, unaccounted
or free, per node and per job;fleet checkevaluates every node with one policy. A
node or a Consul that cannot be read is asource-unavailableERROR (GL-13).- The Nomad matrix runs
fleetagainst the real node, directly and through a Consul dev
agent with the system job's/healthzcheck.
Changed
encoder-saturateduses each card's published cap: none on Quadro RTX 4000, L4, T4
and A10 ("Unrestricted" in NVIDIA's support matrix), 12 on GeForce.--encoder-max
now defaults to 0 (the card's cap);Napplies to every card,-1turns it off. The
old default, 8, flagged a limit the farm's cards do not have (GL-15).hottrusts the driver first: an active thermal slowdown, or 5 °C or less to the
card's own slowdown temperature.--temp-maxdecides only when the driver reports
neither.
Fixed
- The tenant metrics'
joblabel collided with thejobPrometheus attaches to every
target, and was renamedexported_jobon ingestion: the Nomad job never reached
Prometheus under its name. It is nownomad_job. Found by running the dashboard
against a real Prometheus.
gpuledger v0.1.0
The first release. The Nomad side is tested against real agents on every stable minor
from 1.0 to 2.0; the NVIDIA side against a fake nvidia-smi until the run on a real GPU
node (GL-10, now in v0.2.0).
Added
- The Nomad matrix:
integration/nomad_test.goagainst a realnomad agent -devon the
latest patch of every Nomad minor since 1.0, read from releases.hashicorp.com at run
time, plus 1.7.3 — ACLs on, Nomad's example device plugin rebuilt asnvidia/gpu,
Docker tasks in two namespaces and one outside Nomad; every join case,promtool check metricson/metrics,nomad job validateon the system job. On every change and
weekly (GL-21). Green on 1.0.18, 1.1.18, 1.2.16, 1.3.16, 1.4.14, 1.5.17, 1.6.10, 1.7.3,
1.7.7, 1.8.4, 1.9.7, 1.10.5, 1.11.3 and 2.0.7. deploy/nomad/gpuledger.policy.hcl: the ACL policy gpuledger's token needs —
agent:read,node:read,read-jobon the namespaces — the one the matrix tests with.- The README's Compatibility section: what the matrix observed on every version.
- Unit tests for every fix below, the Docker client over a unix socket, nvidia-smi's
failure paths, the exit code under every policy and the HTTP handlers.
Changed
- Tenant series in
/metricscarry acontainer_idlabel. - The system job takes
datacentersas a variable:"*"matches every datacenter from
Nomad 1.5 only. - The ledger JSON carries
nomadReadand, per tenant,allocVisible.
Fixed
- With ACLs, Nomad leaves out of
/v1/node/<id>/allocations, without an error, the
allocations in namespaces the token cannotread-job: every Nomad task there was a
falseunreserved-tenantBAD and its GPU a falseidle. A Nomad container whose
allocation was not returned is now asource-unavailableERROR naming the namespace
and the capability, and with Nomad unread (--no-nomad, or down) no task is judged on
reservations (GL-22). - A process name holding arguments leaked a path element from them (
ffmpeg -i /data/x/match.mp4becamematch.mp4); the name is now cut at the first blank before
the directory is dropped. --exit-onwith an unknown level (Bad,fatal) meant "always exit 0" in a gate; it is
now a usage error, as are--interval≤ 0 (the ticker panicked) and arguments after
the flags (gpuledger --json checkranls)./metricswrote each family's samples interleaved with the others'; families are now
contiguous, HELP and TYPE first. Two tenants the Docker API could not name produced a
duplicate series; tenant series carrycontainer_id.- Tenants holding a GPU without a process were ordered by map iteration; ties are now
broken by container name and id. docker inspectof an id shorter than 12 characters panicked while formatting its error.- Reservations keep NVIDIA GPUs only (vendor
nvidia, typegpu): another plugin's
device with typegpucould not match any UUID. A Nomad 403 names the ACL capability
the token lacks. /healthzansweredokin the body of a 503; it now lists the failing sources.