Skip to content

Releases: HiWay-Media/gpuledger

gpuledger v0.5.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 08:55
18aba73

A page to look at: serve shows its node at /, fleet serve shows the cluster from
one process, both server-rendered, refreshing without script, behind GET-only
endpoints and a strict Content-Security-Policy; the fleet view as a Nomad job, run for
real in the matrix — which found that a Consul service keeps a job off a node without
Consul from Nomad 1.3. The driver side, as before, only against a fake nvidia-smi
(GL-10, v0.6.0).

Upgrading from 0.4.0

  • Every HTTP endpoint answers GET and HEAD only; anything else is now a 405.
  • The job specs take -var consul=false for a cluster without Consul (default on:
    nothing changes where Consul runs).
  • fleet serve and deploy/nomad/gpuledger-fleet.nomad.hcl are new; they listen on 9878.

Added

  • gpuledger fleet serve: polls the fleet every --interval and serves its page (each
    node linked to its own), /fleet JSON, /metrics with the cluster's totals
    (gpuledger_fleet_*) and /healthz, behind the same GET-only, CSP headers; listens
    on 9878 by default (GL-34).
  • deploy/nomad/gpuledger-fleet.nomad.hcl: fleet serve as one service instance,
    discovery as a variable, checksum required. The Nomad matrix runs it with nomad job run on every version, the binary of the commit served locally (GL-35).
  • The node's page: serve answers / with the GPUs, their state and since when, who
    reserved and who holds each, and the findings — server-rendered HTML, light and dark,
    refreshing every --interval without a script (GL-33).
  • Safe to expose: every endpoint GET and HEAD only (405 otherwise), a
    Content-Security-Policy that allows the page's own stylesheet and nothing else,
    nosniff, no-referrer, X-Frame-Options: DENY, no-store; names escaped by
    context, tested with a container named like a script (GL-36).

Fixed

  • On a cluster without Consul, from Nomad 1.3, the system job was never placed: a Consul
    service block adds the constraint ${attr.consul.version} >= 1.8.0. All three job
    specs take -var consul=false, which leaves the registration out. Found by running
    the fleet job in the Nomad matrix.
  • The site's navbar wrapped onto a second line and out of the header once the README
    grew to seven sections. It is one line now whatever the README grows to: the sections
    scroll sideways when they do not fit, with a fade only on the side that hides some,
    GitHub pinned outside the scroll, and only the logo and GitHub below 900 px.

gpuledger v0.4.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 07:52
3135bbe

What the GPUs cost: counters of the seconds each GPU spends in each state and of each
job's GPU-seconds, held or idle; gpuledger report reading them back from Prometheus;
the waste in the dashboard and an alert on it. Every number checked through a real
Prometheus on real Nomad allocations in the matrix; the driver side, as before, only
against a fake nvidia-smi (GL-10, v0.5.0).

Upgrading from 0.3.0

  • Nothing to change: the counters are new series, and report is a new command.
  • serve --history (which the job specs set) keeps the counters across restarts; without
    it they start from zero at each start, which Prometheus reads as a counter reset.
  • The dashboard has three new panels and the rules a new alert; re-import them to get
    them.

Added

  • The cost in the dashboard and the rules: GPU-hours per job over the range, the idle
    share per job, GPUs per state from the counters; GPULedgerJobMostlyIdle (info) when
    a job with at least a GPU-hour reserved over a day left more than half of it idle.
    Unit-tested with promtool test rules, a mutated threshold checked to fail; every
    panel's query run against the real Prometheus in the matrix (GL-32).
  • gpuledger report --prometheus URL --since 7d: GPU-hours per job — reserved, held,
    idle and the idle share, most idle first — and per node and state, from the counters
    in Prometheus; --json under the JSON contract; a bearer token by variable name
    (GL-31).
  • Counters: gpuledger_gpu_state_seconds_total{state} per GPU (GL-29) and
    gpuledger_job_gpu_seconds_total{namespace,nomad_job,use} per job, use held or
    idle (GL-30) — GPU-hours and waste over any window with increase(). Kept in the
    --history file across restarts; gaps and partial reads accrue nothing. The Nomad
    matrix checks them through a real Prometheus on real reservations.

gpuledger v0.3.0

Choose a tag to compare

@github-actions github-actions released this 26 Sep 20:00
95d29e8

A hardened cluster: Nomad and Consul over mutual TLS, a workload identity in place of a
static token, fleet through Nomad's own service discovery, and release binaries with
build provenance. Every item observed in the Nomad matrix on each version that has the
feature; the driver side, as before, only against a fake nvidia-smi (GL-10, v0.4.0).

Upgrading from 0.2.0

  • The system job now requires -var checksum=sha256:… of the binary, from the
    release's checksums file (README: Run it); Nomad refuses a download that does not match.
  • Both job specs pass --nomad-node-id ${node.unique.id}; nothing to do unless you run
    gpuledger by other means with a workload identity before Nomad 1.11.
  • This is the first release with build provenance: gh attestation verify (README: Install).

Added

  • --nomad-node-id: the node's Nomad id given, /v1/agent/self is not called. Both
    job specs pass ${node.unique.id} (GL-25).
  • deploy/nomad/gpuledger.wi.nomad.hcl: the system job with no static token, Nomad
    1.5+ — identity { env = true } and the policy file bound to the job. The Nomad
    matrix runs gpuledger with a task's workload identity on every version from 1.5, and
    observed that /v1/agent/self refuses one before 1.11 (GL-25).
  • Verifiable releases: build provenance attestations on every binary and the checksums
    file, verified by the release workflow before it publishes; gh attestation verify in
    the README's install steps; a dry run of those steps, with a tampered binary that must
    fail, on every change to the release workflow (GL-28).

Changed

  • The system job requires -var checksum=sha256:… and passes it to the artifact, so
    Nomad refuses a binary that does not match the release's checksums file. The Nomad
    matrix checks the job validates with it and is refused without it (GL-28).
  • Consul over TLS for fleet: --consul-ca-cert, --consul-ca-path,
    --consul-client-cert, --consul-client-key, --consul-tls-server-name, defaulting
    to the Consul CLI's variables, and CONSUL_HTTP_SSL for a bare address. The matrix
    runs a Consul dev agent with verify_incoming (GL-27).
  • fleet --nomad-service NAME (and --nomad-namespace): the endpoints from Nomad's
    own service discovery, Nomad 1.3+, with the same address, token and TLS as the rest —
    for clusters without Consul. Observed on every matrix version from 1.3 (GL-26).
  • Nomad over mutual TLS: --nomad-ca-cert, --nomad-ca-path, --nomad-client-cert,
    --nomad-client-key, --nomad-tls-server-name, defaulting to the Nomad CLI's
    NOMAD_CACERT, NOMAD_CAPATH, NOMAD_CLIENT_CERT, NOMAD_CLIENT_KEY,
    NOMAD_TLS_SERVER_NAME. The Nomad matrix runs an agent with verify_https_client on
    every version (GL-24).

gpuledger v0.2.0

Choose a tag to compare

@github-actions github-actions released this 25 Sep 15:31
dfccb41

The whole cluster, and watched: fleet across nodes, since-when per GPU, per-card
thresholds from NVIDIA's own numbers, Podman, findings as metrics with alert rules and a
dashboard, and a JSON contract. Every Nomad-dependent fact re-observed on every stable
Nomad from 1.0.18 to 2.0.7; the driver side is still observed only against a fake
nvidia-smi — GL-10, now in v0.3.0.

Upgrading from 0.1.0

  • The tenant metrics' job label is now nomad_job (Prometheus had been renaming it
    exported_job): queries on gpuledger_tenant_*{job=…} must change.
  • --encoder-max defaults to 0, the card's own cap — none on Quadro and datacenter
    cards. To keep 0.1.0's behaviour, pass --encoder-max 8.
  • hot follows the driver's thermal margin and slowdown flags where the driver reports
    them; --temp-max is the fallback.
  • The system job adds --history on a sticky ephemeral disk, and takes datacenters as
    a variable.

Added

  • The JSON contract: "schema": 1 on ls --json, check --json, /ledger,
    /findings, fleet --json and the history file; the rule for what bumps it in the
    README; golden files of ls, check and fleet ls in testdata/golden that fail
    the tests on any change until regenerated on purpose. fleet shows each node's
    schema, and marks a node on another one in the table (GL-20).
  • gpuledger_findings{node,code,level} — the count of each code at the last refresh,
    with a zero for every code — and gpuledger_worst_level (0 OK … 3 ERROR) (GL-17).
  • deploy/prometheus/gpuledger.rules.yml: source down, unreserved and unmanaged
    tenants, hot, thermal slowdown, reserved-idle over six hours; unit-tested with
    promtool test rules in CI (GL-18).
  • deploy/grafana/gpuledger.json: states, findings, reserved-idle durations,
    utilisation, memory, encoder sessions, temperature and the driver's margin, memory by
    tenant, unreserved Nomad tenants. CI imports it into the latest Grafana; the Nomad
    matrix runs every panel's query and every rule against a real Prometheus (GL-19).
  • A test that every metric and label the rules and the dashboard name exists in
    /metrics.
  • Podman, next to Docker: --podman (default: /run/podman/podman.sock when it exists,
    off, or an endpoint), libpod-<id> cgroups (conmon excluded), and — for Nomad's
    podman driver, which labels nothing without extra_labels — the allocation id from the
    container's name <task>-<alloc id>, trusted only when Nomad returns it; allocFromName
    in the ledger JSON. The Nomad matrix runs a real nomad-driver-podman task on every
    version (GL-16).
  • internal/cards: NVENC engines, generation and session cap per card, with the source.
  • Optional nvidia-smi queries, each on its own so a missing field never breaks the main
    one: temperature.gpu.tlimit and the thermal slowdown flags (clocks_event_reasons.*,
    then clocks_throttle_reasons.*). In /ledger as thermalMarginC and
    thermalSlowdown, in /metrics as gpuledger_gpu_thermal_margin_celsius and
    gpuledger_gpu_thermal_slowdown.
  • --history FILE: serve records each GPU's state and since when, atomically, each
    refresh; ls and check read it. reserved-idle and idle say for how long, ls
    has a state column, /metrics has gpuledger_gpu_state{state} and
    gpuledger_gpu_state_since_timestamp_seconds. A silence over three intervals, or a
    partial read, restarts the clock. The system job keeps the file on a sticky disk (GL-14).
  • The ledger JSON has state and stateSince per GPU; fleet shares the one
    classification (ledger.Classify).
  • gpuledger fleet ls and fleet check: every node's /ledger, from --targets or
    Consul's health API (--consul, default $CONSUL_HTTP_ADDR; the token from the variable
    named by --consul-token-env). Each GPU is counted as held, reserved-idle, unaccounted
    or free, per node and per job; fleet check evaluates every node with one policy. A
    node or a Consul that cannot be read is a source-unavailable ERROR (GL-13).
  • The Nomad matrix runs fleet against the real node, directly and through a Consul dev
    agent with the system job's /healthz check.

Changed

  • encoder-saturated uses each card's published cap: none on Quadro RTX 4000, L4, T4
    and A10 ("Unrestricted" in NVIDIA's support matrix), 12 on GeForce. --encoder-max
    now defaults to 0 (the card's cap); N applies to every card, -1 turns it off. The
    old default, 8, flagged a limit the farm's cards do not have (GL-15).
  • hot trusts the driver first: an active thermal slowdown, or 5 °C or less to the
    card's own slowdown temperature. --temp-max decides only when the driver reports
    neither.

Fixed

  • The tenant metrics' job label collided with the job Prometheus attaches to every
    target, and was renamed exported_job on ingestion: the Nomad job never reached
    Prometheus under its name. It is now nomad_job. Found by running the dashboard
    against a real Prometheus.

gpuledger v0.1.0

Choose a tag to compare

@github-actions github-actions released this 24 Sep 08:38
9470cd4

The first release. The Nomad side is tested against real agents on every stable minor
from 1.0 to 2.0; the NVIDIA side against a fake nvidia-smi until the run on a real GPU
node (GL-10, now in v0.2.0).

Added

  • The Nomad matrix: integration/nomad_test.go against a real nomad agent -dev on the
    latest patch of every Nomad minor since 1.0, read from releases.hashicorp.com at run
    time, plus 1.7.3 — ACLs on, Nomad's example device plugin rebuilt as nvidia/gpu,
    Docker tasks in two namespaces and one outside Nomad; every join case, promtool check metrics on /metrics, nomad job validate on the system job. On every change and
    weekly (GL-21). Green on 1.0.18, 1.1.18, 1.2.16, 1.3.16, 1.4.14, 1.5.17, 1.6.10, 1.7.3,
    1.7.7, 1.8.4, 1.9.7, 1.10.5, 1.11.3 and 2.0.7.
  • deploy/nomad/gpuledger.policy.hcl: the ACL policy gpuledger's token needs —
    agent:read, node:read, read-job on the namespaces — the one the matrix tests with.
  • The README's Compatibility section: what the matrix observed on every version.
  • Unit tests for every fix below, the Docker client over a unix socket, nvidia-smi's
    failure paths, the exit code under every policy and the HTTP handlers.

Changed

  • Tenant series in /metrics carry a container_id label.
  • The system job takes datacenters as a variable: "*" matches every datacenter from
    Nomad 1.5 only.
  • The ledger JSON carries nomadRead and, per tenant, allocVisible.

Fixed

  • With ACLs, Nomad leaves out of /v1/node/<id>/allocations, without an error, the
    allocations in namespaces the token cannot read-job: every Nomad task there was a
    false unreserved-tenant BAD and its GPU a false idle. A Nomad container whose
    allocation was not returned is now a source-unavailable ERROR naming the namespace
    and the capability, and with Nomad unread (--no-nomad, or down) no task is judged on
    reservations (GL-22).
  • A process name holding arguments leaked a path element from them (ffmpeg -i /data/x/match.mp4 became match.mp4); the name is now cut at the first blank before
    the directory is dropped.
  • --exit-on with an unknown level (Bad, fatal) meant "always exit 0" in a gate; it is
    now a usage error, as are --interval ≤ 0 (the ticker panicked) and arguments after
    the flags (gpuledger --json check ran ls).
  • /metrics wrote each family's samples interleaved with the others'; families are now
    contiguous, HELP and TYPE first. Two tenants the Docker API could not name produced a
    duplicate series; tenant series carry container_id.
  • Tenants holding a GPU without a process were ordered by map iteration; ties are now
    broken by container name and id.
  • docker inspect of an id shorter than 12 characters panicked while formatting its error.
  • Reservations keep NVIDIA GPUs only (vendor nvidia, type gpu): another plugin's
    device with type gpu could not match any UUID. A Nomad 403 names the ACL capability
    the token lacks.
  • /healthz answered ok in the body of a 503; it now lists the failing sources.