Skip to content

Releases: neuralmagic/gpu-pruner

v0.5.0

Choose a tag to compare

@wseaton wseaton released this 31 Aug 02:38
c046066

First release under the neuralmagic org.

New

  • LeaderWorkerSet support (leaderworkerset.x-k8s.io): new l resource flag, on by default. An idle StatefulSet owned by an LWS scales the whole LWS, since scaling only the member StatefulSet is undone by the LWS controller.
  • --idle-threshold (default 0.01): replaces the strict == 0 idle comparison to account for the DCGM GR_ENGINE_ACTIVE noise floor.
  • --exclude-namespaces / --exclude-pods: regex negative matchers injected into every metric selector of the pruning query.
  • Prometheus metrics: opt-in --metrics-addr endpoint with query/scale counters and gauges, plus Service/ServiceMonitor manifests.
  • Slack notifications with ack grace period (experimental): notify-then-wait with Keep 4h/8h/24h buttons. The interaction endpoint requires SLACK_SIGNING_SECRET and verifies Slack request signatures.
  • Web dashboard (experimental): opt-in --dashboard-addr serves stats and an idle-GPU-hours leaderboard; the Prometheus relay is restricted to instant queries. Deploy behind authenticated ingress.
  • Grafana dashboard: GPU health, DCGM, and idle/active leaderboard panels.
  • ClusterRole trimmed to the verbs the pruner actually uses.

Credits

Most of the feature work in this release was contributed by @fuddin-bit, integrated with fixes and hardening.

Full changelog: v0.3.0...v0.5.0

v0.3.0

Choose a tag to compare

@wseaton wseaton released this 06 Nov 15:59
aa1ab70
  • configurable resource level pruning, eg. Notebook only
  • add support for OTEL metrics for observability
  • better logging

Docker Images on Quay.io

0.2.0: Merge pull request #1 from wseaton/event-support

Choose a tag to compare

@wseaton wseaton released this 17 Jul 00:10
729ac7c
adding support for events