Repository navigation
Releases: neuralmagic/gpu-pruner
Releases · neuralmagic/gpu-pruner
Release list
v0.5.0
First release under the neuralmagic org.
New
- LeaderWorkerSet support (
leaderworkerset.x-k8s.io): newlresource flag, on by default. An idle StatefulSet owned by an LWS scales the whole LWS, since scaling only the member StatefulSet is undone by the LWS controller. --idle-threshold(default 0.01): replaces the strict== 0idle comparison to account for the DCGM GR_ENGINE_ACTIVE noise floor.--exclude-namespaces/--exclude-pods: regex negative matchers injected into every metric selector of the pruning query.- Prometheus metrics: opt-in
--metrics-addrendpoint with query/scale counters and gauges, plus Service/ServiceMonitor manifests. - Slack notifications with ack grace period (experimental): notify-then-wait with Keep 4h/8h/24h buttons. The interaction endpoint requires
SLACK_SIGNING_SECRETand verifies Slack request signatures. - Web dashboard (experimental): opt-in
--dashboard-addrserves stats and an idle-GPU-hours leaderboard; the Prometheus relay is restricted to instant queries. Deploy behind authenticated ingress. - Grafana dashboard: GPU health, DCGM, and idle/active leaderboard panels.
- ClusterRole trimmed to the verbs the pruner actually uses.
Credits
Most of the feature work in this release was contributed by @fuddin-bit, integrated with fixes and hardening.
Full changelog: v0.3.0...v0.5.0
v0.3.0
- configurable resource level pruning, eg.
Notebookonly - add support for OTEL metrics for observability
- better logging
0.2.0: Merge pull request #1 from wseaton/event-support
adding support for events