Skip to content

v0.5.0

Latest

Choose a tag to compare

@wseaton wseaton released this 31 Aug 02:38
· 2 commits to main since this release
c046066

First release under the neuralmagic org.

New

  • LeaderWorkerSet support (leaderworkerset.x-k8s.io): new l resource flag, on by default. An idle StatefulSet owned by an LWS scales the whole LWS, since scaling only the member StatefulSet is undone by the LWS controller.
  • --idle-threshold (default 0.01): replaces the strict == 0 idle comparison to account for the DCGM GR_ENGINE_ACTIVE noise floor.
  • --exclude-namespaces / --exclude-pods: regex negative matchers injected into every metric selector of the pruning query.
  • Prometheus metrics: opt-in --metrics-addr endpoint with query/scale counters and gauges, plus Service/ServiceMonitor manifests.
  • Slack notifications with ack grace period (experimental): notify-then-wait with Keep 4h/8h/24h buttons. The interaction endpoint requires SLACK_SIGNING_SECRET and verifies Slack request signatures.
  • Web dashboard (experimental): opt-in --dashboard-addr serves stats and an idle-GPU-hours leaderboard; the Prometheus relay is restricted to instant queries. Deploy behind authenticated ingress.
  • Grafana dashboard: GPU health, DCGM, and idle/active leaderboard panels.
  • ClusterRole trimmed to the verbs the pruner actually uses.

Credits

Most of the feature work in this release was contributed by @fuddin-bit, integrated with fixes and hardening.

Full changelog: v0.3.0...v0.5.0