Skip to content

v1.0.0

Pre-release
Pre-release

Choose a tag to compare

@solusops solusops released this 06 Dec 13:19
· 208 commits to main since this release
e65abb5

What It Does

Real-time GPU monitoring (utilization, memory, temp, power) + intelligent benchmarking with auto-scaling workloads and performance baselines. Web dashboard + CLI interface.

Use Cases

  • Power Users: Wanting to test their new GPUs, workflows etc. AIO monitoring and benchmarking tool.

Known Issues

  • Per-process GPU %: Only CUDA workloads (graphics processes show 0%)
  • Benchmark tab: Requires PyTorch/CuPy + CUDA Toolkit
  • Setup: Windows PowerShell only (manual venv for Linux/Mac)
  • Scope: Single-machine monitoring (multi-node in v2.0)

Installation

Extract ZIP → .\setup.ps1 → python health_monitor.py → http://localhost:8090

Requires: Python 3.8+, NVIDIA GPU with drivers

What's Changed

  • Modulize the files by @DataBoySu in #1
  • First version by @DataBoySu in #2

New Contributors

  • @DataBoySu made their first contribution in #1

Full Changelog: https://github.com/DataBoySu/cluster-monitor/commits/v1.0.0