Skip to content

2026.02.09

Nichamon edited this page May 6, 2026 · 2 revisions

The meeting covered a presentation that provided a technical deep-dive into collecting per-job resource metrics (CPU, memory, IO, network) in multi-tenant environments, examining different data sources (cgroup v2, /proc/ files, and future eBPF). The community reached consensus that cgroup v2-based collection for CPU, memory, and IO has no deployment blockers, with job-level granularity sufficient for initial implementation and configurable cgroup directory patterns needed to support different job schedulers.

Here is the link to the presentation slides: https://github.com/ovis-hpc/ovis-wiki/blob/main/Per-Job%20Monitoring%20Data%20Collection%20with%20LDMS%20--%20published-1.pdf

Main

LDMSCON

Tutorials are available at the conference websites

D/SOS Documentation

LDMS v4 Documentation

Basic

Configurations

Features & Functionalities

Working Examples

Development

Reference Docs

Building

Cray Specific
RPMs
  • Coming soon!

Adding to the code base

Testing

Misc

Man Pages

  • Man pages currently not posted, but they are available in the source and build

LDMS Documentation (v3 branches)

V3 has been deprecated and will be removed soon

Basic

Reference Docs

Building

General
Cray Specific

Configuring

Running

  • Running

Tutorial

Clone this wiki locally