Skip to content

Mess Profiler

Victor Xirau Guardans edited this page Sep 29, 2026 · 4 revisions

Mess Profiler is a convenience wrapper around performance counter tools like perf, likwid, vtune, or others that leverages the Mess benchmark's automatic counter discovery to profile application memory bandwidth.

Note: You don't need to use Mess Profiler as your actual profiler. It can serve as a counter discovery tool—run it with --dry-run to reveal the correct counters for your system, then use those counters with your preferred profiling tool. The convenience of Mess Profiler is that it outputs data in a format that Plotter-Parser understands for plotting application profiles on bandwidth-latency curves.


Why Mess Profiler?

Measuring memory bandwidth with perf, likwid, vtune, or others typically requires:

  1. Identifying the correct counters for your specific CPU (e.g., uncore_imc_0/cas_count_read/)
  2. Trial and error to find working counter names across different vendors and generations
  3. Parsing tool-specific output formats

Mess Profiler eliminates this overhead. The Mess benchmark already auto-discovers the correct memory bandwidth counters for your system. Mess Profiler reuses this discovery, giving you:

  • Zero configuration: Works out of the box on any supported system
  • Consistent output: Same CSV format regardless of underlying tool
  • Correct counters: Uses the same counters validated by the Mess benchmark

How It Works

  1. Mess Profiler runs the same counter discovery logic as the main Mess benchmark
  2. It identifies the appropriate profiling backend for your system
  3. It wraps the backend tool and parses its output into a consistent format
  4. It inherits CPU/memory bindings from numactl or taskset automatically

Use --dry-run to see exactly what counters were discovered:

./build/bin/mess-profiler --dry-run

This outputs the detected backend and counter names without running any measurements.


Usage

./build/bin/mess-profiler [options] [--] <command> [args...]

Options

Measurement Options

Option Description
-s, --interval <time> Sampling interval (e.g., 100ms, 1s). Default: summary mode
-o, --output <file> Output file. Default: stdout

Targeting Options

Option Description
-a, --system-wide System-wide profiling (all CPUs/sockets)
-p, --pid <pid> Profile existing process by PID
-C, --cpu <list> Profile only specified CPUs (e.g., 0-7,16-23)
-N, --nodes <list> Monitor memory traffic to specified NUMA nodes
--no-inherit Don't inherit binding from parent (numactl/taskset)

Output Options

Option Description
-v, --verbose Verbose output (show counter details)
--csv CSV output (default)
--human Human-readable output
--dry, --dry-run Show discovered counters and exit (no measurement)

Examples

Discover Counters

# See what counters mess-profiler will use on your system
./build/bin/mess-profiler --dry-run

Basic Profiling

# Profile an application with 100ms sampling
./build/bin/mess-profiler -s 100ms ./my_app

# Save output to file
./build/bin/mess-profiler -s 50ms -o bandwidth.csv ./my_app

System-Wide Profiling

# Profile all CPUs for 10 seconds
./build/bin/mess-profiler -a -s 1s sleep 10

With NUMA Binding

# Profile app bound to NUMA node 0, cores 0-7
numactl -m 0 -C 0-7 ./build/bin/mess-profiler -s 100ms ./my_app

# Explicit targeting (alternative)
./build/bin/mess-profiler -C 0-7 -N 0 -s 100ms ./my_app

Output Format

The profiler outputs CSV with the following columns:

Column Description
Timestamp(s) Time since start in seconds
Bandwidth(GB/s) Measured memory bandwidth
ReadBytes Bytes read from memory
WriteBytes Bytes written to memory

Example output:

Timestamp(s),Bandwidth(GB/s),ReadBytes,WriteBytes
0.100,45.2,2415919104,2147483648
0.200,47.8,2550136832,2281701376
0.300,46.5,2483027968,2214592512

Use Case: Application Profiling on Bandwidth-Latency Curves

A key use case is capturing your application's memory bandwidth over time, then overlaying it on Mess benchmark bandwidth-latency curves using Plotter-Parser.

Workflow

  1. Run Mess benchmark to generate bandwidth-latency curves:

    ./build/bin/mess
  2. Profile your application:

    ./build/bin/mess-profiler -s 100ms -o app_profile.csv ./my_app
  3. Visualize with app_plotter:

    python3 utils/app_plotter.py -c measuring/multisequential -p app_profile.csv -n "My App"

This shows where your application sits on the system's bandwidth-latency curves at each point in time, revealing whether it's bandwidth-bound or latency-bound.


Backend Selection

Mess Profiler automatically selects the best available backend:

Backend When Used Counter Type
perf Default on most Linux systems MBOX/IMC counters
likwid HBM systems, when perf lacks MBOX counters MBOX counters via likwid-perfctr
Intel PCM When preferred or others unavailable Memory controller counters
VTune When explicitly selected and in PATH VTune summary uncore counters

VTune Notes

  • Use VTune explicitly with --backend vtune in mess-profiler or --measurer=vtune in mess.
  • VTune support depends on the vtune binary being available in PATH.
  • VTune interval mode has much higher overhead than perf or likwid, because each sample launches a VTune collection.
  • VTune PID-attach mode is not supported in Mess Profiler.

See Also

Clone this wiki locally