Repository navigation
Mess Profiler
Mess Profiler is a convenience wrapper around performance counter tools like perf, likwid, vtune, or others that leverages the Mess benchmark's automatic counter discovery to profile application memory bandwidth.
Note: You don't need to use Mess Profiler as your actual profiler. It can serve as a counter discovery tool—run it with
--dry-runto reveal the correct counters for your system, then use those counters with your preferred profiling tool. The convenience of Mess Profiler is that it outputs data in a format that Plotter-Parser understands for plotting application profiles on bandwidth-latency curves.
Measuring memory bandwidth with perf, likwid, vtune, or others typically requires:
-
Identifying the correct counters for your specific CPU (e.g.,
uncore_imc_0/cas_count_read/) - Trial and error to find working counter names across different vendors and generations
- Parsing tool-specific output formats
Mess Profiler eliminates this overhead. The Mess benchmark already auto-discovers the correct memory bandwidth counters for your system. Mess Profiler reuses this discovery, giving you:
- Zero configuration: Works out of the box on any supported system
- Consistent output: Same CSV format regardless of underlying tool
- Correct counters: Uses the same counters validated by the Mess benchmark
- Mess Profiler runs the same counter discovery logic as the main Mess benchmark
- It identifies the appropriate profiling backend for your system
- It wraps the backend tool and parses its output into a consistent format
- It inherits CPU/memory bindings from
numactlortasksetautomatically
Use --dry-run to see exactly what counters were discovered:
./build/bin/mess-profiler --dry-runThis outputs the detected backend and counter names without running any measurements.
./build/bin/mess-profiler [options] [--] <command> [args...]| Option | Description |
|---|---|
-s, --interval <time> |
Sampling interval (e.g., 100ms, 1s). Default: summary mode |
-o, --output <file> |
Output file. Default: stdout |
| Option | Description |
|---|---|
-a, --system-wide |
System-wide profiling (all CPUs/sockets) |
-p, --pid <pid> |
Profile existing process by PID |
-C, --cpu <list> |
Profile only specified CPUs (e.g., 0-7,16-23) |
-N, --nodes <list> |
Monitor memory traffic to specified NUMA nodes |
--no-inherit |
Don't inherit binding from parent (numactl/taskset) |
| Option | Description |
|---|---|
-v, --verbose |
Verbose output (show counter details) |
--csv |
CSV output (default) |
--human |
Human-readable output |
--dry, --dry-run |
Show discovered counters and exit (no measurement) |
# See what counters mess-profiler will use on your system
./build/bin/mess-profiler --dry-run# Profile an application with 100ms sampling
./build/bin/mess-profiler -s 100ms ./my_app
# Save output to file
./build/bin/mess-profiler -s 50ms -o bandwidth.csv ./my_app# Profile all CPUs for 10 seconds
./build/bin/mess-profiler -a -s 1s sleep 10# Profile app bound to NUMA node 0, cores 0-7
numactl -m 0 -C 0-7 ./build/bin/mess-profiler -s 100ms ./my_app
# Explicit targeting (alternative)
./build/bin/mess-profiler -C 0-7 -N 0 -s 100ms ./my_appThe profiler outputs CSV with the following columns:
| Column | Description |
|---|---|
Timestamp(s) |
Time since start in seconds |
Bandwidth(GB/s) |
Measured memory bandwidth |
ReadBytes |
Bytes read from memory |
WriteBytes |
Bytes written to memory |
Example output:
Timestamp(s),Bandwidth(GB/s),ReadBytes,WriteBytes
0.100,45.2,2415919104,2147483648
0.200,47.8,2550136832,2281701376
0.300,46.5,2483027968,2214592512A key use case is capturing your application's memory bandwidth over time, then overlaying it on Mess benchmark bandwidth-latency curves using Plotter-Parser.
-
Run Mess benchmark to generate bandwidth-latency curves:
./build/bin/mess
-
Profile your application:
./build/bin/mess-profiler -s 100ms -o app_profile.csv ./my_app
-
Visualize with app_plotter:
python3 utils/app_plotter.py -c measuring/multisequential -p app_profile.csv -n "My App"
This shows where your application sits on the system's bandwidth-latency curves at each point in time, revealing whether it's bandwidth-bound or latency-bound.
Mess Profiler automatically selects the best available backend:
| Backend | When Used | Counter Type |
|---|---|---|
| perf | Default on most Linux systems | MBOX/IMC counters |
| likwid | HBM systems, when perf lacks MBOX counters | MBOX counters via likwid-perfctr |
| Intel PCM | When preferred or others unavailable | Memory controller counters |
| VTune | When explicitly selected and in PATH
|
VTune summary uncore counters |
- Use VTune explicitly with
--backend vtuneinmess-profileror--measurer=vtuneinmess. - VTune support depends on the
vtunebinary being available inPATH. - VTune interval mode has much higher overhead than
perforlikwid, because each sample launches a VTune collection. - VTune PID-attach mode is not supported in Mess Profiler.
- Mess-Benchmark - Core benchmark (uses same counter discovery)
- Plotter - Visualizing profiler output on bandwidth-latency curves
- Understand output - Output file formats