Repository navigation
Understanding CLI Arguments
This page provides a comprehensive reference for all Mess Benchmark CLI options and flags.
| Category | Key Options |
|---|---|
| Information |
--help, --version, --dry-run
|
| Profiling |
--no-profile, --folder=<path>
|
| Core Binding |
--total-cores=<N>, --cores=<list>, --bind=<nodes>
|
| Ratio Control | --ratio=<0-100> |
| Pause Control |
--pause=<values>, --tier=<name>, --point-count=<N>
|
| Measurement |
--repetitions=<N>, --measurer=<type>, --add-counters=<list>, --inst-lat
|
| Output | --verbose=<0-3> |
| Advanced | --persistent-trafficgen |
Display usage information and all available options.
./build/bin/mess --helpDisplay version information including build configuration.
./build/bin/mess --versionPreview configuration without running measurements. Shows:
- Detected architecture and ISA
- CPU topology and NUMA configuration
- Selected performance counter backend
- Planned measurement configurations
./build/bin/mess --dry-run
./build/bin/mess --dry-run --verbose=2 # More detailProfiling is enabled by default for CPU and GPU runs. Mess saves raw bandwidth
and latency measurements under measuring/ (or the directory selected with
--folder), so the results remain available for plotting and inspection.
./build/bin/mess # Save measurements (default)
./build/bin/mess --no-profile # Run without retaining measurement files
./build/bin/mess --gpu # GPU measurements are also saved by default
./build/bin/mess --gpu --no-profile # Discard GPU measurement files after the run--no-profile uses temporary measurement files and removes them after the run.
It does not disable the benchmark or hardware counters, and it does not delete
results saved by earlier runs. It replaces the old opt-in --profile flag.
Specify output directory for measurement results. Default: measuring/
./build/bin/mess --folder=my_results
./build/bin/mess --folder=numa0_testMultiple runs with different --folder values allow easy organization and comparison:
./build/bin/mess --folder=baseline
./build/bin/mess --folder=after_tuningSet total number of traffic generator cores to use. Cores are selected automatically from available CPUs.
./build/bin/mess --total-cores=8
./build/bin/mess --total-cores=4 # Fewer cores = less bandwidth saturationExplicitly specify which cores to use for traffic generation. Accepts:
- Single core:
--cores=0 - Range:
--cores=0-7 - List:
--cores=0,2,4,6 - Mixed:
--cores=0-3,8-11
# Use only cores 0-7
./build/bin/mess --cores=0-7
# Use even-numbered cores (avoid hyperthreads)
./build/bin/mess --cores=0,2,4,6,8,10,12,14
# First socket only (example topology)
./build/bin/mess --cores=0-15Bind memory allocation to specific NUMA nodes. Accepts:
- Single node:
--bind=0 - Multiple nodes:
--bind=0,1
# Check NUMA topology first
numactl -H
# Local memory access (cores and memory on same node)
./build/bin/mess --cores=0-7 --bind=0
# Remote memory access (cross-NUMA)
./build/bin/mess --cores=0-7 --bind=1
# Interleaved across nodes
./build/bin/mess --bind=0,1Set the load/store ratio (percentage of load operations). Accepts a single value or comma-separated list:
| Value | Meaning |
|---|---|
100 |
100% loads, 0% stores (pure reads) |
75 |
75% loads, 25% stores |
50 |
50% loads, 50% stores (balanced) |
25 |
25% loads, 75% stores |
0 |
0% loads, 100% stores (pure writes) |
# Pure read workload
./build/bin/mess --ratio=100
# Balanced read/write
./build/bin/mess --ratio=50
# Pure write workload
./build/bin/mess --ratio=0
# Multiple ratios in one run
./build/bin/mess --ratio=100,75,50,25,0By default, Mess sweeps through multiple ratios. Use --ratio to focus on a specific mix or set of mixes.
When --pause is omitted, Mess uses Adaptive Curve-Guided Pause Discovery to select pause values automatically.
The pause parameter controls bandwidth saturation by inserting NOP instruction bubbles between memory operations:
-
pause=0- Maximum bandwidth (no NOPs between operations) - Higher values - Reduced bandwidth (more NOPs = more idle time)
Accepts:
- Single value:
--pause=100 - Comma-separated list:
--pause=0,100,1000,10000
# Single point at maximum bandwidth
./build/bin/mess --ratio=100 --pause=0
# Logarithmic spacing for full bandwidth-latency curve
./build/bin/mess --pause=0,10,100,1000,10000,100000
# Dense sampling around region of interest
./build/bin/mess --pause=0,10,20,30,40,50,60,70,80,90,100
# Sparse sampling for quick overview
./build/bin/mess --pause=0,500,5000,50000Providing --pause disables adaptive pause discovery and measures exactly the listed points. This is particularly useful for Iterative Debugging - start with sparse pause values, then add more points in regions of interest.
Select the adaptive pause discovery preset used when --pause is not provided:
| Tier | Point budget | Notes |
|---|---|---|
lite |
15 points per enabled mode | Fast exploratory runs; defaults to one repetition unless --repetitions is set |
standard |
50 points per enabled mode | Default |
detailed |
200 points per enabled mode | Higher-resolution curves |
./build/bin/mess --tier=lite
./build/bin/mess --tier=standard
./build/bin/mess --tier=detailedOverride the adaptive point budget with a custom value. This is useful when the built-in tiers are too small or too large for a specific machine or queue allocation.
./build/bin/mess --ratio=100 --point-count=75--point-count cannot be combined with --tier or --pause.
Number of repetitions for each measurement point. More repetitions reduce variance but increase runtime.
| Value | Effect |
|---|---|
1 |
Fastest, highest variance |
3 |
Default, good balance |
5-10 |
High stability, longer runtime |
./build/bin/mess --repetitions=1 # Quick
./build/bin/mess --repetitions=5 # Stable
./build/bin/mess --repetitions=10 # Very stableSelect the performance counter backend for bandwidth measurement. Mess automatically detects available tools, but you can explicitly choose:
| Type | Description | Best For |
|---|---|---|
auto |
Automatically select best available tool (default) | General use |
perf |
Linux perf (low overhead, standard counters) | DDR systems, general profiling |
likwid |
LIKWID (better HBM support) | HBM systems (required for HBM) |
vtune |
Intel VTune (driverless uncore summary collection) | Explicit VTune workflows |
pcm |
Intel PCM (CXL and Intel-specific features) | CXL memory, Intel systems |
Why explicit selection?
-
perf: Low overhead, works on most systems with standard DDR memory -
likwid: Required for HBM systems where perf maps CAS_COUNT to wrong counters (MBOX instead of HBM registers) -
vtune: Optional explicit backend when VTune is installed and visible inPATH; higher overhead than perf/likwid -
pcm: Required for CXL memory systems
# Auto-detect (default behavior)
./build/bin/mess
# Force LIKWID for HBM systems
./build/bin/mess --measurer=likwid
# Use perf explicitly
./build/bin/mess --measurer=perf
# Use VTune explicitly
./build/bin/mess --measurer=vtune
# Check which tool will be used
./build/bin/mess --measurer=likwid --dry-runMigration from old syntax:
-
--likwidis deprecated. Use--measurer=likwidinstead.
Add extra performance counters to bandwidth measurements. These counters are recorded alongside bandwidth/latency data and can be used for deeper performance analysis. These are automatically aggregated for all the working cores, providing a whole picture of system's performance metrics.
# Add instruction and cycle counts
./build/bin/mess --add-counters=cycles,instructions
# Add cache statistics
./build/bin/mess --add-counters=cache-misses,cache-references
# Multiple counters
./build/bin/mess --add-counters=instructions,cycles,cache-missesAdditional performance counters are measured using the selected tool (e.g., perf, likwid).
Available counters depend on your system. Check with:
perf list Or if using likwid as measurer:
likwid-perfctr -eNote: When measuring with likwid, inputs of add-counters must be event names. Registers and counters will be automatically detected:
# Measure L3 Bandwidth
./build/bin/mess --measurer=likwid --add-counters=L2_LINES_IN_ALL,L2_LINES_OUT_NON_SILENTOr if using VTune as measurer:
vtune -collect-with runsa -knob event-config=? -- /bin/trueNote: When measuring with VTune, inputs of
--add-countersmust be VTune-visible event names.
Migration from old syntax:
-
--extra-perfis deprecated. Use--add-countersinstead.
Enable instruction-level latency sampling using hardware sampling mechanisms: PEBS on Intel and SPE on ARM. This captures latency samples from traffic generator load instructions instead of using a separate pointer-chase process.
# Enable instruction latency sampling
./build/bin/mess --inst-lat
# Combine with adaptive pause discovery
./build/bin/mess --inst-lat --ratio=100 --tier=standard
# Combine with a manual pause list
./build/bin/mess --inst-lat --ratio=100 --pause=0,100,1000See Instruction Sampling for backend requirements, filtering, and output details.
Control output verbosity:
| Level | Output |
|---|---|
0 |
Minimal (errors only) |
1 |
Progress bar, ETA, summary (default) |
2 |
Configuration details, execution steps |
3 |
Raw measurements, subprocess commands, debug data |
4 |
Extra counter detail when profiling |
./build/bin/mess --verbose=0 # Silent
./build/bin/mess --verbose=1 # Progress (default)
./build/bin/mess --verbose=2 # Detailed
./build/bin/mess --verbose=3 # Debug
./build/bin/mess --verbose=4 # Extra counter detailFor troubleshooting, our recomendation is to always use --verbose=3:
./build/bin/mess --verbose=3 2>&1 | tee debug.logThese options are for expert tuning. Use with caution.
Keep the traffic generator processes alive across repetitions for the same ratio and pause value. This increases measurement speed by avoiding process spawn/teardown overhead.
# Faster runs with persistent traffic generator
./build/bin/mess --persistent-trafficgenDefault behavior (without this flag): The traffic generator is killed and respawned between each repetition to ensure statistically independent experiments.
With this flag: The traffic generator stays alive when measuring multiple repetitions of the same (ratio, pause) point. It is only killed when moving to a different configuration.
Use cases:
- Quick iterative testing where speed matters more than strict independence
- Systems where process spawning has high overhead
Trade-off: Measurements within the same point may have subtle correlations since the traffic generator maintains its internal state across repetitions.
# Verify detection
./build/bin/mess --dry-run --verbose=2
# Fast test run
./build/bin/mess --ratio=100# Profile each NUMA node separately
./build/bin/mess --bind=0 --folder=numa0
./build/bin/mess --bind=1 --folder=numa1
# Cross-NUMA (remote memory)
./build/bin/mess --cores=0-7 --bind=1 --folder=cross_numafor cores in 2 4 8 16; do
./build/bin/mess --total-cores=$cores --folder=cores_$cores
done./build/bin/mess --ratio=100 --folder=reads_only
./build/bin/mess --ratio=0 --folder=writes_only
./build/bin/mess --ratio=50 --folder=balanced# Full resolution, high stability
./build/bin/mess --repetitions=5For consistent measurements:
# Disable frequency scaling
sudo cpupower frequency-set -g performance
# Check current governor
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
# Reduce perf paranoid level
echo 0 | sudo tee /proc/sys/kernel/perf_event_paranoid- Iterative Debugging - Incremental refinement workflow
- Mess Benchmark - Benchmark methodology overview
- Understand output - Output file formats