Skip to content

Understanding CLI Arguments

Victor Xirau Guardans edited this page Sep 29, 2026 · 6 revisions

This page provides a comprehensive reference for all Mess Benchmark CLI options and flags.


Quick Reference

Category Key Options
Information --help, --version, --dry-run
Profiling --no-profile, --folder=<path>
Core Binding --total-cores=<N>, --cores=<list>, --bind=<nodes>
Ratio Control --ratio=<0-100>
Pause Control --pause=<values>, --tier=<name>, --point-count=<N>
Measurement --repetitions=<N>, --measurer=<type>, --add-counters=<list>, --inst-lat
Output --verbose=<0-3>
Advanced --persistent-trafficgen

Information Options

--help, -h

Display usage information and all available options.

./build/bin/mess --help

--version, -v

Display version information including build configuration.

./build/bin/mess --version

--dry-run

Preview configuration without running measurements. Shows:

  • Detected architecture and ISA
  • CPU topology and NUMA configuration
  • Selected performance counter backend
  • Planned measurement configurations
./build/bin/mess --dry-run
./build/bin/mess --dry-run --verbose=2  # More detail

Profiling and Output Options

--no-profile

Profiling is enabled by default for CPU and GPU runs. Mess saves raw bandwidth and latency measurements under measuring/ (or the directory selected with --folder), so the results remain available for plotting and inspection.

./build/bin/mess                    # Save measurements (default)
./build/bin/mess --no-profile       # Run without retaining measurement files
./build/bin/mess --gpu              # GPU measurements are also saved by default
./build/bin/mess --gpu --no-profile # Discard GPU measurement files after the run

--no-profile uses temporary measurement files and removes them after the run. It does not disable the benchmark or hardware counters, and it does not delete results saved by earlier runs. It replaces the old opt-in --profile flag.

--folder=<path>

Specify output directory for measurement results. Default: measuring/

./build/bin/mess --folder=my_results
./build/bin/mess --folder=numa0_test

Multiple runs with different --folder values allow easy organization and comparison:

./build/bin/mess --folder=baseline
./build/bin/mess --folder=after_tuning

Core and NUMA Binding

--total-cores=<N>

Set total number of traffic generator cores to use. Cores are selected automatically from available CPUs.

./build/bin/mess --total-cores=8
./build/bin/mess --total-cores=4  # Fewer cores = less bandwidth saturation

--cores=<list>

Explicitly specify which cores to use for traffic generation. Accepts:

  • Single core: --cores=0
  • Range: --cores=0-7
  • List: --cores=0,2,4,6
  • Mixed: --cores=0-3,8-11
# Use only cores 0-7
./build/bin/mess --cores=0-7

# Use even-numbered cores (avoid hyperthreads)
./build/bin/mess --cores=0,2,4,6,8,10,12,14

# First socket only (example topology)
./build/bin/mess --cores=0-15

--bind=<nodes>

Bind memory allocation to specific NUMA nodes. Accepts:

  • Single node: --bind=0
  • Multiple nodes: --bind=0,1
# Check NUMA topology first
numactl -H

# Local memory access (cores and memory on same node)
./build/bin/mess --cores=0-7 --bind=0

# Remote memory access (cross-NUMA)
./build/bin/mess --cores=0-7 --bind=1

# Interleaved across nodes
./build/bin/mess --bind=0,1

Ratio Control

--ratio=<value>

Set the load/store ratio (percentage of load operations). Accepts a single value or comma-separated list:

Value Meaning
100 100% loads, 0% stores (pure reads)
75 75% loads, 25% stores
50 50% loads, 50% stores (balanced)
25 25% loads, 75% stores
0 0% loads, 100% stores (pure writes)
# Pure read workload
./build/bin/mess --ratio=100

# Balanced read/write
./build/bin/mess --ratio=50

# Pure write workload
./build/bin/mess --ratio=0

# Multiple ratios in one run
./build/bin/mess --ratio=100,75,50,25,0

By default, Mess sweeps through multiple ratios. Use --ratio to focus on a specific mix or set of mixes.


Pause Control

When --pause is omitted, Mess uses Adaptive Curve-Guided Pause Discovery to select pause values automatically.

--pause=<values>

The pause parameter controls bandwidth saturation by inserting NOP instruction bubbles between memory operations:

  • pause=0 - Maximum bandwidth (no NOPs between operations)
  • Higher values - Reduced bandwidth (more NOPs = more idle time)

Accepts:

  • Single value: --pause=100
  • Comma-separated list: --pause=0,100,1000,10000
# Single point at maximum bandwidth
./build/bin/mess --ratio=100 --pause=0

# Logarithmic spacing for full bandwidth-latency curve
./build/bin/mess --pause=0,10,100,1000,10000,100000

# Dense sampling around region of interest
./build/bin/mess --pause=0,10,20,30,40,50,60,70,80,90,100

# Sparse sampling for quick overview
./build/bin/mess --pause=0,500,5000,50000

Providing --pause disables adaptive pause discovery and measures exactly the listed points. This is particularly useful for Iterative Debugging - start with sparse pause values, then add more points in regions of interest.

--tier=<lite|standard|detailed>

Select the adaptive pause discovery preset used when --pause is not provided:

Tier Point budget Notes
lite 15 points per enabled mode Fast exploratory runs; defaults to one repetition unless --repetitions is set
standard 50 points per enabled mode Default
detailed 200 points per enabled mode Higher-resolution curves
./build/bin/mess --tier=lite
./build/bin/mess --tier=standard
./build/bin/mess --tier=detailed

--point-count=<N>

Override the adaptive point budget with a custom value. This is useful when the built-in tiers are too small or too large for a specific machine or queue allocation.

./build/bin/mess --ratio=100 --point-count=75

--point-count cannot be combined with --tier or --pause.


Measurement Options

--repetitions=<N>

Number of repetitions for each measurement point. More repetitions reduce variance but increase runtime.

Value Effect
1 Fastest, highest variance
3 Default, good balance
5-10 High stability, longer runtime
./build/bin/mess --repetitions=1   # Quick
./build/bin/mess --repetitions=5   # Stable
./build/bin/mess --repetitions=10  # Very stable

--measurer=<type>

Select the performance counter backend for bandwidth measurement. Mess automatically detects available tools, but you can explicitly choose:

Type Description Best For
auto Automatically select best available tool (default) General use
perf Linux perf (low overhead, standard counters) DDR systems, general profiling
likwid LIKWID (better HBM support) HBM systems (required for HBM)
vtune Intel VTune (driverless uncore summary collection) Explicit VTune workflows
pcm Intel PCM (CXL and Intel-specific features) CXL memory, Intel systems

Why explicit selection?

  • perf: Low overhead, works on most systems with standard DDR memory
  • likwid: Required for HBM systems where perf maps CAS_COUNT to wrong counters (MBOX instead of HBM registers)
  • vtune: Optional explicit backend when VTune is installed and visible in PATH; higher overhead than perf/likwid
  • pcm: Required for CXL memory systems
# Auto-detect (default behavior)
./build/bin/mess

# Force LIKWID for HBM systems
./build/bin/mess --measurer=likwid

# Use perf explicitly
./build/bin/mess --measurer=perf

# Use VTune explicitly
./build/bin/mess --measurer=vtune

# Check which tool will be used
./build/bin/mess --measurer=likwid --dry-run

Migration from old syntax:

  • --likwid is deprecated. Use --measurer=likwid instead.

--add-counters=<counters>

Add extra performance counters to bandwidth measurements. These counters are recorded alongside bandwidth/latency data and can be used for deeper performance analysis. These are automatically aggregated for all the working cores, providing a whole picture of system's performance metrics.

# Add instruction and cycle counts
./build/bin/mess --add-counters=cycles,instructions

# Add cache statistics
./build/bin/mess --add-counters=cache-misses,cache-references

# Multiple counters
./build/bin/mess --add-counters=instructions,cycles,cache-misses

Additional performance counters are measured using the selected tool (e.g., perf, likwid).

Available counters depend on your system. Check with:

perf list 

Or if using likwid as measurer:

likwid-perfctr -e

Note: When measuring with likwid, inputs of add-counters must be event names. Registers and counters will be automatically detected:

# Measure L3 Bandwidth
./build/bin/mess --measurer=likwid --add-counters=L2_LINES_IN_ALL,L2_LINES_OUT_NON_SILENT

Or if using VTune as measurer:

vtune -collect-with runsa -knob event-config=? -- /bin/true

Note: When measuring with VTune, inputs of --add-counters must be VTune-visible event names.

Migration from old syntax:

  • --extra-perf is deprecated. Use --add-counters instead.

--inst-lat

Enable instruction-level latency sampling using hardware sampling mechanisms: PEBS on Intel and SPE on ARM. This captures latency samples from traffic generator load instructions instead of using a separate pointer-chase process.

# Enable instruction latency sampling
./build/bin/mess --inst-lat

# Combine with adaptive pause discovery
./build/bin/mess --inst-lat --ratio=100 --tier=standard

# Combine with a manual pause list
./build/bin/mess --inst-lat --ratio=100 --pause=0,100,1000

See Instruction Sampling for backend requirements, filtering, and output details.


Output and Debugging

--verbose=<level>

Control output verbosity:

Level Output
0 Minimal (errors only)
1 Progress bar, ETA, summary (default)
2 Configuration details, execution steps
3 Raw measurements, subprocess commands, debug data
4 Extra counter detail when profiling
./build/bin/mess --verbose=0  # Silent
./build/bin/mess --verbose=1  # Progress (default)
./build/bin/mess --verbose=2  # Detailed
./build/bin/mess --verbose=3  # Debug
./build/bin/mess --verbose=4  # Extra counter detail

For troubleshooting, our recomendation is to always use --verbose=3:

./build/bin/mess --verbose=3 2>&1 | tee debug.log

Advanced Options

These options are for expert tuning. Use with caution.

--persistent-trafficgen

Keep the traffic generator processes alive across repetitions for the same ratio and pause value. This increases measurement speed by avoiding process spawn/teardown overhead.

# Faster runs with persistent traffic generator
./build/bin/mess --persistent-trafficgen

Default behavior (without this flag): The traffic generator is killed and respawned between each repetition to ensure statistically independent experiments.

With this flag: The traffic generator stays alive when measuring multiple repetitions of the same (ratio, pause) point. It is only killed when moving to a different configuration.

Use cases:

  • Quick iterative testing where speed matters more than strict independence
  • Systems where process spawning has high overhead

Trade-off: Measurements within the same point may have subtle correlations since the traffic generator maintains its internal state across repetitions.


Common Workflows

Quick System Validation

# Verify detection
./build/bin/mess --dry-run --verbose=2

# Fast test run
./build/bin/mess --ratio=100

NUMA Comparison

# Profile each NUMA node separately
./build/bin/mess --bind=0 --folder=numa0
./build/bin/mess --bind=1 --folder=numa1

# Cross-NUMA (remote memory)
./build/bin/mess --cores=0-7 --bind=1 --folder=cross_numa

Core Scaling Study

for cores in 2 4 8 16; do
    ./build/bin/mess --total-cores=$cores --folder=cores_$cores
done

Read/Write Comparison

./build/bin/mess --ratio=100 --folder=reads_only
./build/bin/mess --ratio=0 --folder=writes_only
./build/bin/mess --ratio=50 --folder=balanced

Production Run

# Full resolution, high stability
./build/bin/mess --repetitions=5

Environment Recommendations

For consistent measurements:

# Disable frequency scaling
sudo cpupower frequency-set -g performance

# Check current governor
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor

# Reduce perf paranoid level
echo 0 | sudo tee /proc/sys/kernel/perf_event_paranoid

See Also

Clone this wiki locally