Repository navigation
Traffic Generator
The Traffic Generator is a standalone tool that generates heavy memory bandwidth traffic using pure assembly code. It is the engine behind Mess's bandwidth saturation and can also be used independently for your own testing.
Important: Before diving into the Traffic Generator, it is crucial to understand the distinction between software instructions and hardware operations. Please read Load/Store vs Read/Write.
For detailed customization of the generation kernels, refer to Traffic generator setup.
Unlike traditional bandwidth benchmarks (like STREAM), the Mess Traffic Generator:
- Generates pure assembly code - no compiler overhead, maximum control
- Uses SIMD instructions - wide vector loads/stores for maximum throughput
- Supports configurable load/store ratios - test different read/write mixes
- Implements precise bandwidth control - via NOP-based pause bubbles
The Traffic Generator is primarily used internally by the Mess benchmark, but can be run standalone when you need raw bandwidth saturation without the full benchmark overhead.
The Traffic Generator produces tight assembly loops that perform:
- SIMD memory operations - Vector loads and stores using the widest available registers (AVX-512, SVE, etc.)
- Configurable ratios - Mix of load and store instructions based on the requested ratio
- NOP bubbles - Sequences of NOP instructions inserted between memory operations to reduce bandwidth
Mess calibrates the traffic-generator working-set size during generate_code, before the final binary is installed.
The array must be large enough to stay out of LLC so generated traffic hits DRAM. Too small and bandwidth looks artificially low (cache hits), too large wastes memory and slows setup without improving saturation quality.
- Build a capacity binary sized for the largest candidate that still fits in available memory.
-
Probe candidate sizes: typically 4×, 5×, 6×, 8×, and 12× LLC, plus 20× when larger, and 128× when LLC is ≤ 16 MB. Each probe runs a single-worker (
-w 1), all-reads (-r 100), pause-0 kernel and records self-reported bandwidth. - Select a plateau: Mess looks for sustained near-peak DRAM bandwidth across successive sizes.
-
Recompile the final traffic-generator binary with the selected size baked in as
TrafficGen_ARRAY_SIZE.
If probing fails, Mess falls back to 4× LLC (or 128× LLC on small-LLC embedded systems).
The "pause" value controls bandwidth saturation by inserting NOP instructions (bubbles) between memory operations:
pause=0: [LOAD][STORE][LOAD][STORE]... → Maximum bandwidth
pause=10: [LOAD][STORE][NOP×10][LOAD][STORE]... → Reduced bandwidth
pause=1000: [LOAD][STORE][NOP×1000]... → Very low bandwidth
By varying the pause value, Mess generates different points on the bandwidth-latency curve.
In normal benchmark runs, Mess can choose these pause values automatically through Adaptive Curve-Guided Pause Discovery. The traffic generator itself only receives the final pause value for a point and implements it as assembly-level throttling.
The ratio parameter controls the mix of load vs store operations:
| Ratio | Loads | Stores | Pattern |
|---|---|---|---|
| 100 | 100% | 0% | [LOAD][LOAD][LOAD]... |
| 75 | 75% | 25% | [LOAD][LOAD][LOAD][STORE]... |
| 50 | 50% | 50% | [LOAD][STORE][LOAD][STORE]... |
| 0 | 0% | 100% | [STORE][STORE][STORE]... |
The Traffic Generator includes ISA-specific assembly kernels:
| Architecture | SIMD | Vector Width | Instructions |
|---|---|---|---|
| x86-64 | AVX2 | 256-bit |
vmovaps, vmovntps
|
| x86-64 | AVX-512 | 512-bit |
vmovaps zmm, vmovntps zmm
|
| ARM | NEON_NATIVE | 128-bit |
ldr q, str q
|
| ARM | NEON_PAIR | 256-bit logical pair |
ldp q,q, stp q,q
|
| ARM | SVE | Variable |
ld1d, st1d, stnt1d
|
| Power | VSX | 128-bit |
lxvd2x, stxvd2x
|
| RISC-V | RVV | Variable |
vle, vse
|
The appropriate kernel is selected automatically at compile time.
Again, to tune this please refer to Traffic generator setup.
On ARM, Mess distinguishes between ISA-native NEON traffic and paired-register NEON traffic:
-
NEON_NATIVE uses one 128-bit SIMD register per memory operation (
ldr q,str q) -
NEON_PAIR uses paired 128-bit SIMD registers per memory operation (
ldp q,q,stp q,q)
AUTO currently resolves to NEON_PAIR on ARM.
For SVE, Mess supports:
-
SVE: use the current thread VL -
SVE_MAX: request the maximum thread VL -
SVE128,SVE256,SVE512: request fixed VLs for controlled comparisons
The SVE assembly mnemonic family remains the same across those modes. The active thread VL is configured at benchmark runtime, and the generated kernel uses ptrue p0.d together with ld1d, st1d, and stnt1d.
Mess also supports two ARM addressing styles:
- POST_INCREMENT: use the current base pointer and then advance it
- INDEXED: compute the effective address explicitly from base plus offset
For SVE sequential post-increment, pointer motion is VL-aware via addvl.
The Traffic Generator can be run independently for custom testing:
./build/bin/traffic_gen_multiseq.x [options]| Option | Description |
|---|---|
-r N |
Load/store ratio (0-100, where 100 is all reads) |
-p N |
NOP bubble size (cycles) |
-w N |
Workers (number of array chunks/threads) |
-v N |
Verbosity (0 = quiet, 1 = debug logs) |
# 100% Reads, 0 pause, 1 worker, verbose output
./build/bin/traffic_gen_multiseq.x -r 100 -p 0 -w 1 -v 1# 70% Reads, 50 cycle pause, 4 workers, quiet
./build/bin/traffic_gen_multiseq.x -r 70 -p 50 -w 4 -v 0# 50% bandwidth with pure writes
./build/bin/traffic_generator --ratio=0 --pause=500 --duration=10- System stress testing - Generate sustained memory load
- Thermal testing - CPU/memory thermal behavior under load
- Power measurements - Measure power consumption at different bandwidth levels
- Custom experiments - When you need traffic generation without the full Mess benchmark
The assembly kernels are generated at compile time based on include/Codegen.h and include/KernelTypes.h. Key parameters:
| Parameter | Description |
|---|---|
ratio_granularity |
Precision of load/store ratio (default: 2%) |
ops_per_pause_block |
Memory ops between pause insertions |
num_simd_registers |
Registers to use (affects unrolling) |
use_nontemporal_stores |
Use streaming stores (default: false) |
See Traffic Generator Setup for advanced configuration.
To examine the actual assembly instructions generated for the Traffic Generator:
-
Compile the benchmark (
make) -
Navigate to
src/traffic_gen/src/
In this directory, you will find the generated utility files containing:
- Assembly kernels for each load/store ratio combination
- NOP loops used for implementing pause values
- Architecture-specific optimizations selected at compile time
This allows you to dive into the exact instructions being executed, verify the kernel structure, or understand how the pause mechanism is implemented at the assembly level.
For customization beyond the default values, see Traffic generator setup.
- Mess Benchmark - Full benchmark using Traffic Generator
- Adaptive Curve-Guided Pause Discovery - Automatic benchmark pause selection
- Traffic generator setup - Advanced kernel configuration
- Temporal vs Non-Temporal stores - Store instruction types
- Architecture Support - Platform-specific details