Skip to content

Traffic Generator

Pau Díaz Cuesta edited this page Sep 21, 2026 · 3 revisions

The Traffic Generator is a standalone tool that generates heavy memory bandwidth traffic using pure assembly code. It is the engine behind Mess's bandwidth saturation and can also be used independently for your own testing.


Prerequisites

Important: Before diving into the Traffic Generator, it is crucial to understand the distinction between software instructions and hardware operations. Please read Load/Store vs Read/Write.

For detailed customization of the generation kernels, refer to Traffic generator setup.

Overview

Unlike traditional bandwidth benchmarks (like STREAM), the Mess Traffic Generator:

  • Generates pure assembly code - no compiler overhead, maximum control
  • Uses SIMD instructions - wide vector loads/stores for maximum throughput
  • Supports configurable load/store ratios - test different read/write mixes
  • Implements precise bandwidth control - via NOP-based pause bubbles

The Traffic Generator is primarily used internally by the Mess benchmark, but can be run standalone when you need raw bandwidth saturation without the full benchmark overhead.


How It Works

Assembly Kernels

The Traffic Generator produces tight assembly loops that perform:

  1. SIMD memory operations - Vector loads and stores using the widest available registers (AVX-512, SVE, etc.)
  2. Configurable ratios - Mix of load and store instructions based on the requested ratio
  3. NOP bubbles - Sequences of NOP instructions inserted between memory operations to reduce bandwidth

Array Size Calibration

Mess calibrates the traffic-generator working-set size during generate_code, before the final binary is installed.

Why it matters

The array must be large enough to stay out of LLC so generated traffic hits DRAM. Too small and bandwidth looks artificially low (cache hits), too large wastes memory and slows setup without improving saturation quality.

How calibration works

  1. Build a capacity binary sized for the largest candidate that still fits in available memory.
  2. Probe candidate sizes: typically 4×, 5×, 6×, 8×, and 12× LLC, plus 20× when larger, and 128× when LLC is ≤ 16 MB. Each probe runs a single-worker (-w 1), all-reads (-r 100), pause-0 kernel and records self-reported bandwidth.
  3. Select a plateau: Mess looks for sustained near-peak DRAM bandwidth across successive sizes.
  4. Recompile the final traffic-generator binary with the selected size baked in as TrafficGen_ARRAY_SIZE.

If probing fails, Mess falls back to 4× LLC (or 128× LLC on small-LLC embedded systems).

Pause Mechanism

The "pause" value controls bandwidth saturation by inserting NOP instructions (bubbles) between memory operations:

pause=0:    [LOAD][STORE][LOAD][STORE]...        → Maximum bandwidth
pause=10:   [LOAD][STORE][NOP×10][LOAD][STORE]... → Reduced bandwidth
pause=1000: [LOAD][STORE][NOP×1000]...           → Very low bandwidth

By varying the pause value, Mess generates different points on the bandwidth-latency curve.

In normal benchmark runs, Mess can choose these pause values automatically through Adaptive Curve-Guided Pause Discovery. The traffic generator itself only receives the final pause value for a point and implements it as assembly-level throttling.

Ratio Control

The ratio parameter controls the mix of load vs store operations:

Ratio Loads Stores Pattern
100 100% 0% [LOAD][LOAD][LOAD]...
75 75% 25% [LOAD][LOAD][LOAD][STORE]...
50 50% 50% [LOAD][STORE][LOAD][STORE]...
0 0% 100% [STORE][STORE][STORE]...

Architecture Support

The Traffic Generator includes ISA-specific assembly kernels:

Architecture SIMD Vector Width Instructions
x86-64 AVX2 256-bit vmovaps, vmovntps
x86-64 AVX-512 512-bit vmovaps zmm, vmovntps zmm
ARM NEON_NATIVE 128-bit ldr q, str q
ARM NEON_PAIR 256-bit logical pair ldp q,q, stp q,q
ARM SVE Variable ld1d, st1d, stnt1d
Power VSX 128-bit lxvd2x, stxvd2x
RISC-V RVV Variable vle, vse

The appropriate kernel is selected automatically at compile time.

Again, to tune this please refer to Traffic generator setup.

ARM Notes

On ARM, Mess distinguishes between ISA-native NEON traffic and paired-register NEON traffic:

  • NEON_NATIVE uses one 128-bit SIMD register per memory operation (ldr q, str q)
  • NEON_PAIR uses paired 128-bit SIMD registers per memory operation (ldp q,q, stp q,q)

AUTO currently resolves to NEON_PAIR on ARM.

For SVE, Mess supports:

  • SVE: use the current thread VL
  • SVE_MAX: request the maximum thread VL
  • SVE128, SVE256, SVE512: request fixed VLs for controlled comparisons

The SVE assembly mnemonic family remains the same across those modes. The active thread VL is configured at benchmark runtime, and the generated kernel uses ptrue p0.d together with ld1d, st1d, and stnt1d.

Mess also supports two ARM addressing styles:

  • POST_INCREMENT: use the current base pointer and then advance it
  • INDEXED: compute the effective address explicitly from base plus offset

For SVE sequential post-increment, pointer motion is VL-aware via addvl.


Standalone Usage

The Traffic Generator can be run independently for custom testing:

./build/bin/traffic_gen_multiseq.x [options]

Options

Option Description
-r N Load/store ratio (0-100, where 100 is all reads)
-p N NOP bubble size (cycles)
-w N Workers (number of array chunks/threads)
-v N Verbosity (0 = quiet, 1 = debug logs)

Example: Generate Maximum Read Traffic

# 100% Reads, 0 pause, 1 worker, verbose output
./build/bin/traffic_gen_multiseq.x -r 100 -p 0 -w 1 -v 1

Example: Mixed Traffic with Pause

# 70% Reads, 50 cycle pause, 4 workers, quiet
./build/bin/traffic_gen_multiseq.x -r 70 -p 50 -w 4 -v 0

Example: Generate Controlled Write Traffic

# 50% bandwidth with pure writes
./build/bin/traffic_generator --ratio=0 --pause=500 --duration=10

Use Cases for Standalone Mode

  1. System stress testing - Generate sustained memory load
  2. Thermal testing - CPU/memory thermal behavior under load
  3. Power measurements - Measure power consumption at different bandwidth levels
  4. Custom experiments - When you need traffic generation without the full Mess benchmark

Kernel Generation

The assembly kernels are generated at compile time based on include/Codegen.h and include/KernelTypes.h. Key parameters:

Parameter Description
ratio_granularity Precision of load/store ratio (default: 2%)
ops_per_pause_block Memory ops between pause insertions
num_simd_registers Registers to use (affects unrolling)
use_nontemporal_stores Use streaming stores (default: false)

See Traffic Generator Setup for advanced configuration.


Inspecting the Assembly Code

To examine the actual assembly instructions generated for the Traffic Generator:

  1. Compile the benchmark (make)
  2. Navigate to src/traffic_gen/src/

In this directory, you will find the generated utility files containing:

  • Assembly kernels for each load/store ratio combination
  • NOP loops used for implementing pause values
  • Architecture-specific optimizations selected at compile time

This allows you to dive into the exact instructions being executed, verify the kernel structure, or understand how the pause mechanism is implemented at the assembly level.

For customization beyond the default values, see Traffic generator setup.


See Also

Clone this wiki locally