Skip to content

Traffic Generator Setup

Victor Xirau Guardans edited this page May 28, 2026 · 4 revisions

The traffic generator kernel is configured through include/KernelTypes.h. These settings control how assembly code is generated for memory traffic operations.

For an explanation of what the Traffic Generator binary does, see Traffic Generator.

Here we explain how users can tune the kernel, since it's programmatically generated per-architecture.

Note: Modifying these settings requires recompilation. Most users should not need to change defaults.


KernelConfigMultiSeq Parameters

The KernelConfigMultiSeq struct in include/KernelTypes.h controls the MultiSequential traffic generator behavior:

ratio_granularity

  • Default: 2
  • Purpose: Controls the precision of issued load ratio in percentage steps
  • Details: By default, we use steps of 2% for granularity of issued load ratio. This can be tuned for finer or coarser control over the load/store ratio.

ops_per_pause_block

  • Default: 6
  • Purpose: Number of memory operations executed between pause points
  • Details: The benchmark repeats the structure of "memops - nops - memops - nops". This flag customizes the amount of memory operations before a NOP block.

total_ops

  • Default: 100
  • Purpose: Total operations per kernel iteration
  • Details: The kernel loops between these operations. We keep it at 100 for easy calculation of read ratio.

num_simd_registers

  • Default: 30
  • Purpose: Number of SIMD registers to use when single_registers is false
  • Details: When running with single_registers set to false, we can tune how many registers we want to use so the different loads/stores use different registers for better pipeline utilization.

isa_mode

  • Default: ISAMode::AUTO
  • Purpose: Allows the user to customize the ISA the kernel is compiled for
  • Details: Can be set to match the target architecture (AVX, AVX-512, SVE, NEON, etc.). See ISA Mode Selection below for available options.

enable_interleaving

  • Default: true
  • Purpose: Controls interleaving of operations when using multiple registers
  • Details: When using multiple registers, if we should interleave to maximize parallelism or not between the registers used.

use_nontemporal_stores

  • Default: false
  • Purpose: Controls whether assembly uses non-temporal stores for store instructions
  • Details: If true, the assembly will use non-temporal stores. See Temporal vs Non-Temporal stores for details.

single_registers

  • Default: true
  • Purpose: Controls register usage strategy
  • Details: If true, it will use a single register for all loads, and a single register for all stores. If false, it will use however many registers specified in num_simd_registers.

ISA Mode Selection

ISA selection is done through the ISAMode enum in include/KernelTypes.h. Users can choose the ISA mode in the kernel configuration:

ISA Mode Architecture Description
AUTO Generic Automatic detection (recommended)
AVX512 x86_64 AVX-512 instruction set
AVX2 x86_64 AVX2 instruction set
AVX x86_64 AVX instruction set
SCALAR x86_64 Non-vectorized instructions
NEON_NATIVE ARM64 ISA-native NEON (ldr q / str q)
NEON_PAIR ARM64 Paired NEON (ldp q,q / stp q,q)
SVE ARM64 Use current thread VL with scalable SVE ops
SVE_MAX ARM64 Request maximum thread VL at benchmark runtime
SVE128 ARM64 Request 128-bit thread VL
SVE256 ARM64 Request 256-bit thread VL
SVE512 ARM64 Request 512-bit thread VL
RVV1_0 RISC-V RISC-V Vector 1.0
RVV0_7 RISC-V RISC-V Vector 0.7
VSX Power Vector Scalar Extensions
VMX Power Vector Multimedia Extensions

If the chosen ISA is not available on the system, it will fallback to a supported ISA and warn the user.

On ARM, AUTO currently resolves to NEON_PAIR.

For SVE-width modes, the generated assembly mnemonic family stays the same (ld1d, st1d, stnt1d). Width selection is done at benchmark runtime with prctl(PR_SVE_SET_VL, ...) on the benchmark thread.

addressing_policy

  • Default: AddressingPolicy::AUTO
  • Purpose: Controls how effective addresses are formed in the generated assembly
  • Values:
    • AUTO
    • POST_INCREMENT
    • INDEXED

POST_INCREMENT uses the current base pointer and then advances it. This is the natural streaming form for contiguous loops, and on ARM SVE the pointer motion is VL-aware through addvl.

INDEXED computes an explicit base-plus-offset address for each access. This is useful for controlled comparisons where the address calculation should be separated from post-increment pointer motion.

On ARM, AUTO resolves to POST_INCREMENT for the public traffic generator kernel.

store_line_policy

  • Default: StoreLinePolicy::AUTO
  • Purpose: Controls how many store instructions are emitted per logical write step
  • Values:
    • AUTO
    • ISA_NATIVE
    • CACHELINE_64B
    • CACHELINE_NATIVE

ISA_NATIVE emits one ISA-native store per logical write. CACHELINE_64B emits enough contiguous stores to cover 64 bytes. CACHELINE_NATIVE emits enough contiguous stores to cover the detected cache-line size of the machine.


Architecture-Specific Tuning

For detailed architecture-specific tuning guidance, including register considerations, alignment requirements, and platform-specific optimizations, see Architecture-Support.

The Architecture Support page covers:

  • x86-64 (AVX/AVX-512) tuning considerations
  • ARM (NEON/SVE) vector width and VL handling
  • Power (VSX/VMX) configuration
  • RISC-V (RVV) vector length (VLEN) handling
  • Platform-specific notes and recommendations

Recompilation After Changes

After modifying the kernel-generator configuration:

make clean
make
make install

Debugging Kernel Generation

In some cases, kernel compilation might fail. If compilation fails, you can use the following command to view the actual compilation output and debug why it failed:

make install DEBUG_FLAGS=--debug

This will show the detailed compilation process, including any error messages from the assembler or compiler, helping you identify and fix configuration issues.


See Also

Clone this wiki locally