Repository navigation
Traffic Generator Setup
The traffic generator kernel is configured through include/KernelTypes.h. These settings control how assembly code is generated for memory traffic operations.
For an explanation of what the Traffic Generator binary does, see Traffic Generator.
Here we explain how users can tune the kernel, since it's programmatically generated per-architecture.
Note: Modifying these settings requires recompilation. Most users should not need to change defaults.
The KernelConfigMultiSeq struct in include/KernelTypes.h controls the MultiSequential traffic generator behavior:
-
Default:
2 - Purpose: Controls the precision of issued load ratio in percentage steps
- Details: By default, we use steps of 2% for granularity of issued load ratio. This can be tuned for finer or coarser control over the load/store ratio.
-
Default:
6 - Purpose: Number of memory operations executed between pause points
- Details: The benchmark repeats the structure of "memops - nops - memops - nops". This flag customizes the amount of memory operations before a NOP block.
-
Default:
100 - Purpose: Total operations per kernel iteration
- Details: The kernel loops between these operations. We keep it at 100 for easy calculation of read ratio.
-
Default:
30 -
Purpose: Number of SIMD registers to use when
single_registersis false -
Details: When running with
single_registersset to false, we can tune how many registers we want to use so the different loads/stores use different registers for better pipeline utilization.
-
Default:
ISAMode::AUTO - Purpose: Allows the user to customize the ISA the kernel is compiled for
- Details: Can be set to match the target architecture (AVX, AVX-512, SVE, NEON, etc.). See ISA Mode Selection below for available options.
-
Default:
true - Purpose: Controls interleaving of operations when using multiple registers
- Details: When using multiple registers, if we should interleave to maximize parallelism or not between the registers used.
-
Default:
false - Purpose: Controls whether assembly uses non-temporal stores for store instructions
- Details: If true, the assembly will use non-temporal stores. See Temporal vs Non-Temporal stores for details.
-
Default:
true - Purpose: Controls register usage strategy
-
Details: If true, it will use a single register for all loads, and a single register for all stores. If false, it will use however many registers specified in
num_simd_registers.
ISA selection is done through the ISAMode enum in include/KernelTypes.h. Users can choose the ISA mode in the kernel configuration:
| ISA Mode | Architecture | Description |
|---|---|---|
AUTO |
Generic | Automatic detection (recommended) |
AVX512 |
x86_64 | AVX-512 instruction set |
AVX2 |
x86_64 | AVX2 instruction set |
AVX |
x86_64 | AVX instruction set |
SCALAR |
x86_64 | Non-vectorized instructions |
NEON_NATIVE |
ARM64 | ISA-native NEON (ldr q / str q) |
NEON_PAIR |
ARM64 | Paired NEON (ldp q,q / stp q,q) |
SVE |
ARM64 | Use current thread VL with scalable SVE ops |
SVE_MAX |
ARM64 | Request maximum thread VL at benchmark runtime |
SVE128 |
ARM64 | Request 128-bit thread VL |
SVE256 |
ARM64 | Request 256-bit thread VL |
SVE512 |
ARM64 | Request 512-bit thread VL |
RVV1_0 |
RISC-V | RISC-V Vector 1.0 |
RVV0_7 |
RISC-V | RISC-V Vector 0.7 |
VSX |
Power | Vector Scalar Extensions |
VMX |
Power | Vector Multimedia Extensions |
If the chosen ISA is not available on the system, it will fallback to a supported ISA and warn the user.
On ARM, AUTO currently resolves to NEON_PAIR.
For SVE-width modes, the generated assembly mnemonic family stays the same (ld1d, st1d, stnt1d). Width selection is done at benchmark runtime with prctl(PR_SVE_SET_VL, ...) on the benchmark thread.
-
Default:
AddressingPolicy::AUTO - Purpose: Controls how effective addresses are formed in the generated assembly
-
Values:
AUTOPOST_INCREMENTINDEXED
POST_INCREMENT uses the current base pointer and then advances it. This is the natural streaming form for contiguous loops, and on ARM SVE the pointer motion is VL-aware through addvl.
INDEXED computes an explicit base-plus-offset address for each access. This is useful for controlled comparisons where the address calculation should be separated from post-increment pointer motion.
On ARM, AUTO resolves to POST_INCREMENT for the public traffic generator kernel.
-
Default:
StoreLinePolicy::AUTO - Purpose: Controls how many store instructions are emitted per logical write step
-
Values:
AUTOISA_NATIVECACHELINE_64BCACHELINE_NATIVE
ISA_NATIVE emits one ISA-native store per logical write. CACHELINE_64B emits enough contiguous stores to cover 64 bytes. CACHELINE_NATIVE emits enough contiguous stores to cover the detected cache-line size of the machine.
For detailed architecture-specific tuning guidance, including register considerations, alignment requirements, and platform-specific optimizations, see Architecture-Support.
The Architecture Support page covers:
- x86-64 (AVX/AVX-512) tuning considerations
- ARM (NEON/SVE) vector width and VL handling
- Power (VSX/VMX) configuration
- RISC-V (RVV) vector length (VLEN) handling
- Platform-specific notes and recommendations
After modifying the kernel-generator configuration:
make clean
make
make installIn some cases, kernel compilation might fail. If compilation fails, you can use the following command to view the actual compilation output and debug why it failed:
make install DEBUG_FLAGS=--debugThis will show the detailed compilation process, including any error messages from the assembler or compiler, helping you identify and fix configuration issues.
- Traffic Generator - Overview of traffic generation
- Temporal vs Non-Temporal stores - Store type configuration
- Architecture Support - Supported platforms