Repository navigation
Architecture Support
This page documents supported architectures and provides a developer guide for adding support for new systems.
Mess automatically detects and supports the following architectures:
| Architecture | Status | SIMD | Notes |
|---|---|---|---|
| x86-64 CPUs | Supported | AVX2, AVX-512 | Intel and AMD processors |
| ARM CPUs | Supported | NEON, SVE | Includes Neoverse, Graviton, Apple Silicon |
| Power CPUs | Supported | VSX | Power8 and newer |
| RISC-V CPUs | WIP | RVV 1.0 | Assembly generation, latency measurement and counter detection supported; bandwidth measurement pending (requires approximation via alternative counters) |
| GPUs | Pending | - | Under active development |
Note: Adding a new system should not be necessary in most cases, since Mess automatically detects every configuration. From time to time, we encounter exotic systems that weren't accounted for. In that case, use this guide.
Mess automatically detects your architecture at compile time and selects appropriate:
- SIMD instruction sets
- Register configurations
- Memory alignment requirements
- Performance counter backends
To see what was detected:
./build/bin/mess --dry-run --verbose=2- AVX2: Used automatically on x86-64 processors unless configured otherwise
- AVX-512: Available for manual selection
- Automatic detection via CPUID
-
NEON_NATIVE: 128-bit ISA-native NEON traffic using
ldr q/str q -
NEON_PAIR: paired-register NEON traffic using
ldp q,q/stp q,q -
AUTO: currently resolves to
NEON_PAIR - SVE: scalable vector support with runtime VL handling
- SVE_MAX: requests the maximum thread VL at benchmark runtime
- SVE128 / SVE256 / SVE512: fixed-width SVE benchmark modes
- SVE kernels use
ptrue p0.dtogether withld1d,st1d, andstnt1d - SVE sequential post-increment addressing is VL-aware through
addvl
- VSX: Required (Power8+)
- Endianness: Little-endian supported
RISC-V support is a work in progress. The following features are available:
- Assembly generation: RVV 1.0 and RVV 0.7 vector kernels
- Latency measurement: Pointer chase based latency via perf
- Counter detection: SiFive-specific counter strategy
Not yet available:
- Bandwidth measurement: RISC-V platforms we've tested (e.g., SiFive) do not expose uncore memory controller counters (CAS counters) via perf. Bandwidth measurement requires approximation using alternative counter sets (e.g., L1/L2 cache events), which is not yet implemented. This was available in original Mess 1.0 via raw perf events and self-timed stream output.
This section provides a developer guide for adding support for new CPU architectures or systems.
Mess uses a plugin-based architecture system. To add a new system, you need to implement three main interfaces:
| Interface | Purpose |
|---|---|
| Architecture | Main entry point and factory class |
| KernelAssembler | Handles assembly code generation |
| BandwidthCounterStrategy | Handles hardware performance counter discovery |
┌──────────────────────────────────────────────────────────────┐
│ ArchitectureRegistry │
│ ↓ │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ Architecture (interface) │ │
│ │ - getName() │ │
│ │ - supports(CPUCapabilities) │ │
│ │ - createAssembler(KernelConfig) │ │
│ │ - createCounterStrategy(CPUCapabilities) │ │
│ │ - getSupportedISAs() │ │
│ │ - selectBestISA(CPUCapabilities) │ │
│ │ - getUpiScalingFactor(CPUCapabilities) │ │
│ └──────────────────────────────────────────────────────┘ │
│ ↓ ↓ │
│ ┌──────────────────────┐ ┌────────────────────────────┐ │
│ │ KernelAssembler │ │ BandwidthCounterStrategy │ │
│ │ - generateLoad() │ │ - detectCasCounters() │ │
│ │ - generateStore() │ │ - getTlbMissCounters() │ │
│ │ - generateNop() │ │ │ │
│ └──────────────────────┘ └────────────────────────────┘ │
└──────────────────────────────────────────────────────────────┘
Create new directories for your architecture:
src/arch/<arch_name>/
include/arch/<arch_name>/
#include "architecture/Architecture.h"
class MyArch : public Architecture {
public:
std::string getName() const override { return "myarch"; }
bool supports(const CPUCapabilities& caps) const override;
std::unique_ptr<KernelAssembler> createAssembler(
const KernelConfig& config) const override;
std::unique_ptr<BandwidthCounterStrategy> createCounterStrategy(
const CPUCapabilities& caps) const override;
std::vector<std::shared_ptr<ISA>> getSupportedISAs() const override;
std::shared_ptr<ISA> selectBestISA(
const CPUCapabilities& caps) const override;
double getUpiScalingFactor(const CPUCapabilities& caps) const override;
};#include "architecture/KernelAssembler.h"
class MyAssembler : public KernelAssembler {
public:
MyAssembler(const KernelConfig& config);
// Implement pure virtual methods
std::string generateLoad(int reg, int offset) override;
std::string generateStore(int reg, int offset) override;
std::string generateNop() override;
// ... other methods
};Implement your architecture classes in src/arch/<arch_name>/.
Key implementation details:
| Component | Requirements |
|---|---|
| Assembler | Must generate valid assembly for the target architecture. Use the provided config to handle unrolling and register usage. |
| Counters | Implement detectCasCounters and getTlbMissCounters to map abstract events to hardware-specific raw event codes. |
| ISA | Define supported ISAs (e.g., AVX, NEON) and their vector widths. |
| Scaling | Implement getUpiScalingFactor to convert interconnect units to cache-line-equivalent traffic (usually 1.0 for local memory). |
In your <ArchName>Architecture.cpp, add the registration macro at the end:
#include "architecture/ArchitectureRegistry.h"
// ... implementation ...
static ArchitectureRegistrar<MyArch> my_arch_registrar;This automatically registers your architecture with the system at startup.
Add your new source files to Makefile:
ARCH_SRCS += src/arch/<arch_name>/<ArchName>Architecture.cpp \
src/arch/<arch_name>/<ArchName>Assembler.cppFind the ARCH_SRCS variable definition in the main Makefile and append your files there.
# Build
make clean && make
# Check detection
./build/bin/mess --dry-run --verbose=3
# Test single point
./build/bin/mess --ratio=100 --pause=0 --verbose=3To add support for a specific microarchitecture variant (e.g., Zen4 for x86, or Neoverse for ARM):
In your createAssembler or createCounterStrategy methods, check caps.model_name or caps.uarch to return a specialized subclass:
std::unique_ptr<KernelAssembler> X86Architecture::createAssembler(
const KernelConfig& config) const
{
if (config.uarch == "zen4") {
return std::make_unique<Zen4Assembler>(config);
}
return std::make_unique<X86Assembler>(config);
}Inherit from the base architecture assembler/counters and override only what's needed:
class Zen4Assembler : public X86Assembler {
public:
Zen4Assembler(const KernelConfig& config) : X86Assembler(config) {}
// Override specific instruction generation if needed
std::string generateLoad(int reg, int offset) override {
// Zen4-specific optimization
}
};make clean && make./build/bin/mess --dry-run --verbose=3Check that:
- Architecture is correctly detected
- ISA selection is appropriate
- Performance counters are found
./build/bin/mess --ratio=100 --pause=0 --verbose=3Verify:
- Non-zero bandwidth measurement
- Reasonable latency (~80-200ns unloaded)
- No errors or warnings
./build/bin/messCheck that:
- All ratios produce valid bandwidth-latency curves
- Bandwidth decreases as pause increases
- Latency increases as bandwidth increases
To contribute architecture support:
- Fork the repository
- Create a feature branch
- Implement following the steps above
- Test on the target architecture
- Submit a merge request
Contact mess@bsc.es for guidance.
- Traffic Generator - How traffic generation works
- Traffic generator setup - Kernel configuration
- FAQ - Common build issues