Skip to content

Architecture Support

Victor Xirau Guardans edited this page Sep 29, 2026 · 6 revisions

This page documents supported architectures and provides a developer guide for adding support for new systems.


Supported Architectures

Mess automatically detects and supports the following architectures:

Architecture Status SIMD Notes
x86-64 CPUs Supported AVX2, AVX-512 Intel and AMD processors
ARM CPUs Supported NEON, SVE Includes Neoverse, Graviton, Apple Silicon
Power CPUs Supported VSX Power8 and newer
RISC-V CPUs WIP RVV 1.0 Assembly generation, latency measurement and counter detection supported; bandwidth measurement pending (requires approximation via alternative counters)
GPUs Pending - Under active development

Note: Adding a new system should not be necessary in most cases, since Mess automatically detects every configuration. From time to time, we encounter exotic systems that weren't accounted for. In that case, use this guide.


Auto-Detection

Mess automatically detects your architecture at compile time and selects appropriate:

  • SIMD instruction sets
  • Register configurations
  • Memory alignment requirements
  • Performance counter backends

To see what was detected:

./build/bin/mess --dry-run --verbose=2

Platform-Specific Notes

x86-64

  • AVX2: Used automatically on x86-64 processors unless configured otherwise
  • AVX-512: Available for manual selection
  • Automatic detection via CPUID

ARM

  • NEON_NATIVE: 128-bit ISA-native NEON traffic using ldr q / str q
  • NEON_PAIR: paired-register NEON traffic using ldp q,q / stp q,q
  • AUTO: currently resolves to NEON_PAIR
  • SVE: scalable vector support with runtime VL handling
  • SVE_MAX: requests the maximum thread VL at benchmark runtime
  • SVE128 / SVE256 / SVE512: fixed-width SVE benchmark modes
  • SVE kernels use ptrue p0.d together with ld1d, st1d, and stnt1d
  • SVE sequential post-increment addressing is VL-aware through addvl

Power

  • VSX: Required (Power8+)
  • Endianness: Little-endian supported

RISC-V (WIP)

RISC-V support is a work in progress. The following features are available:

  • Assembly generation: RVV 1.0 and RVV 0.7 vector kernels
  • Latency measurement: Pointer chase based latency via perf
  • Counter detection: SiFive-specific counter strategy

Not yet available:

  • Bandwidth measurement: RISC-V platforms we've tested (e.g., SiFive) do not expose uncore memory controller counters (CAS counters) via perf. Bandwidth measurement requires approximation using alternative counter sets (e.g., L1/L2 cache events), which is not yet implemented. This was available in original Mess 1.0 via raw perf events and self-timed stream output.

Adding Support for New Architectures

This section provides a developer guide for adding support for new CPU architectures or systems.

Plugin Architecture Overview

Mess uses a plugin-based architecture system. To add a new system, you need to implement three main interfaces:

Interface Purpose
Architecture Main entry point and factory class
KernelAssembler Handles assembly code generation
BandwidthCounterStrategy Handles hardware performance counter discovery
┌──────────────────────────────────────────────────────────────┐
│                    ArchitectureRegistry                      │
│                           ↓                                  │
│  ┌──────────────────────────────────────────────────────┐    │
│  │               Architecture (interface)               │    │
│  │  - getName()                                         │    │
│  │  - supports(CPUCapabilities)                         │    │
│  │  - createAssembler(KernelConfig)                     │    │
│  │  - createCounterStrategy(CPUCapabilities)            │    │
│  │  - getSupportedISAs()                                │    │
│  │  - selectBestISA(CPUCapabilities)                    │    │
│  │  - getUpiScalingFactor(CPUCapabilities)              │    │
│  └──────────────────────────────────────────────────────┘    │
│              ↓                              ↓                │
│  ┌──────────────────────┐    ┌────────────────────────────┐  │
│  │   KernelAssembler    │    │ BandwidthCounterStrategy   │  │
│  │  - generateLoad()    │    │  - detectCasCounters()     │  │
│  │  - generateStore()   │    │  - getTlbMissCounters()    │  │
│  │  - generateNop()     │    │                            │  │
│  └──────────────────────┘    └────────────────────────────┘  │
└──────────────────────────────────────────────────────────────┘

Step 1: Create Architecture Files

Create new directories for your architecture:

src/arch/<arch_name>/
include/arch/<arch_name>/

Header File: include/arch/<arch_name>/<ArchName>Architecture.h

#include "architecture/Architecture.h"

class MyArch : public Architecture {
public:
    std::string getName() const override { return "myarch"; }
    bool supports(const CPUCapabilities& caps) const override;

    std::unique_ptr<KernelAssembler> createAssembler(
        const KernelConfig& config) const override;
    std::unique_ptr<BandwidthCounterStrategy> createCounterStrategy(
        const CPUCapabilities& caps) const override;

    std::vector<std::shared_ptr<ISA>> getSupportedISAs() const override;
    std::shared_ptr<ISA> selectBestISA(
        const CPUCapabilities& caps) const override;

    double getUpiScalingFactor(const CPUCapabilities& caps) const override;
};

Header File: include/arch/<arch_name>/<ArchName>Assembler.h

#include "architecture/KernelAssembler.h"

class MyAssembler : public KernelAssembler {
public:
    MyAssembler(const KernelConfig& config);

    // Implement pure virtual methods
    std::string generateLoad(int reg, int offset) override;
    std::string generateStore(int reg, int offset) override;
    std::string generateNop() override;
    // ... other methods
};

Step 2: Implement the Classes

Implement your architecture classes in src/arch/<arch_name>/.

Key implementation details:

Component Requirements
Assembler Must generate valid assembly for the target architecture. Use the provided config to handle unrolling and register usage.
Counters Implement detectCasCounters and getTlbMissCounters to map abstract events to hardware-specific raw event codes.
ISA Define supported ISAs (e.g., AVX, NEON) and their vector widths.
Scaling Implement getUpiScalingFactor to convert interconnect units to cache-line-equivalent traffic (usually 1.0 for local memory).

Step 3: Register the Architecture

In your <ArchName>Architecture.cpp, add the registration macro at the end:

#include "architecture/ArchitectureRegistry.h"

// ... implementation ...

static ArchitectureRegistrar<MyArch> my_arch_registrar;

This automatically registers your architecture with the system at startup.

Step 4: Update Build System

Add your new source files to Makefile:

ARCH_SRCS += src/arch/<arch_name>/<ArchName>Architecture.cpp \
             src/arch/<arch_name>/<ArchName>Assembler.cpp

Find the ARCH_SRCS variable definition in the main Makefile and append your files there.

Step 5: Verify

# Build
make clean && make

# Check detection
./build/bin/mess --dry-run --verbose=3

# Test single point
./build/bin/mess --ratio=100 --pause=0 --verbose=3

Adding a Microarchitecture Variant

To add support for a specific microarchitecture variant (e.g., Zen4 for x86, or Neoverse for ARM):

1. Modify the Architecture Class

In your createAssembler or createCounterStrategy methods, check caps.model_name or caps.uarch to return a specialized subclass:

std::unique_ptr<KernelAssembler> X86Architecture::createAssembler(
    const KernelConfig& config) const
{
    if (config.uarch == "zen4") {
        return std::make_unique<Zen4Assembler>(config);
    }
    return std::make_unique<X86Assembler>(config);
}

2. Implement Specialized Classes

Inherit from the base architecture assembler/counters and override only what's needed:

class Zen4Assembler : public X86Assembler {
public:
    Zen4Assembler(const KernelConfig& config) : X86Assembler(config) {}

    // Override specific instruction generation if needed
    std::string generateLoad(int reg, int offset) override {
        // Zen4-specific optimization
    }
};

Testing New Architectures

Build Test

make clean && make

Dry-run Validation

./build/bin/mess --dry-run --verbose=3

Check that:

  • Architecture is correctly detected
  • ISA selection is appropriate
  • Performance counters are found

Single-Point Test

./build/bin/mess --ratio=100 --pause=0 --verbose=3

Verify:

  • Non-zero bandwidth measurement
  • Reasonable latency (~80-200ns unloaded)
  • No errors or warnings

Full Validation

./build/bin/mess

Check that:

  • All ratios produce valid bandwidth-latency curves
  • Bandwidth decreases as pause increases
  • Latency increases as bandwidth increases

Contributing

To contribute architecture support:

  1. Fork the repository
  2. Create a feature branch
  3. Implement following the steps above
  4. Test on the target architecture
  5. Submit a merge request

Contact mess@bsc.es for guidance.


See Also

Clone this wiki locally