Skip to content

v2.1.0

Latest

Choose a tag to compare

@bsc-mem bsc-mem released this 29 Sep 13:23

September 29th 2026

Added

  • Consumer Platform Support: Expanded Mess to consumer laptops, desktops and workstations, alongside server systems:
    • Apple Silicon Macs: Native macOS bandwidth measurement through IOReport, without an additional counter driver.
    • Windows x86-64: Intel VTune, Intel client IMC and AMD uProf bandwidth backends, with vendor tools and Administrator access.
    • Linux PCs: Improved backend discovery and fallback for supported Intel and AMD systems with accessible bandwidth counters.
  • Mess GUI: Desktop interface for configuring runs, launching bandwidth-latency curves and inspecting results. Installers for Linux, macOS (Apple Silicon) and Windows will be available on the Mess GUI download page.
  • NVIDIA GPU Benchmark: Generate GPU bandwidth-latency curves with --gpu, configurable read/write ratios and adaptive pause discovery. Requires the NVIDIA driver and CUDA Toolkit 12.6+ with nvcc and CUPTI; GPU kernels are compiled at runtime and DRAM counters are collected through CUPTI.
  • Automatic Traffic-Array Calibration: Select traffic-generator array sizes during kernel generation using measured bandwidth, last-level cache capacity and memory limits, with a fallback when calibration cannot complete.
  • Mess Score: Plotter summaries now include a score combining bandwidth efficiency and latency growth under load, normalized against theoretical bandwidth and baseline latency.

Improved

  • Profiling by Default: CPU and GPU runs now save measurement files automatically. Use --no-profile to discard them after the run; this replaces --profile.
  • Instruction-Latency Sampling: Expanded the optional --inst-lat workflow for measuring latency directly from traffic-generator loads using Intel PEBS or ARM SPE on supported Linux systems, with latency statistics and percentile plots. Improved memory-source filtering for HBM and CXL measurements.
  • ISA Tuning and Assembly Generation:
    • Expanded scalar and vector kernel choices, including SSE2, AVX, AVX2 and AVX-512 on x86, and native/paired NEON modes on ARM.
    • Improved SVE handling with runtime vector-length selection, maximum-length and fixed-width modes, and vector-length-aware pointer updates.
    • Extended AddressingPolicy and StoreLinePolicy tuning for address generation and ISA-native or cache-line-sized store operations.
  • Hardware Detection: Added AMD Zen 5 bandwidth-counter discovery and improved CPU identification, last-level cache capacity detection across cache slices, and Apple Silicon memory-configuration detection.
  • Bandwidth Backends:
    • Intel VTune: Improved discovery, counter aggregation, report parsing and collection fallback; available on Linux and supported Windows systems.
    • Intel PCM / CXL: Improved detection of CXL memory-node binding and automatic PCM selection on Intel Linux systems. CXL bandwidth measurement uses the PCM backend and requires supported counters.
  • Adaptive Pause Discovery and Run Tiers: Refined curve-guided sampling around transitions, with validation of suspicious bends. Select --tier=lite, standard or detailed for point budgets of 15, 50 or 200, or use --point-count=N for a custom budget.
  • Reworked Plotter: Improved curve processing, smoothing and comparison across runs, with CSV/JSON data exports, PDF/PNG figures and instruction-sampling percentile plots.
  • Measurement Stability: Stabilization now uses theoretical and achievable bandwidth estimates, accounting for worker count and memory placement, to scale warmup and noise thresholds. Improved low-throughput handling and traffic-generator health checks.
  • Code Maintainability and Portability: Improved separation of platform services, measurement backends and parsers; more reliable process cleanup, bounded command collection, build resource discovery and error reporting.

Fixed

  • Latency Timing: Improved CPU-frequency calibration and consistent latency conversion between runtime measurements and the plotter, including elapsed-time fallback when cycle counters are unavailable.
  • Perf Interval Accounting: Aligned interval counter values with their timestamps to avoid incorrect bandwidth normalization.
  • Curve Processing: Preserved curves below 50% read ratio, improved latency-outlier and bandwidth-zigzag handling, and restricted zero-bandwidth anchors to stable near-idle data.
  • Build Paths: Improved handling of paths containing spaces and generated resources when running from a different working directory.

Compatibility

  • CPU bandwidth measurement depends on counters supported by the processor, operating system and selected backend.
  • macOS bandwidth measurement requires Apple Silicon. --bind, --inst-lat and --add-counters are unavailable on macOS and Windows.
  • Windows backends are implemented; native Windows driver and benchmark-curve validation remains pending.

Mess GUI downloads

GUI version 1.0.0-beta · source commit 00bb6647c125bc7b4bc8b3379a2d8f80703784d5