Repository navigation
September 29th 2026
Added
- Consumer Platform Support: Expanded Mess to consumer laptops, desktops and workstations, alongside server systems:
- Apple Silicon Macs: Native macOS bandwidth measurement through IOReport, without an additional counter driver.
- Windows x86-64: Intel VTune, Intel client IMC and AMD uProf bandwidth backends, with vendor tools and Administrator access.
- Linux PCs: Improved backend discovery and fallback for supported Intel and AMD systems with accessible bandwidth counters.
- Mess GUI: Desktop interface for configuring runs, launching bandwidth-latency curves and inspecting results. Installers for Linux, macOS (Apple Silicon) and Windows will be available on the Mess GUI download page.
- NVIDIA GPU Benchmark: Generate GPU bandwidth-latency curves with
--gpu, configurable read/write ratios and adaptive pause discovery. Requires the NVIDIA driver and CUDA Toolkit 12.6+ withnvccand CUPTI; GPU kernels are compiled at runtime and DRAM counters are collected through CUPTI. - Automatic Traffic-Array Calibration: Select traffic-generator array sizes during kernel generation using measured bandwidth, last-level cache capacity and memory limits, with a fallback when calibration cannot complete.
- Mess Score: Plotter summaries now include a score combining bandwidth efficiency and latency growth under load, normalized against theoretical bandwidth and baseline latency.
Improved
- Profiling by Default: CPU and GPU runs now save measurement files automatically. Use
--no-profileto discard them after the run; this replaces--profile. - Instruction-Latency Sampling: Expanded the optional
--inst-latworkflow for measuring latency directly from traffic-generator loads using Intel PEBS or ARM SPE on supported Linux systems, with latency statistics and percentile plots. Improved memory-source filtering for HBM and CXL measurements. - ISA Tuning and Assembly Generation:
- Expanded scalar and vector kernel choices, including SSE2, AVX, AVX2 and AVX-512 on x86, and native/paired NEON modes on ARM.
- Improved SVE handling with runtime vector-length selection, maximum-length and fixed-width modes, and vector-length-aware pointer updates.
- Extended
AddressingPolicyandStoreLinePolicytuning for address generation and ISA-native or cache-line-sized store operations.
- Hardware Detection: Added AMD Zen 5 bandwidth-counter discovery and improved CPU identification, last-level cache capacity detection across cache slices, and Apple Silicon memory-configuration detection.
- Bandwidth Backends:
- Intel VTune: Improved discovery, counter aggregation, report parsing and collection fallback; available on Linux and supported Windows systems.
- Intel PCM / CXL: Improved detection of CXL memory-node binding and automatic PCM selection on Intel Linux systems. CXL bandwidth measurement uses the PCM backend and requires supported counters.
- Adaptive Pause Discovery and Run Tiers: Refined curve-guided sampling around transitions, with validation of suspicious bends. Select
--tier=lite,standardordetailedfor point budgets of 15, 50 or 200, or use--point-count=Nfor a custom budget. - Reworked Plotter: Improved curve processing, smoothing and comparison across runs, with CSV/JSON data exports, PDF/PNG figures and instruction-sampling percentile plots.
- Measurement Stability: Stabilization now uses theoretical and achievable bandwidth estimates, accounting for worker count and memory placement, to scale warmup and noise thresholds. Improved low-throughput handling and traffic-generator health checks.
- Code Maintainability and Portability: Improved separation of platform services, measurement backends and parsers; more reliable process cleanup, bounded command collection, build resource discovery and error reporting.
Fixed
- Latency Timing: Improved CPU-frequency calibration and consistent latency conversion between runtime measurements and the plotter, including elapsed-time fallback when cycle counters are unavailable.
- Perf Interval Accounting: Aligned interval counter values with their timestamps to avoid incorrect bandwidth normalization.
- Curve Processing: Preserved curves below 50% read ratio, improved latency-outlier and bandwidth-zigzag handling, and restricted zero-bandwidth anchors to stable near-idle data.
- Build Paths: Improved handling of paths containing spaces and generated resources when running from a different working directory.
Compatibility
- CPU bandwidth measurement depends on counters supported by the processor, operating system and selected backend.
- macOS bandwidth measurement requires Apple Silicon.
--bind,--inst-latand--add-countersare unavailable on macOS and Windows. - Windows backends are implemented; native Windows driver and benchmark-curve validation remains pending.
Mess GUI downloads
GUI version 1.0.0-beta · source commit 00bb6647c125bc7b4bc8b3379a2d8f80703784d5