Skip to content

GALAXY v0.4.0 — Deterministic Native CPU Runtime and Scaling Evidence

Choose a tag to compare

@EmergentMonk EmergentMonk released this 14 Sep 19:18
Immutable release. Only release title and notes can be modified.
6f17a73

RELEASE NOTES:

GALAXY v0.4.0 freezes the first formally reproducible native-CPU performance milestone for the project.

This release combines the deterministic galaxy instrument, native GPU runtime, exact-u64 logical addressing, native multicore CPU runtime, BAM32/Q2.30 lookup-table backend, local Ryzen validation, and independent qBraid/Azure EPYC scaling evidence into one archival release.

The emphasis of v0.4.0 is not simply higher benchmark numbers. It is reproducible execution: exact source identity, deterministic checksums, explicit runtime receipts, preserved hardware/topology evidence, and conservative boundaries around what the measurements do and do not establish.

Highlights

  • Native deterministic CPU runtime for GALAXY
  • Exact positive-u64 logical population support
  • Float and BAM-LUT CPU backends
  • Deterministic scalar/parallel checksum parity
  • Multicore scaling through 96 available logical CPUs
  • Ryzen 9 5950X local validation
  • AMD EPYC 7763 qBraid/Azure replication
  • 32 / 48 / 64 / 96 worker scaling study
  • NUMA and SMT topology capture
  • Controlled CPU-affinity experiments
  • Preserved raw benchmark receipts and evidence archive
  • Explicit affinity-command provenance
  • Existing Vulkan/CUDA GPU runtime retained
  • Reproducibility and scientific claim boundaries documented

Native CPU runtime

The native CPU runtime evolves independent resident test particles in GALAXY's fixed potential while preserving the project's deterministic logical-addressing contract.

The runtime supports:

  • exact positive-u64 logical populations up to:

    18,446,744,073,709,551,615

  • bounded resident populations up to:

    16,777,216

  • deterministic proportional logical-ID sampling

  • split-u64-hash32-avalanche-v1 identity addressing

  • Float/libm execution

  • BAM32/Q2.30 LUT execution

  • deterministic contiguous worker partitioning

  • scalar and parallel checksum comparison

  • fixed worker-count measurement

  • interleaved Float/BAM timing

  • machine-readable CPU runtime receipts

  • strict verification mode

The CPU implementation does not replace GALAXY's GPU runtime. It provides an independently measurable and reproducible native execution path.

BAM32 lookup-table backend

v0.4.0 includes the deterministic BAM-LUT CPU path derived from the isolated retro-math reference work.

The LUT path uses:

  • BAM32 angular representation
  • Q2.30 fixed-point lookup values
  • 16,384 LUT entries
  • deterministic conversion
  • sampled LUT diagnostic validation

The current diagnostic evaluates 8,193 angles and observed a sampled maximum absolute Q30 error of:

255

This value is a sampled implementation diagnostic, not a mathematical maximum over the full 2^32 BAM angle space.

Local Ryzen 9 5950X validation

The native CPU runtime was validated locally on:

AMD Ryzen 9 5950X
32 logical CPUs
Ubuntu Linux

The local experiment reached the full resident cap of:

16,777,216 particles

with deterministic scalar/parallel checksum parity.

A primary 8,388,608-resident / 32-worker result measured approximately:

Backend Scalar median Parallel median Measured speedup
Float 1.623 s 97.37 ms 16.67x
BAM-LUT 504.6 ms 84.56 ms 5.97x

The broader local matrix showed Float continuing to benefit through 32 workers while BAM-LUT plateaued substantially earlier at large resident sizes.

That observation motivated the wider qBraid experiment.

qBraid / Azure EPYC 7763 replication

The cloud replication was performed manually on a qBraid Nanoacademic Medium CPU instance.

Observed guest topology:

AMD EPYC 7763 64-Core Processor

with:

  • 96 online logical CPUs
  • 48 exposed cores
  • 2 threads per exposed core
  • 2 NUMA nodes
  • 192 MiB aggregate L3 reported by lscpu
  • approximately 377 GiB RAM
  • Microsoft full virtualization
  • Linux 6.8 Azure kernel

The guest explicitly exposed adjacent SMT sibling pairs:

0-1, 2-3, ... , 94-95

The exact merged PR #8 CPU runtime commit was verified before benchmarking.

All native CPU runtime tests passed, and verification confirmed:

  • available_parallelism=96
  • effective_workers=96
  • deterministic u64 boundary fingerprint
  • Float scalar/parallel checksum parity
  • BAM-LUT scalar/parallel checksum parity
  • sampled LUT diagnostic consistency

EPYC 8M worker-scaling result

For:

  • logical population = 18,446,744,073,709,551,615
  • resident particles = 8,388,608
  • frames = 8
  • repeats = 5
  • seed = 303

the unpinned worker sweep produced:

Workers Float parallel Float speedup BAM-LUT parallel BAM-LUT speedup
32 121.02 ms 20.30x 45.06 ms 16.65x
48 77.07 ms 31.71x 31.46 ms 24.10x
64 92.23 ms 26.65x 38.58 ms 19.53x
96 81.49 ms 29.92x 35.52 ms 21.22x

The best measured unpinned result for both backends occurred at 48 workers.

Performance was not monotonic beyond that point.

All worker counts preserved identical deterministic backend checksums.

Topology and affinity experiment

A repeat-matched 48-worker study was then used to test whether the 48-worker optimum corresponded simply to the guest's 48 exposed physical cores.

It did not.

Measured 48-worker results:

Configuration Float parallel BAM-LUT parallel
Unpinned 77.22 ms 31.38 ms
NUMA node 0 only 78.12 ms 33.01 ms
NUMA node 1 only 77.10 ms 34.33 ms
One SMT thread per exposed core across both NUMA nodes 115.54 ms 47.84 ms

The cross-NUMA one-thread-per-core configuration was roughly 50% slower than the unpinned result for both backends.

By contrast, restricting the process to one NUMA domain—24 exposed cores / 48 logical CPUs—retained performance close to the unrestricted baseline.

This result demonstrates strong topology sensitivity in this environment.

It is consistent with NUMA locality, first-touch placement, remote-memory traffic, cache effects, or memory-bandwidth interactions.

The available evidence does not isolate which of those mechanisms caused the difference because hardware memory-traffic counters were not collected.

Affinity provenance

The historical CPU runtime receipt schema records available and effective worker counts but does not encode Linux CPU affinity.

For the completed EPYC affinity experiment, v0.4.0 therefore preserves a separate affinity manifest binding each retained benchmark receipt to the exact taskset invocation used by the operator.

Future GALAXY affinity experiments are required to capture the kernel-visible allowed CPU list inside the constrained process before launching the benchmark.

This keeps the distinction clear between:

  • receipt-native evidence
  • operator command provenance
  • topology observations
  • inferred architectural interpretation

Evidence archive

The qBraid EPYC evidence is preserved under:

evidence/qbraid/EPYC7763-20260914/

including:

  • host topology capture
  • SMT sibling mapping
  • 32-worker receipt
  • 48-worker receipt
  • 64-worker receipt
  • 96-worker receipt
  • 7-repeat unpinned 48-worker confirmation
  • NUMA node 0 affinity receipt
  • NUMA node 1 affinity receipt
  • cross-NUMA one-thread-per-core receipt
  • affinity command manifest
  • evidence documentation

Original evidence archive SHA-256:

5c0474b9537a0ee34493a58c22b368f4846ad12f41633430d4444a678561e87a

GPU runtime

GALAXY continues to retain its native GPU runtime alongside the new CPU evidence.

Supported paths include:

  • Rust / wgpu / Vulkan
  • NVIDIA CUDA / CuPy RawKernel
  • memory-bounded u64 tiled execution
  • deterministic logical addressing beyond 2^32
  • hardware validation receipts
  • fixed-potential independent-particle dynamics

The CPU and GPU runtimes are separate implementations and should not be interpreted as direct end-to-end performance equivalents without a dedicated controlled comparison.

Scientific scope

GALAXY remains a deterministic galaxy dynamics and visualization instrument.

The native runtimes evolve independent test particles in prescribed gravitational potentials.

GALAXY does not currently implement:

  • pairwise stellar gravity
  • evolving self-gravity
  • hydrodynamics
  • gas evolution
  • star formation
  • self-consistent N-body dynamics
  • cross-particle force coupling

Large logical populations describe deterministic address spaces from which bounded resident populations are sampled or tiled.

A logical population of u64::MAX therefore does not mean that 18.4 quintillion particles are simultaneously allocated in memory.

Claim boundary

This release supports the following statement:

On the tested qBraid EPYC 7763 guest, GALAXY's 8,388,608-resident native CPU workload measured its best unpinned result at 48 workers in the tested 32/48/64/96-worker sweep for both Float and BAM-LUT, while preserving deterministic scalar/parallel checksums. A separate affinity experiment found a substantial penalty when one logical thread per exposed core was forced across both NUMA domains, while node-local 48-thread execution remained near the unpinned baseline.

This release does not claim that:

  • 96 cloud vCPUs are 96 physical cores
  • the unpinned 48-worker run occupied exactly 48 physical cores
  • BAM-LUT is proven memory-bandwidth bound
  • NUMA first-touch is proven to cause the observed affinity penalty
  • absolute EPYC and Ryzen performance are directly comparable
  • u64::MAX particles are simultaneously resident
  • GALAXY is a self-consistent N-body solver
  • the CPU runtime universally outperforms the GPU runtime
  • the sampled LUT diagnostic is a global mathematical error bound

Reproducibility milestone

v0.4.0 is intended to serve as the frozen baseline for the next CPU-runtime architecture phase.

Future work can now compare against a stable reference before exploring:

  • persistent worker pools
  • reduced thread-creation overhead
  • topology-aware scheduling
  • NUMA-aware placement
  • explicit affinity policies
  • stronger receipt-native topology provenance

Those changes belong to a later release so that the performance effect of the new architecture can be measured against this frozen baseline.

Release identity

Version:

v0.4.0

Release commit:

6f17a734b9241359d36a9bf3d208b8527a456327

Repository:

QSOLKCB/GALAXY

This release is intended as the source snapshot for the corresponding Zenodo software record.