GALAXY v0.4.0 — Deterministic Native CPU Runtime and Scaling Evidence
RELEASE NOTES:
GALAXY v0.4.0 freezes the first formally reproducible native-CPU performance milestone for the project.
This release combines the deterministic galaxy instrument, native GPU runtime, exact-u64 logical addressing, native multicore CPU runtime, BAM32/Q2.30 lookup-table backend, local Ryzen validation, and independent qBraid/Azure EPYC scaling evidence into one archival release.
The emphasis of v0.4.0 is not simply higher benchmark numbers. It is reproducible execution: exact source identity, deterministic checksums, explicit runtime receipts, preserved hardware/topology evidence, and conservative boundaries around what the measurements do and do not establish.
Highlights
- Native deterministic CPU runtime for GALAXY
- Exact positive-u64 logical population support
- Float and BAM-LUT CPU backends
- Deterministic scalar/parallel checksum parity
- Multicore scaling through 96 available logical CPUs
- Ryzen 9 5950X local validation
- AMD EPYC 7763 qBraid/Azure replication
- 32 / 48 / 64 / 96 worker scaling study
- NUMA and SMT topology capture
- Controlled CPU-affinity experiments
- Preserved raw benchmark receipts and evidence archive
- Explicit affinity-command provenance
- Existing Vulkan/CUDA GPU runtime retained
- Reproducibility and scientific claim boundaries documented
Native CPU runtime
The native CPU runtime evolves independent resident test particles in GALAXY's fixed potential while preserving the project's deterministic logical-addressing contract.
The runtime supports:
-
exact positive-u64 logical populations up to:
18,446,744,073,709,551,615 -
bounded resident populations up to:
16,777,216 -
deterministic proportional logical-ID sampling
-
split-u64-hash32-avalanche-v1identity addressing -
Float/libm execution
-
BAM32/Q2.30 LUT execution
-
deterministic contiguous worker partitioning
-
scalar and parallel checksum comparison
-
fixed worker-count measurement
-
interleaved Float/BAM timing
-
machine-readable CPU runtime receipts
-
strict verification mode
The CPU implementation does not replace GALAXY's GPU runtime. It provides an independently measurable and reproducible native execution path.
BAM32 lookup-table backend
v0.4.0 includes the deterministic BAM-LUT CPU path derived from the isolated retro-math reference work.
The LUT path uses:
- BAM32 angular representation
- Q2.30 fixed-point lookup values
- 16,384 LUT entries
- deterministic conversion
- sampled LUT diagnostic validation
The current diagnostic evaluates 8,193 angles and observed a sampled maximum absolute Q30 error of:
255
This value is a sampled implementation diagnostic, not a mathematical maximum over the full 2^32 BAM angle space.
Local Ryzen 9 5950X validation
The native CPU runtime was validated locally on:
AMD Ryzen 9 5950X
32 logical CPUs
Ubuntu Linux
The local experiment reached the full resident cap of:
16,777,216 particles
with deterministic scalar/parallel checksum parity.
A primary 8,388,608-resident / 32-worker result measured approximately:
| Backend | Scalar median | Parallel median | Measured speedup |
|---|---|---|---|
| Float | 1.623 s | 97.37 ms | 16.67x |
| BAM-LUT | 504.6 ms | 84.56 ms | 5.97x |
The broader local matrix showed Float continuing to benefit through 32 workers while BAM-LUT plateaued substantially earlier at large resident sizes.
That observation motivated the wider qBraid experiment.
qBraid / Azure EPYC 7763 replication
The cloud replication was performed manually on a qBraid Nanoacademic Medium CPU instance.
Observed guest topology:
AMD EPYC 7763 64-Core Processor
with:
- 96 online logical CPUs
- 48 exposed cores
- 2 threads per exposed core
- 2 NUMA nodes
- 192 MiB aggregate L3 reported by
lscpu - approximately 377 GiB RAM
- Microsoft full virtualization
- Linux 6.8 Azure kernel
The guest explicitly exposed adjacent SMT sibling pairs:
0-1, 2-3, ... , 94-95
The exact merged PR #8 CPU runtime commit was verified before benchmarking.
All native CPU runtime tests passed, and verification confirmed:
available_parallelism=96effective_workers=96- deterministic u64 boundary fingerprint
- Float scalar/parallel checksum parity
- BAM-LUT scalar/parallel checksum parity
- sampled LUT diagnostic consistency
EPYC 8M worker-scaling result
For:
- logical population =
18,446,744,073,709,551,615 - resident particles =
8,388,608 - frames =
8 - repeats =
5 - seed =
303
the unpinned worker sweep produced:
| Workers | Float parallel | Float speedup | BAM-LUT parallel | BAM-LUT speedup |
|---|---|---|---|---|
| 32 | 121.02 ms | 20.30x | 45.06 ms | 16.65x |
| 48 | 77.07 ms | 31.71x | 31.46 ms | 24.10x |
| 64 | 92.23 ms | 26.65x | 38.58 ms | 19.53x |
| 96 | 81.49 ms | 29.92x | 35.52 ms | 21.22x |
The best measured unpinned result for both backends occurred at 48 workers.
Performance was not monotonic beyond that point.
All worker counts preserved identical deterministic backend checksums.
Topology and affinity experiment
A repeat-matched 48-worker study was then used to test whether the 48-worker optimum corresponded simply to the guest's 48 exposed physical cores.
It did not.
Measured 48-worker results:
| Configuration | Float parallel | BAM-LUT parallel |
|---|---|---|
| Unpinned | 77.22 ms | 31.38 ms |
| NUMA node 0 only | 78.12 ms | 33.01 ms |
| NUMA node 1 only | 77.10 ms | 34.33 ms |
| One SMT thread per exposed core across both NUMA nodes | 115.54 ms | 47.84 ms |
The cross-NUMA one-thread-per-core configuration was roughly 50% slower than the unpinned result for both backends.
By contrast, restricting the process to one NUMA domain—24 exposed cores / 48 logical CPUs—retained performance close to the unrestricted baseline.
This result demonstrates strong topology sensitivity in this environment.
It is consistent with NUMA locality, first-touch placement, remote-memory traffic, cache effects, or memory-bandwidth interactions.
The available evidence does not isolate which of those mechanisms caused the difference because hardware memory-traffic counters were not collected.
Affinity provenance
The historical CPU runtime receipt schema records available and effective worker counts but does not encode Linux CPU affinity.
For the completed EPYC affinity experiment, v0.4.0 therefore preserves a separate affinity manifest binding each retained benchmark receipt to the exact taskset invocation used by the operator.
Future GALAXY affinity experiments are required to capture the kernel-visible allowed CPU list inside the constrained process before launching the benchmark.
This keeps the distinction clear between:
- receipt-native evidence
- operator command provenance
- topology observations
- inferred architectural interpretation
Evidence archive
The qBraid EPYC evidence is preserved under:
evidence/qbraid/EPYC7763-20260914/
including:
- host topology capture
- SMT sibling mapping
- 32-worker receipt
- 48-worker receipt
- 64-worker receipt
- 96-worker receipt
- 7-repeat unpinned 48-worker confirmation
- NUMA node 0 affinity receipt
- NUMA node 1 affinity receipt
- cross-NUMA one-thread-per-core receipt
- affinity command manifest
- evidence documentation
Original evidence archive SHA-256:
5c0474b9537a0ee34493a58c22b368f4846ad12f41633430d4444a678561e87a
GPU runtime
GALAXY continues to retain its native GPU runtime alongside the new CPU evidence.
Supported paths include:
- Rust /
wgpu/ Vulkan - NVIDIA CUDA / CuPy RawKernel
- memory-bounded u64 tiled execution
- deterministic logical addressing beyond 2^32
- hardware validation receipts
- fixed-potential independent-particle dynamics
The CPU and GPU runtimes are separate implementations and should not be interpreted as direct end-to-end performance equivalents without a dedicated controlled comparison.
Scientific scope
GALAXY remains a deterministic galaxy dynamics and visualization instrument.
The native runtimes evolve independent test particles in prescribed gravitational potentials.
GALAXY does not currently implement:
- pairwise stellar gravity
- evolving self-gravity
- hydrodynamics
- gas evolution
- star formation
- self-consistent N-body dynamics
- cross-particle force coupling
Large logical populations describe deterministic address spaces from which bounded resident populations are sampled or tiled.
A logical population of u64::MAX therefore does not mean that 18.4 quintillion particles are simultaneously allocated in memory.
Claim boundary
This release supports the following statement:
On the tested qBraid EPYC 7763 guest, GALAXY's 8,388,608-resident native CPU workload measured its best unpinned result at 48 workers in the tested 32/48/64/96-worker sweep for both Float and BAM-LUT, while preserving deterministic scalar/parallel checksums. A separate affinity experiment found a substantial penalty when one logical thread per exposed core was forced across both NUMA domains, while node-local 48-thread execution remained near the unpinned baseline.
This release does not claim that:
- 96 cloud vCPUs are 96 physical cores
- the unpinned 48-worker run occupied exactly 48 physical cores
- BAM-LUT is proven memory-bandwidth bound
- NUMA first-touch is proven to cause the observed affinity penalty
- absolute EPYC and Ryzen performance are directly comparable
u64::MAXparticles are simultaneously resident- GALAXY is a self-consistent N-body solver
- the CPU runtime universally outperforms the GPU runtime
- the sampled LUT diagnostic is a global mathematical error bound
Reproducibility milestone
v0.4.0 is intended to serve as the frozen baseline for the next CPU-runtime architecture phase.
Future work can now compare against a stable reference before exploring:
- persistent worker pools
- reduced thread-creation overhead
- topology-aware scheduling
- NUMA-aware placement
- explicit affinity policies
- stronger receipt-native topology provenance
Those changes belong to a later release so that the performance effect of the new architecture can be measured against this frozen baseline.
Release identity
Version:
v0.4.0
Release commit:
6f17a734b9241359d36a9bf3d208b8527a456327
Repository:
QSOLKCB/GALAXY
This release is intended as the source snapshot for the corresponding Zenodo software record.