GALAXY v0.6.0 — Bit-Exact Stream–Reduce–Discard CPU Memory Architecture
Release Notes
GALAXY v0.6.0 formalizes the first completed CPU memory-wall phase after the host-aware execution architecture frozen in v0.5.0.
The central result is simple:
On the tested AMD Ryzen 9 5950X, GALAXY reduced its declared CPU algorithmic working set from 1,310,720 bytes (1,280 KiB; 1.25 MiB) to 245,760 bytes (240 KiB), an 81.25% reduction, while preserving the exact deterministic checksum and measuring a slightly lower median runtime under the tested workload.
This release freezes the evidence from PE #15 — Stream → Reduce → Discard.
The work does not replace GALAXY's canonical correctness model. Instead, it demonstrates that substantially less transient CPU state can be materialized while retaining the existing deterministic BAM-LUT, logical-addressing, and wrapping-u64 reduction contracts.
Release identity
Version: v0.6.0
Source commit: fa1c76fb49664ae4cdd6dc090cccd702399c2b60
Previous release: v0.5.0
Previous Zenodo DOI: 10.5281/zenodo.22774400
v0.6.0 contains the merged PE #15 memory-wall work from PR #15.
Headline result
Primary tested host:
AMD Ryzen 9 5950X 16-Core Processor
Linux x86_64
32 available logical CPUs
Primary workload:
logical population = 18,446,744,073,709,551,615
resident particles = 1,048,576
frames = 8
workers = 32
repeats = 5
seed = 303
Baseline:
reduction = materialized contribution array
LUT = full 16K sine + cosine Q2.30
microtile = 1024 particles
working set = 1,310,720 bytes
= 1,280 KiB
= 1.25 MiB
median = 9.521277 ms
peak VmHWM = 4,940 KiB
checksum = 1de0cecbb0955e44
Selected memory-efficient candidate:
reduction = fused hash + wrapping-u64 reduce
LUT = full 16K sine + cosine Q2.30
microtile = 128 particles
working set = 245,760 bytes
= 240 KiB
median = 9.424647 ms
peak VmHWM = 3,924 KiB
checksum = 1de0cecbb0955e44
Measured change:
algorithmic working-set reduction = 81.25%
working set = 5.3333x smaller
median runtime change = -1.0149%
process VmHWM reduction = 20.57%
The runtime difference is host- and workload-specific evidence and is not claimed as a universal speedup.
PE #15 — Stream → Reduce → Discard
The v0.5.0 worker-local SoA path retained a complete u64 contribution array for every active tile:
particle fields
↓
x/y projection arrays
↓
u64 contribution array
↓
wrapping-u64 reduction
↓
discard
PE #15 tested whether that intermediate contribution field could be removed entirely:
particle fields
↓
x/y projection arrays
↓
hash
↓
wrapping-u64 partial reduction
↓
discard
The fused path therefore reduces transient tiled storage from:
36 bytes / particle
to:
28 bytes / particle
while preserving the exact final checksum.
This is possible because GALAXY's contribution evidence is reduced by wrapping addition modulo 2^64, allowing transient contribution values to be folded into a deterministic partial sum instead of retained as a complete array.
Cache-sized microtiles
PE #15 tested:
1024
512
256
128
64
particles per worker microtile.
All microtile sizes preserved the same deterministic checksum.
For the fused path using the existing full 16K LUT:
Microtile Algorithmic working set Median
1024 1,048,576 B 9.390 ms
512 589,824 B 9.383 ms
256 360,448 B 9.429 ms
128 245,760 B 9.425 ms
64 188,416 B 9.546 ms
The 128-particle microtile represents the strongest measured balance between reduced working state and retained runtime on the tested Ryzen 9 5950X.
The 64-particle candidate reduced the algorithmic working set further to:
188,416 bytes
184 KiB
which is:
85.625% below the baseline
6.9565x smaller
while its median runtime was only approximately 0.26% above the baseline in this five-repeat run.
That result remains useful follow-up evidence but is not treated as establishing a universally superior tile size.
Fastest measured configuration
The fastest cell in the 30-case Ryzen matrix was:
reduction = fused hash/reduce
LUT = full 16K sine + cosine
microtile = 512
median = 9.383207 ms
working set = 589,824 bytes
Compared with the materialized 1024-particle baseline:
working-set reduction = 55%
median runtime change = -1.45%
This demonstrates that fused reduction can reduce memory substantially without requiring smaller microtiles to obtain a measured runtime benefit on the tested host.
Exact compressed BAM LUT experiments
v0.6.0 also preserves two exact LUT-compression experiments.
The canonical LUT stores:
16,384 cosine values
16,384 sine values
Q2.30 / i32
for:
131,072 bytes
Exact cosine-derived representation
The cosine candidate stores the complete cosine table plus exact signed sine corrections.
Measured storage:
98,304 bytes
Every reconstructed canonical LUT sample is checked exactly before execution.
Exact quarter-wave representation
The quarter-wave candidate stores:
4,097 canonical cosine samples
+
one-byte correction codes
+
a small exact signed correction palette
The tested representation used:
49,380 bytes
112 correction-palette entries
All 16,384 canonical sine/cosine sample pairs reconstructed exactly.
Representative non-table BAM interpolation results were also required to match the canonical LUT exactly.
No approximate trigonometric substitution is accepted.
Compressed LUT performance result
The compressed LUTs succeeded as exact memory representations but did not outperform the full LUT on the tested Ryzen 9 5950X.
For the fused 64-particle microtile:
full LUT:
median = 9.546157 ms
working set = 188,416 bytes
cosine-corrected:
median ≈ 9.85 ms
quarter-wave corrected:
median = 14.011176 ms
working set = 106,724 bytes
The quarter-wave candidate achieved the smallest complete algorithmic working set in the matrix:
106,724 bytes
≈ 104.2 KiB
This is approximately:
91.86% less memory
12.28x smaller
than the baseline.
However, its median runtime was approximately 47.2% higher than the baseline on this host.
Under the PE #15 acceptance boundary this is a memory success but a production HOLD.
The result is retained as useful negative evidence: exact LUT compression is possible, but on this Ryzen 9 5950X the reconstruction work costs more than the saved LUT traffic is worth.
Exactness result
The complete benchmark matrix contained:
2 reduction modes
×
3 LUT representations
×
5 microtile sizes
=
30 benchmark cells
Every cell produced:
checksum = 1de0cecbb0955e44
The PE #15 verifier additionally required:
exact full/cosine/quarter-wave LUT sample parity;
representative interpolation parity;
exact fused versus materialized reduction equality;
exact streaming-reference checksum parity;
exact parity across every required microtile;
deterministic worker-index reduction.
A memory or timing improvement is invalid if any exactness gate fails.
RSS measurement
Each matrix candidate was launched in an isolated process.
On Linux the receipt records:
/proc/self/status
VmHWM
This avoids the process-wide high-water contamination that would occur if multiple memory candidates were measured sequentially inside one process.
For the principal 128-particle candidate:
baseline VmHWM = 4,940 KiB
candidate VmHWM = 3,924 KiB
reduction = 1,016 KiB
≈ 20.57%
RSS remains an operating-system/process measurement and should not be confused with the explicitly calculated algorithmic working-set value.
Native code generation
The fused reduction path retains useful host-native SIMD code generation on the tested Zen 3 CPU.
Inspection of the exported fused reduction symbol showed vector integer operations including instructions from the VEX/AVX2-family execution path.
The experiment therefore did not obtain its memory reduction by falling back to a purely scalar contribution implementation.
ISA observations remain compiler-, build-, and host-specific.
Architecture conclusion
The PE #15 evidence supports the following CPU architecture:
deterministic logical range
↓
small worker-local microtile
↓
procedural particle generation
↓
BAM-LUT projection
↓
SIMD-friendly contribution hash
↓
immediate wrapping-u64 reduction
↓
discard transient tile
The important result is not simply a smaller allocation.
The experiment demonstrates that GALAXY does not need to retain every intermediate contribution in memory in order to preserve its deterministic evidence contract.
The strongest measured v0.6.0 candidate keeps the existing full LUT and attacks the larger source of transient memory directly:
do not materialize what can be reduced exactly
Relationship to v0.5.0
v0.5.0 established:
SIMD feasibility
→ bounded worker-local SoA
→ guarded production integration
→ persistent worker pools
→ topology-aware scheduling
→ calibrated host-aware execution
v0.6.0 adds:
stream
→ fuse contribution hashing and reduction
→ use cache-sized microtiles
→ discard transient state
The canonical and existing production galaxy-cpu command surfaces from v0.5.0 remain available.
PE #15 was implemented as an evidence-first experimental path so memory architecture could be evaluated without silently redefining the established runtime.
Roadmap boundary
During PE #15, other planned optimization work was deliberately deferred.
That includes:
heterogeneous CPU + GPU scheduling;
NUMA-local pools;
multi-GPU execution;
output-pipeline restructuring;
symmetry/orbit compression;
wider memory-budgeted auto-selection.
The purpose was to isolate one question:
How little CPU state does GALAXY actually need to retain while performing the exact existing workload?
v0.6.0 freezes the answer obtained from that experiment before additional execution architectures are layered on top.
Scientific and performance boundary
GALAXY remains a deterministic galaxy dynamics and visualization instrument.
The v0.6.0 memory work changes execution architecture and evidence storage, not the physical model.
This release does not claim:
that 128 particles is a universally optimal microtile;
that 512 particles is universally the fastest tile;
that fused reduction is universally faster on all CPUs;
that quarter-wave LUT compression is generally beneficial;
that Linux VmHWM equals algorithmic working-set size;
that the approximately 1.01% measured Ryzen improvement is a universal speedup;
that the complete logical u64 population is simultaneously resident;
that GALAXY is a self-consistent N-body solver.
The performance measurements are specific to the recorded host and workload.
The bit-exact checksum result is the relevant deterministic implementation property.
Reproducibility
Primary evidence archive:
GALAXY-memory-wall-Ryzen5950X-20260915T200828Z.zip
Archive SHA-256:
b2c0fee5ecb7a720e0da4d017ca0d3eac69e7d0db6f76283839f7c7c6d3515ff
The archive contains:
the 30-cell matrix;
individual JSON receipts;
individual benchmark logs;
host/toolchain evidence;
exactness verification output;
fused/materialized native disassembly;
SHA-256 integrity manifest.
Release progression
v0.4.0
deterministic native CPU baseline
+ Ryzen / EPYC scaling evidence
↓
v0.5.0
SIMD-aware bounded SoA
+ persistent workers
+ topology-aware scheduling
+ calibrated host-auto policy
↓
v0.6.0
Stream → Reduce → Discard
+ fused hash/reduction
+ cache-sized microtiles
+ exact compressed LUT experiments
+ 81.25% primary algorithmic working-set reduction
Release identity
GALAXY v0.6.0
Bit-Exact Stream–Reduce–Discard CPU Memory Architecture
Git tag:
v0.6.0
Source commit:
fa1c76fb49664ae4cdd6dc090cccd702399c2b60
Previous version:
v0.5.0
Previous Zenodo DOI:
10.5281/zenodo.22774400