Skip to content

Releases: QSOLKCB/GALAXY

GALAXY v0.8.0 — Parallel GPU Barnes–Hut Construction and Hardware Evidence Harness

Choose a tag to compare

@EmergentMonk EmergentMonk released this 27 Sep 16:50
Immutable release. Only release title and notes can be modified.
e051324

GALAXY v0.8.0 advances the resident Barnes–Hut self-gravity programme from the correctness-first GPU tree builder frozen in v0.7.0 to a parallel BH #2D construction path, while adding a fail-closed framework for capturing future real-hardware scaling evidence.

This release also promotes the browser’s planar Barnes–Hut N-body simulator to the default GALAXY landing experience while preserving the original rotation-law instrument separately.

Release identity

Version:          v0.8.0
Source commit:    e051324b1263f20ebb88d0e585992c7c1838504a
Previous release: v0.7.0
Previous commit:  7fe4dc63d40bb4afbb93f51c07d05b49f21e9756
Merged PR:        #24

v0.8.0 contains 107 commits beyond the v0.7.0 release boundary.

BH #2D — Parallel GPU tree construction

The principal compute milestone is a new parallel GPU Barnes–Hut tree builder.

Added:

  • workgroup-parallel root-bounds reduction;
  • parallel Morton-code generation;
  • eight stable 4-bit LSD radix passes;
  • deterministic sparse level-order cells;
  • parallel range and topology construction;
  • reverse-depth parallel mass and centre-of-mass aggregation;
  • a separate galaxy-bh-gpu-tree-parallel verifier;
  • resident workloads up to 65,536 bodies for the BH #2D path.

The frozen tree representation uses:

ordering:
stable-lsd-radix-4bit-morton-code;
resident body index retained for equal Morton keys

layout:
sparse-level-order-slot=depth*N+group_start

Those representation declarations are now validated as part of the evidence contract rather than treated as descriptive metadata.

Existing Barnes–Hut paths remain executable oracles

BH #2D does not replace the earlier correctness ladder.

The release retains:

  • BH #2A flat f64 CPU reference;
  • BH #2B2 host-built evolving GPU path;
  • BH #2C serialized device-side tree builder;
  • bounded exact direct-force probes.

For workloads within the configured oracle limits, the parallel builder is checked against these frozen reference surfaces.

Large workloads explicitly record oracle skips rather than hiding expensive CPU work inside benchmark measurements.

Real-hardware evidence harness

Added:

scripts/bench-bh2d-hardware.py

The harness performs a source-pinned resident-body scaling sweep through the BH #2D executable using --require-hardware.

The default sweep covers:

512
1024
2048
4096
8192
16384
32768
65536

Each accepted run is bound to:

  • one Git revision;
  • one selected GPU identity;
  • the enumerated adapter count;
  • GPU backend and device type;
  • Cargo and rustc identities;
  • compiler/linker/tool hashes;
  • Cargo configuration;
  • dependency-source hashes;
  • workload parameters;
  • receipt hashes;
  • raw run-log hashes;
  • benchmark samples;
  • correctness and topology invariants.

The resulting manifest schema is:

galaxy.bh2d-hardware-scaling-manifest.v1

Fail-closed provenance hardening

The hardware-evidence path received extensive adversarial review and now fails closed across source, build, loader, receipt and topology boundaries.

Clean launch boundary

Real evidence capture must begin through a dedicated clean launcher.

POSIX:

sh scripts/bench-bh2d-hardware-launch.sh \
  --output runs/bh2d-hardware-sweep \
  --adapter 0

Native Windows:

scripts\bench-bh2d-hardware-launch.cmd ...

The POSIX launcher starts isolated Python through a newly constructed environment rather than trusting loader variables visible after Python has already started.

Direct evidence capture through the Python file is rejected.

The manifest records the Python executable and its SHA-256 together with:

launch_boundary = clean-environment-launcher-v1

Platform-aware toolchain provenance

The earlier Unix-only tool assumption has been replaced by platform-specific tooling.

Linux and macOS record their native compiler/linker/Git tools.

Windows supports the active MSVC/LLVM environment without requiring a POSIX cc, while selected tools are resolved, hashed and used to construct the narrowed build PATH.

Native Windows CI also verifies that launcher failures propagate the Python process exit status correctly.

Build-environment isolation

Evidence capture rejects or removes build-affecting environment controls including:

  • LD_* loader injection;
  • the DYLD_* namespace;
  • Vulkan ICD, driver and layer overrides;
  • additive Vulkan driver/layer paths;
  • Vulkan loader enable/disable selectors;
  • Rust and Cargo compiler/wrapper overrides;
  • target-specific linker/runner flags;
  • Cargo profile overrides;
  • C/C++ compiler and linker flags.

Dependency-source binding

Cargo dependency trees are fingerprinted independently of their physical cache location.

Registry and Git dependencies remain hashed even if CARGO_HOME is located beneath the repository checkout.

Dependency fingerprints include:

  • relative path;
  • file type;
  • permission/mode bits;
  • file contents.

Executable-bit changes therefore alter the dependency tree fingerprint.

Git and source-tree provenance

The harness:

  • binds Git operations to the explicit checkout;
  • disables replacement objects;
  • rejects index flags that could conceal changes;
  • compares raw tracked bytes against index blobs;
  • detects executable-mode changes;
  • rejects unexpected compiler-affecting files;
  • validates effective Cargo configuration;
  • checks source and toolchain provenance throughout the sweep.

Strict receipt validation

BH #2D receipts now undergo substantially stronger structural validation.

Checks include:

  • strict UTF-8 JSON;
  • rejection of NaN and infinity constants;
  • rejection of duplicate JSON object names;
  • integer-type enforcement for counters;
  • workload binding;
  • hardware measurement classification;
  • adapter index/count consistency;
  • numeric and textual adapter selector parity with the Rust producer;
  • allowed backend set: Vulkan, Metal or DX12;
  • software-adapter detection;
  • driver provenance;
  • zero host tree rebuilds;
  • zero host particle readbacks during evolution;
  • force-solve and tree-build counts;
  • tree checksums;
  • frozen ordering and layout declarations;
  • leaf/internal-cell fanout constraints;
  • per-depth occupancy constraints;
  • bucket-size constraints;
  • shared root-to-leaf path capacity;
  • force-error RMS/max consistency;
  • state-error zero/nonzero consistency;
  • benchmark sample and median recomputation;
  • BH #2C comparison/skip boundaries;
  • exact BH #2D allocation formula.

Failed verifier runs retain cryptographically bound diagnostics.

A failed manifest records the run log path and SHA-256, and when an unusable receipt exists its path and hash are retained as well.

This also applies when the verifier exits successfully but emits a missing, malformed or validator-rejected receipt.

Browser N-body home page

index.html is now the default interactive planar Barnes–Hut N-body simulator.

The default scene uses 768 mutually interacting resident bodies, with selectable populations from 128 to 2,048.

The browser experience includes:

  • binary encounter, rotating-disc and cold-collapse presets;
  • pause and resume;
  • deterministic single-step;
  • seed control;
  • orbit and zoom controls;
  • Barnes–Hut tree overlay;
  • direct-force auditing;
  • glow sprites and short trajectory trails;
  • reduced-motion handling;
  • hidden-tab suspension;
  • refresh-rate-independent physics.

Physics advances on a fixed 60 Hz schedule rather than using display refresh as the simulation clock.

Overload drops excess catch-up work instead of enlarging the physical timestep.

Preserved browser instruments

The previous prescribed-field rotation-law application remains available at:

rotation-lab.html

The detailed Barnes–Hut laboratory remains at:

barnes-hut.html

This keeps the historical rotation-law interface available while allowing the resident N-body simulator to become GALAXY’s default browser entrypoint.

Validation

The merged implementation passed the relevant repository validation surfaces, including:

  • native Rust runtime tests;
  • Barnes–Hut CPU and GPU reference validation;
  • BH #2D host-side contract tests;
  • clean-launch dry-run validation;
  • native Windows launcher failure propagation;
  • Python 3.13 and 3.14 runner checks;
  • browser N-body integration tests;
  • refresh-rate invariance;
  • pause/single-step behavior;
  • visibility and reduced-motion timing;
  • static Pages distribution checks;
  • native CPU runtime;
  • CPU host-auto promotion;
  • retro CPU portability;
  • CPU memory-wall validation.

The Native GPU CI path also executes the Barnes–Hut stack through software Vulkan for correctness validation.

Evidence boundary

This release does not claim measured hardware GPU speedup.

Software Vulkan remains correctness and integration evidence only.

The new harness establishes the machinery required to collect defensible real-GPU evidence, but a completed real-hardware manifest has not yet been produced as part of this release.

Therefore:

  • BH #2D parallel construction is implemented;
  • its correctness/provenance contract is frozen;
  • real-hardware scaling evidence remains pending;
  • BH #2E production promotion remains a separate decision.

A completed future hardware manifest will support claims only for its recorded:

  • source revision;
  • host;
  • toolchain;
  • dependency tree;
  • adapter;
  • workload;
  • benchmark samples.

What remains deliberately deferred

This release does not claim:

  • production-quality BH #2D performance;
  • measured real-GPU speedup;
  • multi-GPU Barnes–Hut partitioning;
  • deterministic distributed tree construction;
  • logical-u64 mutually interact...
Read more

GALAXY v0.7.0 - GALAXY — Barnes–Hut Self-Gravity and GPU Tree Construction

Choose a tag to compare

@EmergentMonk EmergentMonk released this 27 Sep 05:35
Immutable release. Only release title and notes can be modified.
7fe4dc6

This release substantially expands GALAXY beyond its established prescribed-field and logical-u64 particle runtimes by adding a separate, opt-in resident self-gravity execution family.

The new Barnes–Hut path progresses from a deterministic CPU reference through Morton-ordered flat trees, executable GPU force traversal, persistent multi-step GPU evolution, and finally GPU-side tree construction.

The existing browser, native CPU, fixed-potential GPU, and logical-u64 execution contracts remain intact.

Major additions

Resident Barnes–Hut self-gravity

Added a deterministic planar Barnes–Hut N-body reference implementation with:

  • mutually coupled resident bodies;
  • deterministic quadtree construction;
  • total-mass and centre-of-mass aggregation;
  • classical Barnes–Hut s / d < theta opening control;
  • explicit rejection of aggregate cells containing the target body;
  • Plummer-style gravitational softening;
  • kick–drift–kick leapfrog integration;
  • deterministic disc and collision fixtures;
  • bounded-depth handling for coincident bodies;
  • exact O(N²) direct-force verification.

This execution family is intentionally separate from GALAXY's logical-u64 tiled runtime because self-gravity couples the complete resident population.

Independent direct-force oracle

Added a separate all-pairs force evaluator so the Barnes–Hut approximation does not validate itself.

Validation includes:

  • theta-zero parity against direct forces;
  • bounded RMS and worst-relative error gates;
  • deterministic probe selection;
  • bounded direct-force probes for larger workloads;
  • explicit avoidance of hidden O(N²) verification costs in high-count GPU runs.

Barnes–Hut browser laboratory

Added a dedicated browser entrypoint:

barnes-hut.html

The laboratory provides:

  • live evolving self-gravity;
  • quadtree visualization;
  • tree-depth overlays;
  • force-term accounting;
  • direct-force probe audits;
  • theta, softening, bucket and timestep controls;
  • deterministic seeded presets;
  • binary-disc encounter;
  • rotating-disc and cold-collapse scenarios.

The visualizer makes the tree approximation and direct-force error surface observable rather than treating Barnes–Hut as a black box.

Native Barnes–Hut CPU runtime

Added the standalone nbody/ Rust crate.

It provides:

  • recursive Barnes–Hut reference traversal;
  • exact direct-force evaluation;
  • deterministic fixtures;
  • leapfrog integration;
  • checksums;
  • force and tree statistics;
  • verification commands;
  • machine-readable receipts.

New commands include:

verify

run

and later flat-tree verification and probe commands.

BH #2A — Morton / flat-tree substrate

Added a GPU-oriented pointer-free Barnes–Hut representation.

Features include:

  • 16-bit-per-axis coordinate quantization;
  • 32-bit Morton/Z-order keys;
  • stable spatial ordering;
  • preservation of resident order for equal Morton keys;
  • explicit body-to-Morton-position mapping;
  • flat cell arrays;
  • explicit u32 child indices;
  • contiguous Morton body ranges;
  • bottom-up mass and centre-of-mass aggregation;
  • iterative stack traversal;
  • deterministic topology checksums.

The flat-tree implementation retains f64 arithmetic so topology migration is verified independently from later GPU precision changes.

New flat-tree verification

Added:

verify-flat

and:

flat-probe

These verify:

  • theta-zero parity with the direct O(N²) oracle;
  • theta-0.5 accuracy;
  • agreement with the recursive Barnes–Hut implementation;
  • stable equal-key ordering;
  • configured maximum-depth handling;
  • exact root range and total mass preservation;
  • repeatable topology checksums.

Tree construction and traversal timing are reported separately.

BH #2B1 — GPU transfer ABI and force traversal

Added a frozen f32/u32 GPU-facing Barnes–Hut ABI.

Record sizes are explicitly defined for:

  • settings;
  • resident bodies;
  • Morton entries;
  • flat cells;
  • acceleration outputs.

A canonical packed-byte fixture verifies cross-backend layout consistency.

Vulkan / WGSL Barnes–Hut traversal

Added real Barnes–Hut force traversal through wgpu and WGSL.

The shader performs:

  • iterative flat-tree traversal;
  • leaf direct-force evaluation;
  • Barnes–Hut aggregate acceptance;
  • Morton-range self-exclusion;
  • deterministic child visitation;
  • f32 acceleration output.

This is an actual GPU compute path, not a decorative or disconnected accelerator layer.

CUDA parity surface

Added matched CUDA Barnes–Hut traversal source with:

  • equivalent transfer-record layout;
  • compile-time size assertions;
  • matching traversal semantics;
  • target-range self-exclusion.

Host-side CI verifies CUDA ABI/source parity without falsely claiming CUDA execution on non-NVIDIA runners.

GPU traversal verifier

Added:

galaxy-bh-gpu

The verifier reports:

  • CPU tree construction;
  • f32 packing;
  • GPU transfer;
  • GPU traversal dispatch;
  • GPU readback;
  • full GPU-vs-flat-CPU comparison;
  • bounded deterministic direct-force probes.

Large particle counts no longer trigger a complete O(N²) CPU solve before GPU execution.

BH #2B2 — evolving GPU self-gravity

Added persistent GPU-resident N-body state.

New GPU state includes:

  • position;
  • mass;
  • velocity;
  • persistent acceleration buffers.

Added WGSL integration kernels for:

  • half-kick + drift;
  • final half-kick.

The resulting multi-step execution contract is:

GPU force -> GPU kick/drift -> tree rebuild -> GPU force -> GPU final kick

Boundary acceleration is reused, so an N-step run performs N+1 force solves, rather than recomputing both force boundaries independently every step.

New evolution executable

Added:

galaxy-bh-evolve

Features include:

  • disc and collision presets;
  • multi-step GPU evolution;
  • persistent GPU state;
  • explicit tree-rebuild synchronization boundaries;
  • final direct-force probes;
  • state and topology checksums;
  • conservation diagnostics;
  • bounded full CPU trajectory comparison.

Receipts include:

  • force-solve count;
  • tree-rebuild count;
  • topology changes;
  • state checksums;
  • total mass drift;
  • centre-of-mass drift;
  • linear-momentum drift;
  • angular-momentum drift;
  • separate CPU/GPU stage timings.

BH #2C — GPU tree construction

GALAXY now constructs the Barnes–Hut tree on the GPU for the BH #2C execution path.

The device-side tree build performs:

  • GPU root-bounds calculation;
  • parallel Morton key generation;
  • deterministic GPU bitonic ordering;
  • resident-body sorted-position assignment;
  • flat-cell topology construction;
  • bottom-up mass and centre-of-mass aggregation;
  • direct hand-off to the existing GPU force traversal.

During BH #2C evolution there are:

  • zero host particle readbacks inside the step loop;
  • zero CPU tree rebuilds inside the step loop.

The persistent drifted GPU state feeds the GPU tree builder directly.

Deterministic GPU ordering

Morton entries are ordered by:

(Morton code, body index)

This supplies a deterministic total ordering and preserves the intent of the stable host ordering when Morton keys collide.

GPU topology contract

BH #2C intentionally defines a new f32 GPU topology representation rather than claiming bit-identical cell numbering with the f64 host tree.

The GPU tree retains the important scientific invariants:

  • contiguous sorted body ranges;
  • explicit child links;
  • bounded Morton depth;
  • target-range membership;
  • deterministic ordering;
  • bottom-up aggregates.

Correctness is established by force and trajectory comparison against the frozen BH #2A/B2B2 references.

BH #2C verifier

Added:

galaxy-bh-gpu-tree

The verifier performs complete multi-step self-gravity with GPU-built trees and records:

  • GPU tree-build count;
  • force-solve count;
  • zero host rebuild/readback assertions;
  • tree buffer sizes;
  • cell, leaf and maximum-depth statistics;
  • root bounds;
  • deterministic repeated-tree checksums;
  • GPU-vs-BH #2A force error;
  • bounded direct-force error;
  • full GPU-vs-f64 trajectory error;
  • stage-specific GPU tree-construction timings.

The initial BH #2C builder is intentionally correctness-first and limited to 4,096 resident bodies.

Control-heavy bounds, ordering, topology and aggregate stages are currently serialized GPU kernels. This is a device-ownership and correctness milestone, not a production GPU-tree performance claim.

Validation results

The deterministic Mesa Vulkan CI fixture successfully executed the complete GPU Barnes–Hut stack.

For a 128-body, three-step self-gravity run:

  • host tree rebuilds during BH #2C evolution: 0
  • host particle readbacks during steps: 0
  • force solves: 4
  • evolution tree builds: 4
  • GPU tree overflow: none
  • repeated unchanged-state GPU tree checksum: identical

Final GPU-tree force versus BH #2A flat f64 reference:

  • RMS relative error: approximately 4.22e-7
  • maximum relative error: approximately 2.72e-6

Final bounded exact direct-force probes:

  • RMS relative error: approximately 0.879%
  • maximum relative error: approximately 2.14%

Three-step GPU trajectory versus the f64 BH #2A reference:

  • position RMS relative error: approximately 5.17e-8
  • position maximum relative error: approximately 1.52e-7
  • velocity RMS relative error: approximately 9.19e-8
  • velocity maximum relative error: approximately 2.92e-7

These figures are correctness evidence from Mesa software Vulkan and are not hardware GPU performance claims.

CI and reproducibility

The native GPU workflow now validates:

  • Barnes–Hut Rust tests;
  • transfer-record ABI;
  • CUDA source/layout parity;
  • BH #2B1 GPU traversal;
  • BH #2B2 ...
Read more

GALAXY v0.6.0 — Bit-Exact Stream–Reduce–Discard CPU Memory Architecture

Choose a tag to compare

@EmergentMonk EmergentMonk released this 15 Sep 20:23
Immutable release. Only release title and notes can be modified.
fa1c76f

Release Notes

GALAXY v0.6.0 formalizes the first completed CPU memory-wall phase after the host-aware execution architecture frozen in v0.5.0.

The central result is simple:

On the tested AMD Ryzen 9 5950X, GALAXY reduced its declared CPU algorithmic working set from 1,310,720 bytes (1,280 KiB; 1.25 MiB) to 245,760 bytes (240 KiB), an 81.25% reduction, while preserving the exact deterministic checksum and measuring a slightly lower median runtime under the tested workload.

This release freezes the evidence from PE #15 — Stream → Reduce → Discard.

The work does not replace GALAXY's canonical correctness model. Instead, it demonstrates that substantially less transient CPU state can be materialized while retaining the existing deterministic BAM-LUT, logical-addressing, and wrapping-u64 reduction contracts.

Release identity

Version: v0.6.0
Source commit: fa1c76fb49664ae4cdd6dc090cccd702399c2b60
Previous release: v0.5.0
Previous Zenodo DOI: 10.5281/zenodo.22774400

v0.6.0 contains the merged PE #15 memory-wall work from PR #15.

Headline result

Primary tested host:

AMD Ryzen 9 5950X 16-Core Processor
Linux x86_64
32 available logical CPUs

Primary workload:

logical population = 18,446,744,073,709,551,615
resident particles = 1,048,576
frames             = 8
workers            = 32
repeats            = 5
seed               = 303

Baseline:

reduction          = materialized contribution array
LUT                = full 16K sine + cosine Q2.30
microtile          = 1024 particles
working set        = 1,310,720 bytes
                   = 1,280 KiB
                   = 1.25 MiB
median             = 9.521277 ms
peak VmHWM         = 4,940 KiB
checksum           = 1de0cecbb0955e44

Selected memory-efficient candidate:

reduction          = fused hash + wrapping-u64 reduce
LUT                = full 16K sine + cosine Q2.30
microtile          = 128 particles
working set        = 245,760 bytes
                   = 240 KiB
median             = 9.424647 ms
peak VmHWM         = 3,924 KiB
checksum           = 1de0cecbb0955e44

Measured change:

algorithmic working-set reduction = 81.25%
working set                        = 5.3333x smaller
median runtime change              = -1.0149%
process VmHWM reduction            = 20.57%

The runtime difference is host- and workload-specific evidence and is not claimed as a universal speedup.

PE #15 — Stream → Reduce → Discard

The v0.5.0 worker-local SoA path retained a complete u64 contribution array for every active tile:

particle fields
    ↓
x/y projection arrays
    ↓
u64 contribution array
    ↓
wrapping-u64 reduction
    ↓
discard

PE #15 tested whether that intermediate contribution field could be removed entirely:

particle fields
    ↓
x/y projection arrays
    ↓
hash
    ↓
wrapping-u64 partial reduction
    ↓
discard

The fused path therefore reduces transient tiled storage from:

36 bytes / particle

to:

28 bytes / particle

while preserving the exact final checksum.

This is possible because GALAXY's contribution evidence is reduced by wrapping addition modulo 2^64, allowing transient contribution values to be folded into a deterministic partial sum instead of retained as a complete array.

Cache-sized microtiles

PE #15 tested:

1024
512
256
128
64

particles per worker microtile.

All microtile sizes preserved the same deterministic checksum.

For the fused path using the existing full 16K LUT:

Microtile	Algorithmic working set	Median
1024	1,048,576 B	9.390 ms
512	589,824 B	9.383 ms
256	360,448 B	9.429 ms
128	245,760 B	9.425 ms
64	188,416 B	9.546 ms

The 128-particle microtile represents the strongest measured balance between reduced working state and retained runtime on the tested Ryzen 9 5950X.

The 64-particle candidate reduced the algorithmic working set further to:

188,416 bytes
184 KiB

which is:

85.625% below the baseline
6.9565x smaller

while its median runtime was only approximately 0.26% above the baseline in this five-repeat run.

That result remains useful follow-up evidence but is not treated as establishing a universally superior tile size.

Fastest measured configuration

The fastest cell in the 30-case Ryzen matrix was:

reduction = fused hash/reduce
LUT       = full 16K sine + cosine
microtile = 512
median    = 9.383207 ms
working set = 589,824 bytes

Compared with the materialized 1024-particle baseline:

working-set reduction = 55%
median runtime change  = -1.45%

This demonstrates that fused reduction can reduce memory substantially without requiring smaller microtiles to obtain a measured runtime benefit on the tested host.

Exact compressed BAM LUT experiments

v0.6.0 also preserves two exact LUT-compression experiments.

The canonical LUT stores:

16,384 cosine values
16,384 sine values
Q2.30 / i32

for:

131,072 bytes
Exact cosine-derived representation

The cosine candidate stores the complete cosine table plus exact signed sine corrections.

Measured storage:

98,304 bytes

Every reconstructed canonical LUT sample is checked exactly before execution.

Exact quarter-wave representation

The quarter-wave candidate stores:

4,097 canonical cosine samples
+
one-byte correction codes
+
a small exact signed correction palette

The tested representation used:

49,380 bytes
112 correction-palette entries

All 16,384 canonical sine/cosine sample pairs reconstructed exactly.

Representative non-table BAM interpolation results were also required to match the canonical LUT exactly.

No approximate trigonometric substitution is accepted.

Compressed LUT performance result

The compressed LUTs succeeded as exact memory representations but did not outperform the full LUT on the tested Ryzen 9 5950X.

For the fused 64-particle microtile:

full LUT:
  median = 9.546157 ms
  working set = 188,416 bytes

cosine-corrected:
  median ≈ 9.85 ms

quarter-wave corrected:
  median = 14.011176 ms
  working set = 106,724 bytes

The quarter-wave candidate achieved the smallest complete algorithmic working set in the matrix:

106,724 bytes
≈ 104.2 KiB

This is approximately:

91.86% less memory
12.28x smaller

than the baseline.

However, its median runtime was approximately 47.2% higher than the baseline on this host.

Under the PE #15 acceptance boundary this is a memory success but a production HOLD.

The result is retained as useful negative evidence: exact LUT compression is possible, but on this Ryzen 9 5950X the reconstruction work costs more than the saved LUT traffic is worth.

Exactness result

The complete benchmark matrix contained:

2 reduction modes
×
3 LUT representations
×
5 microtile sizes
=
30 benchmark cells

Every cell produced:

checksum = 1de0cecbb0955e44

The PE #15 verifier additionally required:

exact full/cosine/quarter-wave LUT sample parity;
representative interpolation parity;
exact fused versus materialized reduction equality;
exact streaming-reference checksum parity;
exact parity across every required microtile;
deterministic worker-index reduction.

A memory or timing improvement is invalid if any exactness gate fails.

RSS measurement

Each matrix candidate was launched in an isolated process.

On Linux the receipt records:

/proc/self/status
VmHWM

This avoids the process-wide high-water contamination that would occur if multiple memory candidates were measured sequentially inside one process.

For the principal 128-particle candidate:

baseline VmHWM = 4,940 KiB
candidate VmHWM = 3,924 KiB

reduction = 1,016 KiB
          ≈ 20.57%

RSS remains an operating-system/process measurement and should not be confused with the explicitly calculated algorithmic working-set value.

Native code generation

The fused reduction path retains useful host-native SIMD code generation on the tested Zen 3 CPU.

Inspection of the exported fused reduction symbol showed vector integer operations including instructions from the VEX/AVX2-family execution path.

The experiment therefore did not obtain its memory reduction by falling back to a purely scalar contribution implementation.

ISA observations remain compiler-, build-, and host-specific.

Architecture conclusion

The PE #15 evidence supports the following CPU architecture:

deterministic logical range
        ↓
small worker-local microtile
        ↓
procedural particle generation
        ↓
BAM-LUT projection
        ↓
SIMD-friendly contribution hash
        ↓
immediate wrapping-u64 reduction
        ↓
discard transient tile

The important result is not simply a smaller allocation.

The experiment demonstrates that GALAXY does not need to retain every intermediate contribution in memory in order to preserve its deterministic evidence contract.

The strongest measured v0.6.0 candidate keeps the existing full LUT and attacks the larger source of transient memory directly:

do not materialize what can be reduced exactly
Relationship to v0.5.0

v0.5.0 established:

SIMD feasibility
→ bounded worker-local SoA
→ guarded production integration
→ persistent worker pools
→ topology-aware scheduling
→ calibrated host-aware execution

v0.6.0 adds:

stream
→ fuse contribution hashing and reduction
→ use cache-sized microtiles
→ discard transient state

The canonical and existing production galaxy-cpu command surfaces from v0.5.0 remain available.

PE #15 was implemented as an evidence-first experimental path so memory architecture could be evaluated without silently redefining the established runtime.

Roadmap boundary

During PE #15, other planned optimization work was deliberately deferred.

That includes:

heterogeneous CPU + GPU scheduling;
NUMA-local pools;
multi-GPU execution;
output-pipeline restructuring;
symmetry/orbit ...
Read more

GALAXY v0.5.0 — Deterministic Host-Aware CPU Execution and Memory-Bounded Runtime Architecture

Choose a tag to compare

@EmergentMonk EmergentMonk released this 15 Sep 17:25
Immutable release. Only release title and notes can be modified.
b2e8603

#Release Notes

GALAXY v0.5.0 advances the native CPU runtime introduced in v0.4.0 into a layered, deterministic optimization architecture.

The central change is not one isolated benchmark improvement. The CPU runtime now has a complete progression from SIMD feasibility testing, through bounded worker-local memory, guarded production integration, persistent worker pools, topology-aware scheduling, and finally calibrated host-aware automatic promotion.

Throughout this work the canonical execution path remains available as the correctness oracle. Optimized paths are required to preserve exact deterministic checksum parity and fail closed when that contract is violated.

This release incorporates the work from PRs #10 through #14.

Highlights

  • Deterministic SIMD/autovectorization probe
  • AVX2 / AVX-512 code-generation evidence tooling
  • Worker-local structure-of-arrays CPU execution
  • Bounded per-worker particle tiles instead of full-resident optimized state
  • Production bench-soa / verify-soa execution paths
  • Persistent worker pools with reusable local buffers
  • Best-effort physical-core topology detection
  • Explicit physical-first and logical/SMT scheduling policies
  • Host-aware automatic runtime and tile calibration
  • Exact streaming canonical BAM-LUT oracle
  • 5% promotion margin before leaving the canonical path
  • Effective tile-shape-faithful calibration
  • Requested frame-depth-faithful calibration
  • Full persistent startup + teardown lifecycle accounting
  • Explicit topology, oracle, lifecycle and RSS evidence scopes
  • Cross-platform native CPU CI on Linux x86-64, Linux ARM64, macOS ARM64 and Windows x86-64
  • New memory-wall / heterogeneous execution roadmap

CPU optimization sequence

v0.5.0 freezes the following CPU architecture progression:

PR #10 SIMD / autovectorization feasibility
PR #11 worker-local bounded SoA execution
PR #12 guarded production SoA integration
PR #13 persistent workers + topology-aware scheduling
PR #14 calibrated host-aware promotion with fail-closed parity

Each stage preserved the established deterministic BAM-LUT and wrapping-u64 checksum contract before the next stage was allowed to build on it.

PR #10 — Deterministic CPU SIMD probe

PR #10 introduced an isolated deterministic probe for GALAXY's contribution-hash stage.

The probe:

converts the contribution inputs into a structure-of-arrays batch form;
compares every result against the canonical scalar contribution function;
requires repeated checksum stability;
records compile-time and runtime CPU feature evidence;
compares generic x86-64 and target-cpu=native builds;
extracts the batch symbol for disassembly inspection;
records XMM / YMM / ZMM evidence when available;
rejects missing or malformed disassembly evidence rather than silently passing.

The local Ryzen 9 5950X evidence recorded:

generic median 34,942,130 ns
native median 10,742,437 ns
native speedup 3.252719099x
time reduction 69.256491%
checksum e3f89b31c3f4c4e8

The generic build used legacy SSE2 packed operations while the native Zen 3 build used AVX2/VEX packed operations.

This result was deliberately treated as an isolated kernel result, not an end-to-end GALAXY speedup.

PR #11 — Worker-local bounded SoA execution

PR #11 moved the SIMD-friendly contribution shape into a complete BAM-LUT execution experiment.

The reference path retained a full resident array-of-structures representation.

The optimized path instead uses:

logical particle range
↓
bounded worker-local compact SoA tile
↓
BAM-LUT projection
↓
SIMD-friendly contribution batch
↓
wrapping-u64 local reduction
↓
deterministic worker reduction

Each worker owns compact particle fields:

id_lo
id_hi
radius_q16
initial_bam
delta_bam

plus bounded x/y/contribution scratch storage.

The important memory change is architectural:

reference working state ∝ total resident population

worker-local SoA working state ∝ workers × tile size

The optimized path therefore no longer requires a complete resident particle array merely to execute the deterministic workload.

The experiment also added:

generic/native worker and tile sweeps;
reference-vs-SoA checksum gates;
Linux peak-RSS evidence in isolated benchmark processes;
exact worker-count invariance tests;
actual allocated tile-capacity reporting;
decoded SIMD evidence for the production-shaped hash batch;
clean benchmark-output directory requirements.

A timing result is invalid if checksum parity fails.

PR #12 — Guarded production SoA integration

The successful worker-local SoA path was integrated into the production galaxy-cpu binary without silently replacing the canonical runtime.

New commands:

galaxy-cpu verify-soa [--workers N] [--tile N]

galaxy-cpu bench-soa
[--logical U64] [--resident N] [--frames N]
[--workers N] [--tile N] [--repeats N]
[--seed U32] [--receipt PATH]

The existing canonical commands remain:

galaxy-cpu verify
galaxy-cpu bench

Integrated SoA receipts identify the path explicitly:

schema = galaxy.cpu-runtime-soa-receipt.v1
runtime = galaxy-cpu
execution_mode = worker-local-soa-guarded
guarded_opt_in = true
canonical_fallback = bench

The optimized path remains explicit and auditable rather than becoming an invisible default.

The production dispatcher/build arrangement was also hardened so the canonical runtime and optimized modules are exposed through one controlled galaxy-cpu command surface.

PR #13 — Persistent topology-aware worker execution

PR #13 removed repeated worker creation from steady-state optimized execution.

New commands:

galaxy-cpu verify-soa-pool
[--workers N] [--tile N]
[--schedule physical-first|logical]

galaxy-cpu bench-soa-pool
[--logical U64] [--resident N] [--frames N]
[--workers N] [--tile N]
[--schedule physical-first|logical]
[--repeats N] [--seed U32] [--receipt PATH]

Persistent workers:

are created once;
own deterministic contiguous resident ranges;
keep one reusable local SoA tile;
reuse that storage across warm-up and measured trials;
perform the same deterministic particle generation, BAM-LUT projection and contribution hashing as the spawned SoA path;
may complete in arbitrary order;
are always reduced in deterministic worker-index order.

Two scheduling policies are explicit:

physical-first
logical

physical-first uses detected physical-core availability where reliable.

logical permits SMT workers explicitly.

Topology detection is best-effort:

Linux: allowed CPU set plus sysfs package/core topology;
macOS: hw.physicalcpu;
unsupported/unavailable physical topology: explicit logical fallback.

This release does not equate topology detection with CPU affinity or NUMA pinning.

PR #14 — Calibrated host-aware automatic promotion

PR #14 adds the automatic selection layer above the canonical, spawned-SoA and persistent-SoA paths.

New commands:

galaxy-cpu verify-auto [--workers N]

galaxy-cpu bench-auto
[--logical U64] [--resident N] [--frames N]
[--workers N] [--repeats N]
[--seed U32] [--receipt PATH]

Policy identifier:

calibrated-host-auto-v1

The tuner does not use CPU model-name lookup tables.

It measures the actual host and workload.

Candidate families include:

canonical
spawned SoA
persistent physical-first SoA
persistent logical/SMT SoA

across the evidence-backed tile set:

1024
4096
16384
65536
Tile-faithful calibration

A nominal tile size is not treated as measured unless calibration actually exercises that effective per-worker tile size.

For example, a nominal 65,536-particle tile is not considered a 65,536-particle calibration when worker partitioning only gives the worker 256 particles.

The calibration resident population therefore begins from a bounded base but expands when necessary to preserve the requested workload's effective tile shape.

It never expands beyond the requested resident population.

Policy:

tile_shape_policy = expand-resident-for-effective-tile-v1

Every tiled candidate records:

calibration_effective_tile_particles
requested_effective_tile_particles
tile_shape_match

A mismatch fails tuning closed.

Frame-faithful calibration

Calibration also preserves the requested frame depth.

This matters because particle generation and local-tile fill are primarily per-particle costs, while BAM projection and contribution hashing repeat per frame.

A four-frame calibration cannot safely represent the cost mix of an arbitrary deep-frame workload merely by multiplying the result afterward.

Policy:

frame_calibration_policy = preserve-requested-depth-v1

Therefore:

calibration_frames = requested_frames
Full-requested-workload score projection

Candidate timing evidence is compared on the same requested-workload scale.

calibration_work =
calibration_resident × calibration_frames

requested_work =
requested_resident × requested_frames

projected_median =
ceil(
calibration_median ×
requested_work /
calibration_work
)

Policy:

score_projection = linear-particle-frame-v1

This projection is an explicit tuning heuristic, not a theorem of perfectly linear runtime scaling.

Persistent lifecycle scoring

Persistent candidates now account for the complete pool lifecycle.

Startup includes:

worker creation;
worker-local allocation;
explicit first-touch / page commitment of all local SoA and scratch buffers.

Teardown includes:

shutdown signalling;
worker joins;
worker-local buffer destruction and deallocation.

The scored lifecycle is:

full_pool_lifecycle =
full_pool_startup +
full_pool_teardown

and:

persistent_score =
projected_median +
full_pool_lifecycle / requested...

Read more

GALAXY v0.4.0 — Deterministic Native CPU Runtime and Scaling Evidence

Choose a tag to compare

@EmergentMonk EmergentMonk released this 14 Sep 19:18
Immutable release. Only release title and notes can be modified.
6f17a73

RELEASE NOTES:

GALAXY v0.4.0 freezes the first formally reproducible native-CPU performance milestone for the project.

This release combines the deterministic galaxy instrument, native GPU runtime, exact-u64 logical addressing, native multicore CPU runtime, BAM32/Q2.30 lookup-table backend, local Ryzen validation, and independent qBraid/Azure EPYC scaling evidence into one archival release.

The emphasis of v0.4.0 is not simply higher benchmark numbers. It is reproducible execution: exact source identity, deterministic checksums, explicit runtime receipts, preserved hardware/topology evidence, and conservative boundaries around what the measurements do and do not establish.

Highlights

  • Native deterministic CPU runtime for GALAXY
  • Exact positive-u64 logical population support
  • Float and BAM-LUT CPU backends
  • Deterministic scalar/parallel checksum parity
  • Multicore scaling through 96 available logical CPUs
  • Ryzen 9 5950X local validation
  • AMD EPYC 7763 qBraid/Azure replication
  • 32 / 48 / 64 / 96 worker scaling study
  • NUMA and SMT topology capture
  • Controlled CPU-affinity experiments
  • Preserved raw benchmark receipts and evidence archive
  • Explicit affinity-command provenance
  • Existing Vulkan/CUDA GPU runtime retained
  • Reproducibility and scientific claim boundaries documented

Native CPU runtime

The native CPU runtime evolves independent resident test particles in GALAXY's fixed potential while preserving the project's deterministic logical-addressing contract.

The runtime supports:

  • exact positive-u64 logical populations up to:

    18,446,744,073,709,551,615

  • bounded resident populations up to:

    16,777,216

  • deterministic proportional logical-ID sampling

  • split-u64-hash32-avalanche-v1 identity addressing

  • Float/libm execution

  • BAM32/Q2.30 LUT execution

  • deterministic contiguous worker partitioning

  • scalar and parallel checksum comparison

  • fixed worker-count measurement

  • interleaved Float/BAM timing

  • machine-readable CPU runtime receipts

  • strict verification mode

The CPU implementation does not replace GALAXY's GPU runtime. It provides an independently measurable and reproducible native execution path.

BAM32 lookup-table backend

v0.4.0 includes the deterministic BAM-LUT CPU path derived from the isolated retro-math reference work.

The LUT path uses:

  • BAM32 angular representation
  • Q2.30 fixed-point lookup values
  • 16,384 LUT entries
  • deterministic conversion
  • sampled LUT diagnostic validation

The current diagnostic evaluates 8,193 angles and observed a sampled maximum absolute Q30 error of:

255

This value is a sampled implementation diagnostic, not a mathematical maximum over the full 2^32 BAM angle space.

Local Ryzen 9 5950X validation

The native CPU runtime was validated locally on:

AMD Ryzen 9 5950X
32 logical CPUs
Ubuntu Linux

The local experiment reached the full resident cap of:

16,777,216 particles

with deterministic scalar/parallel checksum parity.

A primary 8,388,608-resident / 32-worker result measured approximately:

Backend Scalar median Parallel median Measured speedup
Float 1.623 s 97.37 ms 16.67x
BAM-LUT 504.6 ms 84.56 ms 5.97x

The broader local matrix showed Float continuing to benefit through 32 workers while BAM-LUT plateaued substantially earlier at large resident sizes.

That observation motivated the wider qBraid experiment.

qBraid / Azure EPYC 7763 replication

The cloud replication was performed manually on a qBraid Nanoacademic Medium CPU instance.

Observed guest topology:

AMD EPYC 7763 64-Core Processor

with:

  • 96 online logical CPUs
  • 48 exposed cores
  • 2 threads per exposed core
  • 2 NUMA nodes
  • 192 MiB aggregate L3 reported by lscpu
  • approximately 377 GiB RAM
  • Microsoft full virtualization
  • Linux 6.8 Azure kernel

The guest explicitly exposed adjacent SMT sibling pairs:

0-1, 2-3, ... , 94-95

The exact merged PR #8 CPU runtime commit was verified before benchmarking.

All native CPU runtime tests passed, and verification confirmed:

  • available_parallelism=96
  • effective_workers=96
  • deterministic u64 boundary fingerprint
  • Float scalar/parallel checksum parity
  • BAM-LUT scalar/parallel checksum parity
  • sampled LUT diagnostic consistency

EPYC 8M worker-scaling result

For:

  • logical population = 18,446,744,073,709,551,615
  • resident particles = 8,388,608
  • frames = 8
  • repeats = 5
  • seed = 303

the unpinned worker sweep produced:

Workers Float parallel Float speedup BAM-LUT parallel BAM-LUT speedup
32 121.02 ms 20.30x 45.06 ms 16.65x
48 77.07 ms 31.71x 31.46 ms 24.10x
64 92.23 ms 26.65x 38.58 ms 19.53x
96 81.49 ms 29.92x 35.52 ms 21.22x

The best measured unpinned result for both backends occurred at 48 workers.

Performance was not monotonic beyond that point.

All worker counts preserved identical deterministic backend checksums.

Topology and affinity experiment

A repeat-matched 48-worker study was then used to test whether the 48-worker optimum corresponded simply to the guest's 48 exposed physical cores.

It did not.

Measured 48-worker results:

Configuration Float parallel BAM-LUT parallel
Unpinned 77.22 ms 31.38 ms
NUMA node 0 only 78.12 ms 33.01 ms
NUMA node 1 only 77.10 ms 34.33 ms
One SMT thread per exposed core across both NUMA nodes 115.54 ms 47.84 ms

The cross-NUMA one-thread-per-core configuration was roughly 50% slower than the unpinned result for both backends.

By contrast, restricting the process to one NUMA domain—24 exposed cores / 48 logical CPUs—retained performance close to the unrestricted baseline.

This result demonstrates strong topology sensitivity in this environment.

It is consistent with NUMA locality, first-touch placement, remote-memory traffic, cache effects, or memory-bandwidth interactions.

The available evidence does not isolate which of those mechanisms caused the difference because hardware memory-traffic counters were not collected.

Affinity provenance

The historical CPU runtime receipt schema records available and effective worker counts but does not encode Linux CPU affinity.

For the completed EPYC affinity experiment, v0.4.0 therefore preserves a separate affinity manifest binding each retained benchmark receipt to the exact taskset invocation used by the operator.

Future GALAXY affinity experiments are required to capture the kernel-visible allowed CPU list inside the constrained process before launching the benchmark.

This keeps the distinction clear between:

  • receipt-native evidence
  • operator command provenance
  • topology observations
  • inferred architectural interpretation

Evidence archive

The qBraid EPYC evidence is preserved under:

evidence/qbraid/EPYC7763-20260914/

including:

  • host topology capture
  • SMT sibling mapping
  • 32-worker receipt
  • 48-worker receipt
  • 64-worker receipt
  • 96-worker receipt
  • 7-repeat unpinned 48-worker confirmation
  • NUMA node 0 affinity receipt
  • NUMA node 1 affinity receipt
  • cross-NUMA one-thread-per-core receipt
  • affinity command manifest
  • evidence documentation

Original evidence archive SHA-256:

5c0474b9537a0ee34493a58c22b368f4846ad12f41633430d4444a678561e87a

GPU runtime

GALAXY continues to retain its native GPU runtime alongside the new CPU evidence.

Supported paths include:

  • Rust / wgpu / Vulkan
  • NVIDIA CUDA / CuPy RawKernel
  • memory-bounded u64 tiled execution
  • deterministic logical addressing beyond 2^32
  • hardware validation receipts
  • fixed-potential independent-particle dynamics

The CPU and GPU runtimes are separate implementations and should not be interpreted as direct end-to-end performance equivalents without a dedicated controlled comparison.

Scientific scope

GALAXY remains a deterministic galaxy dynamics and visualization instrument.

The native runtimes evolve independent test particles in prescribed gravitational potentials.

GALAXY does not currently implement:

  • pairwise stellar gravity
  • evolving self-gravity
  • hydrodynamics
  • gas evolution
  • star formation
  • self-consistent N-body dynamics
  • cross-particle force coupling

Large logical populations describe deterministic address spaces from which bounded resident populations are sampled or tiled.

A logical population of u64::MAX therefore does not mean that 18.4 quintillion particles are simultaneously allocated in memory.

Claim boundary

This release supports the following statement:

On the tested qBraid EPYC 7763 guest, GALAXY's 8,388,608-resident native CPU workload measured its best unpinned result at 48 workers in the tested 32/48/64/96-worker sweep for both Float and BAM-LUT, while preserving deterministic scalar/parallel checksums. A separate affinity experiment found a substantial penalty when one logical thread per exposed core was forced across both NUMA domains, while node-local 48-thread execution remained near the unpinned baseline.

This release does not claim that:

  • 96 cloud vCPUs are 96 physical cores
  • the unpinned 48-worker run occupied exactly 48 physical cores
  • BAM-LUT is proven memory-bandwidth bound
  • NUMA first-touch is proven to cause the observed affinity penalty
  • absolute EPYC and Ryzen performance are directly comparable
  • u64::MAX particles are simultaneously resident
  • GALAXY is a self-consistent N-body solver
  • the CPU runtime universally outperforms the GPU runtime
  • the sampled LUT diagnostic is a global mathematical error bound

Reproducibility milestone

v0.4.0 is intended to serve as the frozen baseline f...

Read more