Skip to content

Architectural Decisions

MasterLaplace edited this page Jul 14, 2026 · 4 revisions

Architectural Decisions

This page documents the key architectural decisions (ADRs) made during the development of LplPlugin, along with their rationale.


6.1 FlatAtomicsHashMap vs Alternatives (boost, folly, std)

Choice: Custom lock-free hash map with contiguous pool.

Reasons:

  • Chunks are heavy objects (~1KB+) β†’ we want them in a contiguous pool, not scattered on the heap
  • 64-bit atomic entries (42-bit key + 22-bit index) enable CAS on a single atomic<uint64_t>
  • get() is wait-free without mutex, crucial for reads during physics tick
  • No external dependency (boost forbidden, zero-dep philosophy)
  • forEach copies _activeSlots under lock β†’ safe iteration without blocking mutations

6.2 SpinLock vs std::mutex

Choice: std::atomic_flag + _mm_pause() instead of std::mutex.

Reasons:

  • Critical sections are ultra-short (~10 instructions): swap-and-pop, push_back
  • std::mutex makes a futex system call under contention β†’ unpredictable latency (10-100Β΅s)
  • SpinLock: constant ~100ns latency under short contention
  • One SpinLock per chunk β†’ distributed contention (no global hot lock)

6.3 SoA vs AoS for Component Storage

Choice: Structure of Arrays (SoA), not Array of Structures (AoS).

Reasons:

  • Physics iterates over positions then velocities separately β†’ cache-friendly linear access
  • CUDA: perfect memory coalescence (adjacent threads access adjacent addresses)
  • Double buffering applies naturally (one [2] array per component)
  • AoS wastes cache bandwidth when reading only 1 out of N components

6.4 Morton + Bias vs Signed Coordinates

Choice: Morton keys on biased coordinates (adding an offset of 2²⁰).

Reasons:

  • Chunk coordinates are signed (world centered on origin)
  • Morton only works on positive integers
  • The 2²⁰ bias covers [-1M, +1M] while preserving spatial locality
  • Alternative (signed coords + hash): loses locality β†’ Morton prefix unusable

6.5 Double Buffering vs Locking

Choice: Two separate buffers for positions, velocities, forces instead of a single mutex-protected buffer.

Reasons:

  • Network reads the read buffer while physics writes to the write buffer β†’ zero contention
  • No lock during state broadcast (composed of thousands of entities)
  • Swap is a simple atomic index toggle + async memcpy
  • Memory cost doubled only for hot data (cold data not doubled)

6.6 Backward Iteration for Swap-and-Pop

Choice: Inter-chunk migration via backward iteration (i = count-1..0).

Reasons:

  • Swap-and-pop moves the last element to the deleted position
  • Forward iteration: swapping element 5 with the last (e.g. 99) skips the element that just arrived at position 5
  • Backward iteration: the swapped element is at an already processed index β†’ no skip
  • Bug fixed in conversation: original code had forward iteration β†’ double-tick bug

6.7 EntityRef (Mutable Proxy) vs Copies

Choice: getEntity() returns an EntityRef containing references directly to the SoA, not copies.

Reasons:

  • Zero copy: modifying the EntityRef directly modifies the arrays
  • Natural API: entity.position = newPos; instead of chunk.setPosition(idx, newPos);
  • Combined access: one call instead of N calls to getPosition(), getVelocity(), etc.

6.8 PinnedAllocator (Zero-Copy GPU) vs cudaMemcpy

Choice: Mapped pinned memory instead of explicit CPU→GPU→CPU copies.

Reasons:

  • cudaMemcpy: 2 copies per frame (CPUβ†’GPU before kernel, GPUβ†’CPU after) β†’ latency + bandwidth
  • Pinned mapped: GPU accesses CPU RAM directly via PCIe β†’ single implicit copy
  • For small entity counts per chunk (<10k), PCIe overhead is negligible
  • The malloc fallback for compilation without CUDA is transparent via #ifdef __CUDACC__

6.9 Per-Chunk ECS (Not Global)

Choice: Each chunk owns its own SoA arrays, not a centralized global ECS (like EnTT).

Reasons:

  • Inter-chunk migration = simple local swap-and-pop + insertion in the new chunk
  • No global fragmentation: each chunk is compact
  • Physics tick is embarrassingly parallel (one chunk = one independent work unit)
  • Future horizontal scaling (Server Meshing): one chunk = one transferable unit to another server

6.10 Kernel Capture (NF_DROP) vs Classic UDP Socket

Choice: Netfilter kernel module that drops packets + mmap ring buffer instead of recvfrom().

Reasons:

  • recvfrom(): 2 copies (kernel β†’ socket buffer β†’ userspace) + syscall overhead (~1Β΅s/call)
  • Kernel module: 1 copy (SKB β†’ ring buffer) + mmap zero-copy to userspace
  • No socket overhead (file descriptors, select/epoll, buffer management)
  • Measured latency: 62.55Β΅s from network packet to ring buffer
  • Measured throughput: 495 packets/s without loss (0% loss over 200k packets)

6.11 Adaptive Sparse Set for Partition::_idToLocal

Choice: Replace std::unordered_map<uint32_t, uint32_t> with a sparse set with lazy allocation and adaptive sizing.

Reasons:

  • unordered_map: repeated dynamic allocations, cache misses, unpredictable hash collisions
  • Naive sparse set (1M): allocates 4MB per chunk β†’ with 100+ chunks = system hang
  • Adaptive sparse set: allocates only on first insertion, sized at min(entityId Γ— 1.5, 1MB)
  • Reduces typical per-chunk memory: 64KB-256KB vs 4MB
  • O(1) guaranteed (no collisions) vs O(1) amortized + hash overhead for unordered_map

Implementation:

// Partition::addEntity() β€” lazy allocation
if (entity.id >= _sparseCapacity) {
    uint32_t newCapacity = entity.id + (entity.id >> 1); // 1.5x growth
    newCapacity = std::min(newCapacity, 262144u);         // Cap at 1MB
    newCapacity = 1u << (32u - __builtin_clz(newCapacity - 1)); // Align power of 2
    _sparseToLocal.resize(newCapacity, INVALID_INDEX);
    _sparseCapacity = newCapacity;
}

Benchmark (sparse set vs unordered_map):

  • Partition::findEntityIndex: ~71 ns (vs 65 ns before) β†’ acceptable trade-off to avoid crashes
  • Global throughput: 38-55 M ops/sec (identical)
  • Memory: 64-256KB/chunk (vs 4MB theoretical with 1M Γ— 4 bytes)

6.12 Empty Chunk Garbage Collection

Choice: Inline GC every frame in WorldPartition::step(), deleting chunks with 0 entities via FlatAtomicsHashMap::remove().

Problem solved:

  • Migrating entities leave a chunk via checkAndMigrate() and are re-inserted into a new chunk via getOrCreateChunk()
  • The old (potentially empty) chunk was never deleted from the hash map
  • forEach() iterated all chunks including empty ones β†’ physicsTick() acquired a SpinLock + rebuilt octree for 0 entities
  • Result: monotonic chunk growth (17 β†’ 36+ in 30s) and linear physics step degradation (2ms β†’ 9ms+)

Solution:

// WorldPartition::step() β€” after re-inserting migrants
std::vector<uint64_t> emptyKeys;
_partitions.forEach([&](Partition &p) {
    if (p.getEntityCount() == 0 && p.getMortonKey() != 0)
        emptyKeys.push_back(p.getMortonKey());
});
for (uint64_t key : emptyKeys)
    _partitions.remove(key);  // tombstone + slot recycling

Complementary optimizations:

  • Early return in physicsTick() and checkAndMigrate(): if (_ids.empty()) return; after SpinLock
  • Morton key tracking: _mortonKey member of Partition, transferred via move semantics

Results:

Metric Before GC After GC
Chunks (30s, 50 NPCs) 17 β†’ 36+ (monotonic) 17 β†’ 24-28 (oscillating)
Physics step 2ms β†’ 9ms+ (unbounded) 2ms β†’ 3.8ms (capped)
Chunks deleted Never Yes (descends: 28 β†’ 24)

Note: This GC removes empty chunks (0 entities) immediately. Future TTL will keep inactive chunks (with persistent data but no players) in RAM for 5-30s before serialization to disk.


6.13 Physics Step Optimizations

Problem: With 50 NPCs, physics step grew indefinitely (2.7ms β†’ 5.8ms in 60s). Root cause: radix sort in FlatDynamicOctree::rebuild() allocates 4 Γ— uint32_t counts[65536] = {0} on the stack = 1MB of zero-init per chunk per frame. With infinite NPC dispersion, chunk count grew to 42+, i.e. ~42MB memset/frame.

5 optimizations implemented:

# Optimization Impact
1 Octree bypass for chunks ≀ 32 entities β†’ brute-force NΒ² Eliminates radix sort (1MB/chunk) for small chunks
2 Chunk size 255 β†’ 1000 Reduces chunk count from 42 β†’ 4
3 Linear friction (damping 0.995/frame) Prevents infinite NPC dispersion
4 Sleeping entities (|v|Β² < 0.01 for 30 frames) Skip physics + migration for stationary entities
5 ThreadPool CPU (batch dispatch to N workers) Parallelizes physicsTick across active chunks

Before / After (20 seconds, 50 NPCs):

Metric Before After Improvement
Physics Step 2.7 β†’ 5.8 ms (growing) 0.052 – 0.064 ms (stable) ~50-100Γ—
Active chunks 17 β†’ 42+ (growing) 4 (stable) No more dispersion
Transit/frame 3-8 (constant) 0 (stable) Sleeping + friction
Stability Linear degradation Constant plateau βœ…

Tuning constants:

VELOCITY_DAMPING            = 0.995f     // ~26% loss/sec at 60Hz
SLEEP_VELOCITY_SQ_THRESHOLD = 0.01f      // threshold 0.1 m/s
SLEEP_FRAMES_THRESHOLD      = 30         // ~0.5s at 60Hz
BRUTE_FORCE_THRESHOLD       = 32         // = octree leafCapacity

6.14 AABB Collision β€” Iterative Solver

Algorithm (inspired by Jolt Physics):

  1. AABB Detection: overlap on 3 axes (simplified Separating Axis Theorem)
  2. Minimum penetration axis: chooses the axis with smallest interpenetration β†’ collision normal
  3. Positional separation: 100% correction proportional to inverse mass (no Baumgarte/slop)
  4. Newton impulse: j = -(1+e) Γ— v_relΒ·n / (1/mA + 1/mB), restitution coefficient e=0.5
  5. Auto wake: any collision wakes involved entities (_sleeping[i] = false)
  6. Iterative solver: 4 complete passes (SOLVER_ITERATIONS = 4) for multi-body convergence
  7. Adaptive broadphase: brute-force NΒ² for ≀32 entities, octree rebuild() + query() beyond

Key points:

  • Sleeping entities participate in detection (only "two sleeping" case is skipped)
  • Deterministic: no insertion-order dependency, pair (i,j) processed only for i<j
  • Separate ground collision (Pass 1) with bounce y *= -RESTITUTION
  • Performance: ~0.055ms for 50 entities (4 chunks, 4 solver iterations)

Constants:

Constant Value Purpose
RESTITUTION 0.5 Coefficient of restitution
SOLVER_ITERATIONS 4 Constraint solver passes
BRUTE_FORCE_THRESHOLD 32 Brute-force vs octree threshold

6.15 Sequential vs ThreadPool Parallelization

Choice: Global ThreadPool instead of std::for_each(par, ...) or manual thread management.

Reasons:

  • Thread reuse: Avoids the cost of creating/destroying threads every frame (very expensive at 60Hz)
  • Granularity control: Allows dispatching per chunk or per system
  • Automatic load balancing: Workers consume from a shared queue; if one chunk is empty (fast) and another is full (slow), workers naturally distribute the load
  • Unified abstraction: Used by both SystemScheduler and WorldPartition

6.16 Unified Network Class vs NetworkDispatch + Raw Sockets

Choice: Single header-only Network.hpp class replacing the old NetworkDispatch.cpp/.hpp + scattered raw socket code in main.cpp and visual.cpp.

Reasons:

  • Code duplication eliminated: server and client shared the same protocol logic (packet parsing, MSG_* dispatch) but had separate implementations
  • Dual-mode transparency: #ifdef LPL_USE_SOCKET switches between kernel driver and socket fallback at compile time β€” no runtime branching
  • Header-only: no separate .cpp β€” simplifies the build system (removed NetworkDispatch.o from Makefile)
  • Client management centralized: ClientEndpoint tracking, broadcast_state(), and handle_connect() all live in one place
  • Socket fallback for development: allows testing without loading the kernel module (useful for CI, WSL, containers)

6.17 Bidirectional Kernel Module (TX via Kernel Thread)

Choice: Add a TX path to the kernel module instead of using sendto() from userspace.

Reasons:

  • sendto() from userspace: syscall overhead (~1Β΅s/call) + socket buffer management for each packet
  • Kernel TX thread (lpl_tx_worker): sleeps on a wait queue, woken by a single ioctl(LPL_IOC_KICK_TX) call
  • Batch processing: one ioctl wakes the thread which then drains the entire TX ring β€” N packets for the cost of 1 syscall
  • kernel_sendmsg() bypasses the userspace socket layer entirely
  • Symmetry: both RX and TX now go through shared memory ring buffers β€” truly bidirectional zero-copy pipeline

6.18 Flat Module Architecture vs Monolithic engine/

Choice: Decompose the monolithic engine/ + shared/ + plugins/ into independent root-level modules (16 at the time of the refactor, 20 today), each with its own xmake.lua, include/lpl/<name>/, and src/.

Reasons:

  • Monolithic engine/ grew to 19+ header files with tangled dependencies β†’ build times and comprehension suffered
  • Flat modules enforce explicit dependency declarations (each xmake.lua lists only what it needs)
  • Independent compilation: changing lpl-math doesn't recompile lpl-net if the interface is stable
  • SOLID alignment: each module has a single responsibility (container, concurrency, ecs, physics, etc.)
  • Facilitates future selective deployment (server needs lpl-net + lpl-ecs, not lpl-render + lpl-bci)

Migration strategy: Move all legacy code to _legacy/ (preserving git history), then rebuild each module from scratch applying SOLID, Doxygen, and .inl conventions.


6.19 xmake over Make

Choice: Replace the legacy Makefile with xmake as the sole build system.

Reasons:

  • Make required manual -I flag management for each module β†’ error-prone at 20 modules
  • xmake handles dependency resolution between targets natively (add_deps)
  • Built-in package management (add_requires) for Vulkan SDK, ImGui, etc.
  • Cross-compilation support (Android NDK) without custom toolchain files
  • Build modes (debug, release, profile) with one command
  • Lua scripting for custom targets (future: kernel module install)

Trade-off: Lost make install/make uninstall for kernel module (regression #3). Custom xmake targets planned.


6.20 C++23 with -fno-rtti and -fno-exceptions

Choice: Upgrade from C++20 to C++23 with RTTI and exceptions disabled globally.

Reasons:

  • -fno-rtti: eliminates vtable overhead and typeid β€” aligns with data-oriented design (no virtual dispatch in hot paths)
  • -fno-exceptions: guarantees deterministic control flow β€” no hidden try/catch paths, smaller binary, better inlining
  • VULKAN_HPP_NO_EXCEPTIONS: Vulkan C++ bindings return vk::Result instead of throwing
  • C++23 features used: std::expected (replaces exceptions for error handling), consteval, improved constexpr, structured bindings
  • -Werror -Wall -Wextra: zero-tolerance for warnings enforces code quality

6.21 VkWrapper Integration (Vulkan Renderer)

Choice: Port the personal VkWrapper project into the render/ module instead of continuing with OpenGL 2.1.

Reasons:

  • OpenGL 2.1 was a debug placeholder β€” no shader pipeline, no compute, no modern GPU features
  • VkWrapper already had a working Vulkan pipeline (device, swapchain, graphics pipeline) following vulkan-tutorial.com
  • Vulkan provides compute shaders as a future CUDA alternative (portability)
  • ImGui integration (GLFW + Vulkan backend) for debug overlays

Adaptations:

  • Removed all EngineSquared dependencies (entt, spdlog) β€” replaced by lpl-core and lpl-ecs
  • Removed EngineSquared plugin/system abstractions β€” adapted to lpl-engine SystemScheduler
  • Internal headers (src/vk/) kept private to the render module

6.22 Software Rasterizer alongside Vulkan (IRenderer parity)

Choice: Put both a dependency-free CPU SoftwareRasterizer and the VulkanRenderer behind one IRenderer interface, and check they produce the same image via RenderParity.

Reasons:

  • A frame can be produced with no GPU at all (CI, headless capture, the freestanding target where Vulkan may be absent).
  • Parity turns the software path into an oracle: any divergence in the Vulkan path is a bug, caught by comparison.
  • The software rasterizer's Topology uses only +,-,*,/ and a hardware sqrt, so it stays deterministic and portable.

Cost: two backends to maintain; parity testing to keep honest.


6.23 Platform Backend Abstraction (lpl-platform)

Choice: Route host services (clock, display, GPU memory, input) through backend interfaces (IClockBackend, IDisplayBackend, IGpuMemoryBackend, IInputBackend) with concrete linux/ and kernel/ implementations.

Reason: the engine targets both userspace Linux and a freestanding kernel host. Swapping host means swapping a backend, not editing the engine. This is what keeps the hosted and freestanding builds from forking.


6.24 Serialization & Deterministic Replay (lpl-serial)

Choice: Repurpose the serial/ module from a USB serial-port wrapper (that concern moved to bci/source/serial/) into state serialization and deterministic replay: ISerializable, StateSnapshot, ReplayRecorder, ReplayPlayer.

Reason: the simulation is already deterministic (fixed-point maths, fixed tick), so a recorded input + snapshot stream reproduces a run bit-for-bit. That gives debugging, regression tests and the rollback netcode a common substrate instead of three ad-hoc mechanisms.


6.25 RCU Sessions & Rollback Netcode (lpl-net)

Choice: Move SessionManager toward Read-Copy-Update semantics and add a snapshot-ring RollbackStrategy in net/netcode/.

Reasons:

  • Broadcast reads the session set every tick while connections mutate it β€” RCU lets readers proceed without locking out writers.
  • Rollback needs a circular buffer of serialized snapshots (built on lpl-serial) to rewind and re-simulate on late/mispredicted input.

Status: stubs in place; not yet the default path.


6.26 Platform-Scoped Packaging over Source Patching (BrainFlow / liblsl)

Choice: To get brainflow and liblsl into the official xmake-repo, scope each recipe to the platforms it genuinely supports (build from source on Linux/macOS, ship BrainFlow's official prebuilt binaries on Windows x64/x86, skip unsupported combos) instead of patching upstream sources to force-compile everywhere.

Reasons:

  • Inline source rewrites (ATL removal, std::tr1 fixes) were unmergeable and fragile.
  • The test framework skips a package on platforms with no matching on_install, so scoping is clean rather than a failure.
  • -DLSL_OPTIMIZATIONS=OFF drops liblsl's LTO, which otherwise emits clang bitcode that GNU ld can't archive.

Outcome: merged upstream into xmake-io/xmake-repo; bci/ now resolves its deps from the official repo.


← Modules | Next: Performance Results β†’

πŸ“– LplPlugin Wiki

🏠 Home


πŸ“ Project Overview

πŸ”§ Technical Reference

πŸ“Š Status & Results

πŸ—ΊοΈ Planning

πŸŽ“ Research

Clone this wiki locally