Repository navigation
Architectural Decisions
This page documents the key architectural decisions (ADRs) made during the development of LplPlugin, along with their rationale.
Choice: Custom lock-free hash map with contiguous pool.
Reasons:
- Chunks are heavy objects (~1KB+) β we want them in a contiguous pool, not scattered on the heap
- 64-bit atomic entries (42-bit key + 22-bit index) enable CAS on a single
atomic<uint64_t> -
get()is wait-free without mutex, crucial for reads during physics tick - No external dependency (boost forbidden, zero-dep philosophy)
-
forEachcopies_activeSlotsunder lock β safe iteration without blocking mutations
Choice: std::atomic_flag + _mm_pause() instead of std::mutex.
Reasons:
- Critical sections are ultra-short (~10 instructions): swap-and-pop, push_back
-
std::mutexmakes a futex system call under contention β unpredictable latency (10-100Β΅s) - SpinLock: constant ~100ns latency under short contention
- One SpinLock per chunk β distributed contention (no global hot lock)
Choice: Structure of Arrays (SoA), not Array of Structures (AoS).
Reasons:
- Physics iterates over
positionsthenvelocitiesseparately β cache-friendly linear access - CUDA: perfect memory coalescence (adjacent threads access adjacent addresses)
- Double buffering applies naturally (one
[2]array per component) - AoS wastes cache bandwidth when reading only 1 out of N components
Choice: Morton keys on biased coordinates (adding an offset of 2Β²β°).
Reasons:
- Chunk coordinates are signed (world centered on origin)
- Morton only works on positive integers
- The 2Β²β° bias covers [-1M, +1M] while preserving spatial locality
- Alternative (signed coords + hash): loses locality β Morton prefix unusable
Choice: Two separate buffers for positions, velocities, forces instead of a single mutex-protected buffer.
Reasons:
- Network reads the read buffer while physics writes to the write buffer β zero contention
- No lock during state broadcast (composed of thousands of entities)
- Swap is a simple atomic index toggle + async memcpy
- Memory cost doubled only for hot data (cold data not doubled)
Choice: Inter-chunk migration via backward iteration (i = count-1..0).
Reasons:
- Swap-and-pop moves the last element to the deleted position
- Forward iteration: swapping element 5 with the last (e.g. 99) skips the element that just arrived at position 5
- Backward iteration: the swapped element is at an already processed index β no skip
- Bug fixed in conversation: original code had forward iteration β double-tick bug
Choice: getEntity() returns an EntityRef containing references directly to the SoA, not copies.
Reasons:
- Zero copy: modifying the EntityRef directly modifies the arrays
- Natural API:
entity.position = newPos;instead ofchunk.setPosition(idx, newPos); - Combined access: one call instead of N calls to
getPosition(),getVelocity(), etc.
Choice: Mapped pinned memory instead of explicit CPUβGPUβCPU copies.
Reasons:
-
cudaMemcpy: 2 copies per frame (CPUβGPU before kernel, GPUβCPU after) β latency + bandwidth - Pinned mapped: GPU accesses CPU RAM directly via PCIe β single implicit copy
- For small entity counts per chunk (<10k), PCIe overhead is negligible
- The
mallocfallback for compilation without CUDA is transparent via#ifdef __CUDACC__
Choice: Each chunk owns its own SoA arrays, not a centralized global ECS (like EnTT).
Reasons:
- Inter-chunk migration = simple local swap-and-pop + insertion in the new chunk
- No global fragmentation: each chunk is compact
- Physics tick is embarrassingly parallel (one chunk = one independent work unit)
- Future horizontal scaling (Server Meshing): one chunk = one transferable unit to another server
Choice: Netfilter kernel module that drops packets + mmap ring buffer instead of recvfrom().
Reasons:
-
recvfrom(): 2 copies (kernel β socket buffer β userspace) + syscall overhead (~1Β΅s/call) - Kernel module: 1 copy (SKB β ring buffer) + mmap zero-copy to userspace
- No socket overhead (file descriptors, select/epoll, buffer management)
- Measured latency: 62.55Β΅s from network packet to ring buffer
- Measured throughput: 495 packets/s without loss (0% loss over 200k packets)
Choice: Replace std::unordered_map<uint32_t, uint32_t> with a sparse set with lazy allocation and adaptive sizing.
Reasons:
-
unordered_map: repeated dynamic allocations, cache misses, unpredictable hash collisions - Naive sparse set (1M): allocates 4MB per chunk β with 100+ chunks = system hang
- Adaptive sparse set: allocates only on first insertion, sized at min(entityId Γ 1.5, 1MB)
- Reduces typical per-chunk memory: 64KB-256KB vs 4MB
- O(1) guaranteed (no collisions) vs O(1) amortized + hash overhead for unordered_map
Implementation:
// Partition::addEntity() β lazy allocation
if (entity.id >= _sparseCapacity) {
uint32_t newCapacity = entity.id + (entity.id >> 1); // 1.5x growth
newCapacity = std::min(newCapacity, 262144u); // Cap at 1MB
newCapacity = 1u << (32u - __builtin_clz(newCapacity - 1)); // Align power of 2
_sparseToLocal.resize(newCapacity, INVALID_INDEX);
_sparseCapacity = newCapacity;
}Benchmark (sparse set vs unordered_map):
-
Partition::findEntityIndex: ~71 ns (vs 65 ns before) β acceptable trade-off to avoid crashes - Global throughput: 38-55 M ops/sec (identical)
- Memory: 64-256KB/chunk (vs 4MB theoretical with 1M Γ 4 bytes)
Choice: Inline GC every frame in WorldPartition::step(), deleting chunks with 0 entities via FlatAtomicsHashMap::remove().
Problem solved:
- Migrating entities leave a chunk via
checkAndMigrate()and are re-inserted into a new chunk viagetOrCreateChunk() - The old (potentially empty) chunk was never deleted from the hash map
-
forEach()iterated all chunks including empty ones βphysicsTick()acquired a SpinLock + rebuilt octree for 0 entities - Result: monotonic chunk growth (17 β 36+ in 30s) and linear physics step degradation (2ms β 9ms+)
Solution:
// WorldPartition::step() β after re-inserting migrants
std::vector<uint64_t> emptyKeys;
_partitions.forEach([&](Partition &p) {
if (p.getEntityCount() == 0 && p.getMortonKey() != 0)
emptyKeys.push_back(p.getMortonKey());
});
for (uint64_t key : emptyKeys)
_partitions.remove(key); // tombstone + slot recyclingComplementary optimizations:
-
Early return in
physicsTick()andcheckAndMigrate():if (_ids.empty()) return;after SpinLock -
Morton key tracking:
_mortonKeymember of Partition, transferred via move semantics
Results:
| Metric | Before GC | After GC |
|---|---|---|
| Chunks (30s, 50 NPCs) | 17 β 36+ (monotonic) | 17 β 24-28 (oscillating) |
| Physics step | 2ms β 9ms+ (unbounded) | 2ms β 3.8ms (capped) |
| Chunks deleted | Never | Yes (descends: 28 β 24) |
Note: This GC removes empty chunks (0 entities) immediately. Future TTL will keep inactive chunks (with persistent data but no players) in RAM for 5-30s before serialization to disk.
Problem: With 50 NPCs, physics step grew indefinitely (2.7ms β 5.8ms in 60s). Root cause: radix sort in FlatDynamicOctree::rebuild() allocates 4 Γ uint32_t counts[65536] = {0} on the stack = 1MB of zero-init per chunk per frame. With infinite NPC dispersion, chunk count grew to 42+, i.e. ~42MB memset/frame.
5 optimizations implemented:
| # | Optimization | Impact |
|---|---|---|
| 1 | Octree bypass for chunks β€ 32 entities β brute-force NΒ² | Eliminates radix sort (1MB/chunk) for small chunks |
| 2 | Chunk size 255 β 1000 | Reduces chunk count from 42 β 4 |
| 3 | Linear friction (damping 0.995/frame) | Prevents infinite NPC dispersion |
| 4 | Sleeping entities (|v|Β² < 0.01 for 30 frames) | Skip physics + migration for stationary entities |
| 5 | ThreadPool CPU (batch dispatch to N workers) | Parallelizes physicsTick across active chunks |
Before / After (20 seconds, 50 NPCs):
| Metric | Before | After | Improvement |
|---|---|---|---|
| Physics Step | 2.7 β 5.8 ms (growing) | 0.052 β 0.064 ms (stable) | ~50-100Γ |
| Active chunks | 17 β 42+ (growing) | 4 (stable) | No more dispersion |
| Transit/frame | 3-8 (constant) | 0 (stable) | Sleeping + friction |
| Stability | Linear degradation | Constant plateau | β |
Tuning constants:
VELOCITY_DAMPING = 0.995f // ~26% loss/sec at 60Hz
SLEEP_VELOCITY_SQ_THRESHOLD = 0.01f // threshold 0.1 m/s
SLEEP_FRAMES_THRESHOLD = 30 // ~0.5s at 60Hz
BRUTE_FORCE_THRESHOLD = 32 // = octree leafCapacity
Algorithm (inspired by Jolt Physics):
- AABB Detection: overlap on 3 axes (simplified Separating Axis Theorem)
- Minimum penetration axis: chooses the axis with smallest interpenetration β collision normal
- Positional separation: 100% correction proportional to inverse mass (no Baumgarte/slop)
-
Newton impulse:
j = -(1+e) Γ v_relΒ·n / (1/mA + 1/mB), restitution coefficiente=0.5 -
Auto wake: any collision wakes involved entities (
_sleeping[i] = false) -
Iterative solver: 4 complete passes (
SOLVER_ITERATIONS = 4) for multi-body convergence -
Adaptive broadphase: brute-force NΒ² for β€32 entities, octree
rebuild()+query()beyond
Key points:
- Sleeping entities participate in detection (only "two sleeping" case is skipped)
- Deterministic: no insertion-order dependency, pair (i,j) processed only for i<j
- Separate ground collision (Pass 1) with bounce
y *= -RESTITUTION - Performance: ~0.055ms for 50 entities (4 chunks, 4 solver iterations)
Constants:
| Constant | Value | Purpose |
|---|---|---|
RESTITUTION |
0.5 | Coefficient of restitution |
SOLVER_ITERATIONS |
4 | Constraint solver passes |
BRUTE_FORCE_THRESHOLD |
32 | Brute-force vs octree threshold |
Choice: Global ThreadPool instead of std::for_each(par, ...) or manual thread management.
Reasons:
- Thread reuse: Avoids the cost of creating/destroying threads every frame (very expensive at 60Hz)
- Granularity control: Allows dispatching per chunk or per system
- Automatic load balancing: Workers consume from a shared queue; if one chunk is empty (fast) and another is full (slow), workers naturally distribute the load
-
Unified abstraction: Used by both
SystemSchedulerandWorldPartition
Choice: Single header-only Network.hpp class replacing the old NetworkDispatch.cpp/.hpp + scattered raw socket code in main.cpp and visual.cpp.
Reasons:
- Code duplication eliminated: server and client shared the same protocol logic (packet parsing, MSG_* dispatch) but had separate implementations
-
Dual-mode transparency:
#ifdef LPL_USE_SOCKETswitches between kernel driver and socket fallback at compile time β no runtime branching -
Header-only: no separate
.cppβ simplifies the build system (removedNetworkDispatch.ofrom Makefile) -
Client management centralized:
ClientEndpointtracking,broadcast_state(), andhandle_connect()all live in one place - Socket fallback for development: allows testing without loading the kernel module (useful for CI, WSL, containers)
Choice: Add a TX path to the kernel module instead of using sendto() from userspace.
Reasons:
-
sendto()from userspace: syscall overhead (~1Β΅s/call) + socket buffer management for each packet - Kernel TX thread (
lpl_tx_worker): sleeps on a wait queue, woken by a singleioctl(LPL_IOC_KICK_TX)call - Batch processing: one ioctl wakes the thread which then drains the entire TX ring β N packets for the cost of 1 syscall
-
kernel_sendmsg()bypasses the userspace socket layer entirely - Symmetry: both RX and TX now go through shared memory ring buffers β truly bidirectional zero-copy pipeline
Choice: Decompose the monolithic engine/ + shared/ + plugins/ into independent root-level modules (16 at the time of the refactor, 20 today), each with its own xmake.lua, include/lpl/<name>/, and src/.
Reasons:
- Monolithic
engine/grew to 19+ header files with tangled dependencies β build times and comprehension suffered - Flat modules enforce explicit dependency declarations (each
xmake.lualists only what it needs) - Independent compilation: changing
lpl-mathdoesn't recompilelpl-netif the interface is stable - SOLID alignment: each module has a single responsibility (container, concurrency, ecs, physics, etc.)
- Facilitates future selective deployment (server needs
lpl-net+lpl-ecs, notlpl-render+lpl-bci)
Migration strategy: Move all legacy code to _legacy/ (preserving git history), then rebuild each module from scratch applying SOLID, Doxygen, and .inl conventions.
Choice: Replace the legacy Makefile with xmake as the sole build system.
Reasons:
- Make required manual
-Iflag management for each module β error-prone at 20 modules - xmake handles dependency resolution between targets natively (
add_deps) - Built-in package management (
add_requires) for Vulkan SDK, ImGui, etc. - Cross-compilation support (Android NDK) without custom toolchain files
- Build modes (
debug,release,profile) with one command - Lua scripting for custom targets (future: kernel module install)
Trade-off: Lost make install/make uninstall for kernel module (regression #3). Custom xmake targets planned.
Choice: Upgrade from C++20 to C++23 with RTTI and exceptions disabled globally.
Reasons:
-
-fno-rtti: eliminates vtable overhead andtypeidβ aligns with data-oriented design (no virtual dispatch in hot paths) -
-fno-exceptions: guarantees deterministic control flow β no hiddentry/catchpaths, smaller binary, better inlining -
VULKAN_HPP_NO_EXCEPTIONS: Vulkan C++ bindings returnvk::Resultinstead of throwing - C++23 features used:
std::expected(replaces exceptions for error handling),consteval, improvedconstexpr, structured bindings -
-Werror -Wall -Wextra: zero-tolerance for warnings enforces code quality
Choice: Port the personal VkWrapper project into the render/ module instead of continuing with OpenGL 2.1.
Reasons:
- OpenGL 2.1 was a debug placeholder β no shader pipeline, no compute, no modern GPU features
- VkWrapper already had a working Vulkan pipeline (device, swapchain, graphics pipeline) following vulkan-tutorial.com
- Vulkan provides compute shaders as a future CUDA alternative (portability)
- ImGui integration (GLFW + Vulkan backend) for debug overlays
Adaptations:
- Removed all EngineSquared dependencies (entt, spdlog) β replaced by
lpl-coreandlpl-ecs - Removed EngineSquared plugin/system abstractions β adapted to
lpl-engineSystemScheduler - Internal headers (
src/vk/) kept private to the render module
Choice: Put both a dependency-free CPU SoftwareRasterizer and the VulkanRenderer behind one IRenderer interface, and check they produce the same image via RenderParity.
Reasons:
- A frame can be produced with no GPU at all (CI, headless capture, the freestanding target where Vulkan may be absent).
- Parity turns the software path into an oracle: any divergence in the Vulkan path is a bug, caught by comparison.
- The software rasterizer's
Topologyuses only+,-,*,/and a hardwaresqrt, so it stays deterministic and portable.
Cost: two backends to maintain; parity testing to keep honest.
Choice: Route host services (clock, display, GPU memory, input) through backend interfaces (IClockBackend, IDisplayBackend, IGpuMemoryBackend, IInputBackend) with concrete linux/ and kernel/ implementations.
Reason: the engine targets both userspace Linux and a freestanding kernel host. Swapping host means swapping a backend, not editing the engine. This is what keeps the hosted and freestanding builds from forking.
Choice: Repurpose the serial/ module from a USB serial-port wrapper (that concern moved to bci/source/serial/) into state serialization and deterministic replay: ISerializable, StateSnapshot, ReplayRecorder, ReplayPlayer.
Reason: the simulation is already deterministic (fixed-point maths, fixed tick), so a recorded input + snapshot stream reproduces a run bit-for-bit. That gives debugging, regression tests and the rollback netcode a common substrate instead of three ad-hoc mechanisms.
Choice: Move SessionManager toward Read-Copy-Update semantics and add a snapshot-ring RollbackStrategy in net/netcode/.
Reasons:
- Broadcast reads the session set every tick while connections mutate it β RCU lets readers proceed without locking out writers.
- Rollback needs a circular buffer of serialized snapshots (built on
lpl-serial) to rewind and re-simulate on late/mispredicted input.
Status: stubs in place; not yet the default path.
Choice: To get brainflow and liblsl into the official xmake-repo, scope each recipe to the platforms it genuinely supports (build from source on Linux/macOS, ship BrainFlow's official prebuilt binaries on Windows x64/x86, skip unsupported combos) instead of patching upstream sources to force-compile everywhere.
Reasons:
- Inline source rewrites (ATL removal,
std::tr1fixes) were unmergeable and fragile. - The test framework skips a package on platforms with no matching
on_install, so scoping is clean rather than a failure. -
-DLSL_OPTIMIZATIONS=OFFdrops liblsl's LTO, which otherwise emits clang bitcode that GNUldcan't archive.
Outcome: merged upstream into xmake-io/xmake-repo; bci/ now resolves its deps from the official repo.
β Modules | Next: Performance Results β
LplPlugin β Zero-Copy Real-Time VR Engine | Author: MasterLaplace | License: GPL-3.0 | Source Code
This project is a marathon, not a sprint.