Skip to content

Performance Results

MasterLaplace edited this page Jul 14, 2026 · 2 revisions

Performance Results

Network Metrics (Phase 1 — Kernel Ring Buffer)

Metric Value
Packets sent 1,000
Packets received 1,000
Packets lost 0 (0.00%)
Throughput ~495 pkt/s

Frame Metrics

Metric Value
Average frame time 62.55 µs
Min frame time 41.93 µs
Max frame time 241.80 µs
Variance ~10 µs
Framerate ~59.49 FPS (locked by usleep)

Analysis

  • 60 FPS target (16,666 µs/frame): LARGELY EXCEEDED — physics + network consume only 0.38% of the time budget
  • Theoretical potential: ~14,000 FPS if framerate cap is removed
  • Stability: variance <15%, no critical spikes
  • Zero-copy validated: no visible degradation compared to classic memcpy

CPU Benchmark (benchmark.cpp) — historical (Feb 2026)

Historical run from the original benchmark app, kept for comparison. The suite has since been rewritten — see Current Benchmark Suite below for the up-to-date numbers.

Configuration: 10,000 entities, 100 frames, random spatial distribution
Command: make benchmark && ./benchmark [numEntities] [numFrames]

Results — Adaptive Sparse Set (February 13, 2026)

Metric 10k entities 50k entities
Simulation Time 0.026 sec 0.180 sec
Throughput 38.6 M ops/sec 55.6 M ops/sec
Avg FPS 3,915.9 fps 1,117.8 fps
EntityRegistry Lookup 0.059 µs 0.073 µs
Partition::findEntityIndex 70.7 ns 71.7 ns
Active Chunks Created 1,598 1,700
Entities per Chunk (avg) 6.3 29.4
Creation Overhead 0.102 sec 0.677 sec

Lookup Architecture Analysis

Major change (February 13, 2026): Replacement of std::unordered_map<uint32_t, uint32_t> with adaptive sparse set in Partition::_idToLocal.

Reason: Avoid massive 4MB allocation per chunk (1M × 4 bytes) that caused system hangs with 100+ simultaneous chunks.

Solution implemented: Lazy allocation with adaptive sizing:

  • Allocates the sparse set only on first entity insertion
  • Sizes to min(entity.id × 1.5, 1MB) instead of fixed 1M
  • Reduces typical per-chunk memory to 64KB-256KB

Measured performance:

  • O(1) lookup: 70.7 ns (vs 65.8 ns before) — negligible cost
  • No repeated dynamic allocations (unlike unordered_map)
  • Cache-friendly: sequential sparse set access
  • Predictability: guaranteed O(1) (no hash collisions)

Physics Step Optimizations (February 13, 2026)

Problem identified: With 50 NPCs, physics step grew from 2.7ms to 5.8ms+ in 60 seconds (non-scalable).

Root cause: FlatDynamicOctree::rebuild() → radix sort 4 passes × uint32_t counts[65536] = {0} = 1MB of zero-init stack per chunk per frame. With 42 chunks (infinite NPC dispersion) = ~42MB memset/frame.

5 Optimizations Applied

# Optimization Impact
1 Octree bypass for chunks ≤ 32 entities → brute-force N² Eliminates radix sort (1MB/chunk) for small chunks
2 Chunk size 255 → 1000 Reduces chunk count from 42 → 4
3 Linear friction (damping 0.995/frame) Prevents infinite NPC dispersion
4 Sleeping entities (|v|² < 0.01 for 30 frames) Skip physics + migration for stationary entities
5 ThreadPool CPU (batch dispatch to N workers) Parallelizes physicsTick across active chunks

Before / After (20 seconds, 50 NPCs)

Metric Before After Improvement
Physics Step 2.7 → 5.8 ms (growing) 0.052 – 0.064 ms (stable) ~50-100×
Active chunks 17 → 42+ (growing) 4 (stable) No more dispersion
Transit/frame 3-8 (constant) 0 (stable) Sleeping + friction
Stability Linear degradation Constant plateau ✅

Tuning Constants

Constant Value Purpose
VELOCITY_DAMPING 0.995f ~26% loss/sec at 60Hz
SLEEP_VELOCITY_SQ_THRESHOLD 0.01f Threshold 0.1 m/s
SLEEP_FRAMES_THRESHOLD 30 ~0.5s at 60Hz
BRUTE_FORCE_THRESHOLD 32 = octree leafCapacity

Current Benchmark Suite (2026-07-14)

The benchmark app was rewritten since the February 2026 runs above into a broader micro-benchmark suite (lpl-bench harness). These numbers therefore measure different things than the historical tables — keep both; do not equate a row here with a row above. Each line reports mean ± relative stddev, with min and p99, over n samples.

Host: Intel Core Ultra 7 165H, 22 threads, Release build (-O3), CPU backend (no CUDA).

Allocators (100k × 64-byte blocks)

Allocator Mean Notes
Pool (acquire + release) 433.8 µs fastest — free-list reuse
Arena (bump-alloc + reset) 1.325 ms —
malloc + free (libc baseline) 1.528 ms reference

Trigonometry — CORDIC (Fixed32) vs libm

Op Mean Notes
CORDIC sin 1M 12.1 ms deterministic fixed-point
std::sin 1M (double) 4.373 ms libm reference

CORDIC accuracy vs libm: max abs error 1.802e-04, RMS 3.269e-05 over 200,000 samples in [-π/2, π/2]. The point is determinism and no libm dependency, not beating hardware sin.

Physics scalability sweep (integrate + O(n²) collide + sleep, CPU)

Entities Step time Throughput Budget
1k 3.429 ms 2.92e5 ent/s real-time (≥60 fps)
2k 6.78 ms 2.95e5 ent/s real-time (≥60 fps)
5k 19.56 ms 2.56e5 ent/s playable (≥30 fps)
10k 33.05 ms 3.03e5 ent/s playable (≥30 fps)
20k 58.18 ms 3.44e5 ent/s too slow (<30 fps)

The O(n²) collision term is what bends the curve; broad-phase closes the gap (see below). This is the CPU fallback — the CUDA backend is not exercised here.

Data structures & threading

Benchmark Mean Notes
EntityRegistry create/destroy 10k 317.9 µs —
EntityRegistry create 100k 2.066 ms —
WorldPartition build + insert 10k 1.515 ms —
FlatAtomicHashMap insert 16k 329 µs —
FlatAtomicHashMap get 16k 337.8 µs wait-free read
FlatAtomicHashMap forEach 16k 329.5 µs —
FlatAtomicHashMap forEachParallel 16k 1.162 ms thread dispatch overhead dominates at this size
FlatAtomicHashMap insert/remove churn 16k 5.376 ms —
ThreadPool dispatch 10k tasks 28.91 ms —
SparseSet lookup 1k 1.081 µs vs linear scan 3.139 ms (~2900×)

Layout & broad-phase

Benchmark Mean Notes
Sum-x AoS (Vec3[]) 1M 668.2 µs —
Sum-x SoA (float[]) 1M 660.6 µs SoA edges out AoS for the streamed sum
N² AABB collision (2k) 6.696 ms brute force
Broad-phase queryRadius (2k, r=10) 843.8 µs ~8× faster than N²

How to reproduce: xmake build lpl-benchmark then run the binary (e.g. ./build/linux/x86_64/release/lpl-benchmark 10000 100). Numbers vary with host and thermal state — rerun on your own machine before quoting.


← Architectural Decisions | Next: Implementation Status →

Clone this wiki locally