-
-
Notifications
You must be signed in to change notification settings - Fork 1
Performance Results
| Metric | Value |
|---|---|
| Packets sent | 1,000 |
| Packets received | 1,000 |
| Packets lost | 0 (0.00%) |
| Throughput | ~495 pkt/s |
| Metric | Value |
|---|---|
| Average frame time | 62.55 µs |
| Min frame time | 41.93 µs |
| Max frame time | 241.80 µs |
| Variance | ~10 µs |
| Framerate | ~59.49 FPS (locked by usleep) |
- 60 FPS target (16,666 µs/frame): LARGELY EXCEEDED — physics + network consume only 0.38% of the time budget
- Theoretical potential: ~14,000 FPS if framerate cap is removed
- Stability: variance <15%, no critical spikes
-
Zero-copy validated: no visible degradation compared to classic
memcpy
Historical run from the original benchmark app, kept for comparison. The suite has since been rewritten — see Current Benchmark Suite below for the up-to-date numbers.
Configuration: 10,000 entities, 100 frames, random spatial distribution
Command: make benchmark && ./benchmark [numEntities] [numFrames]
| Metric | 10k entities | 50k entities |
|---|---|---|
| Simulation Time | 0.026 sec | 0.180 sec |
| Throughput | 38.6 M ops/sec | 55.6 M ops/sec |
| Avg FPS | 3,915.9 fps | 1,117.8 fps |
| EntityRegistry Lookup | 0.059 µs | 0.073 µs |
| Partition::findEntityIndex | 70.7 ns | 71.7 ns |
| Active Chunks Created | 1,598 | 1,700 |
| Entities per Chunk (avg) | 6.3 | 29.4 |
| Creation Overhead | 0.102 sec | 0.677 sec |
Major change (February 13, 2026): Replacement of std::unordered_map<uint32_t, uint32_t> with adaptive sparse set in Partition::_idToLocal.
Reason: Avoid massive 4MB allocation per chunk (1M × 4 bytes) that caused system hangs with 100+ simultaneous chunks.
Solution implemented: Lazy allocation with adaptive sizing:
- Allocates the sparse set only on first entity insertion
- Sizes to min(entity.id × 1.5, 1MB) instead of fixed 1M
- Reduces typical per-chunk memory to 64KB-256KB
Measured performance:
- O(1) lookup: 70.7 ns (vs 65.8 ns before) — negligible cost
- No repeated dynamic allocations (unlike unordered_map)
- Cache-friendly: sequential sparse set access
- Predictability: guaranteed O(1) (no hash collisions)
Problem identified: With 50 NPCs, physics step grew from 2.7ms to 5.8ms+ in 60 seconds (non-scalable).
Root cause: FlatDynamicOctree::rebuild() → radix sort 4 passes × uint32_t counts[65536] = {0} = 1MB of zero-init stack per chunk per frame. With 42 chunks (infinite NPC dispersion) = ~42MB memset/frame.
| # | Optimization | Impact |
|---|---|---|
| 1 | Octree bypass for chunks ≤ 32 entities → brute-force N² | Eliminates radix sort (1MB/chunk) for small chunks |
| 2 | Chunk size 255 → 1000 | Reduces chunk count from 42 → 4 |
| 3 | Linear friction (damping 0.995/frame) | Prevents infinite NPC dispersion |
| 4 | Sleeping entities (|v|² < 0.01 for 30 frames) | Skip physics + migration for stationary entities |
| 5 | ThreadPool CPU (batch dispatch to N workers) | Parallelizes physicsTick across active chunks |
| Metric | Before | After | Improvement |
|---|---|---|---|
| Physics Step | 2.7 → 5.8 ms (growing) | 0.052 – 0.064 ms (stable) | ~50-100× |
| Active chunks | 17 → 42+ (growing) | 4 (stable) | No more dispersion |
| Transit/frame | 3-8 (constant) | 0 (stable) | Sleeping + friction |
| Stability | Linear degradation | Constant plateau | ✅ |
| Constant | Value | Purpose |
|---|---|---|
VELOCITY_DAMPING |
0.995f | ~26% loss/sec at 60Hz |
SLEEP_VELOCITY_SQ_THRESHOLD |
0.01f | Threshold 0.1 m/s |
SLEEP_FRAMES_THRESHOLD |
30 | ~0.5s at 60Hz |
BRUTE_FORCE_THRESHOLD |
32 | = octree leafCapacity |
The benchmark app was rewritten since the February 2026 runs above into a broader micro-benchmark suite (
lpl-benchharness). These numbers therefore measure different things than the historical tables — keep both; do not equate a row here with a row above. Each line reports mean ± relative stddev, with min and p99, overnsamples.Host: Intel Core Ultra 7 165H, 22 threads, Release build (
-O3), CPU backend (no CUDA).
| Allocator | Mean | Notes |
|---|---|---|
| Pool (acquire + release) | 433.8 µs | fastest — free-list reuse |
| Arena (bump-alloc + reset) | 1.325 ms | — |
malloc + free (libc baseline) |
1.528 ms | reference |
| Op | Mean | Notes |
|---|---|---|
CORDIC sin 1M |
12.1 ms | deterministic fixed-point |
std::sin 1M (double) |
4.373 ms | libm reference |
CORDIC accuracy vs libm: max abs error 1.802e-04, RMS 3.269e-05 over 200,000 samples in [-π/2, π/2]. The point is determinism and no libm dependency, not beating hardware sin.
| Entities | Step time | Throughput | Budget |
|---|---|---|---|
| 1k | 3.429 ms | 2.92e5 ent/s | real-time (≥60 fps) |
| 2k | 6.78 ms | 2.95e5 ent/s | real-time (≥60 fps) |
| 5k | 19.56 ms | 2.56e5 ent/s | playable (≥30 fps) |
| 10k | 33.05 ms | 3.03e5 ent/s | playable (≥30 fps) |
| 20k | 58.18 ms | 3.44e5 ent/s | too slow (<30 fps) |
The O(n²) collision term is what bends the curve; broad-phase closes the gap (see below). This is the CPU fallback — the CUDA backend is not exercised here.
| Benchmark | Mean | Notes |
|---|---|---|
| EntityRegistry create/destroy 10k | 317.9 µs | — |
| EntityRegistry create 100k | 2.066 ms | — |
| WorldPartition build + insert 10k | 1.515 ms | — |
| FlatAtomicHashMap insert 16k | 329 µs | — |
| FlatAtomicHashMap get 16k | 337.8 µs | wait-free read |
| FlatAtomicHashMap forEach 16k | 329.5 µs | — |
| FlatAtomicHashMap forEachParallel 16k | 1.162 ms | thread dispatch overhead dominates at this size |
| FlatAtomicHashMap insert/remove churn 16k | 5.376 ms | — |
| ThreadPool dispatch 10k tasks | 28.91 ms | — |
| SparseSet lookup 1k | 1.081 µs | vs linear scan 3.139 ms (~2900×) |
| Benchmark | Mean | Notes |
|---|---|---|
Sum-x AoS (Vec3[]) 1M |
668.2 µs | — |
Sum-x SoA (float[]) 1M |
660.6 µs | SoA edges out AoS for the streamed sum |
| N² AABB collision (2k) | 6.696 ms | brute force |
Broad-phase queryRadius (2k, r=10) |
843.8 µs | ~8× faster than N² |
How to reproduce:
xmake build lpl-benchmarkthen run the binary (e.g../build/linux/x86_64/release/lpl-benchmark 10000 100). Numbers vary with host and thermal state — rerun on your own machine before quoting.
← Architectural Decisions | Next: Implementation Status →
LplPlugin — Zero-Copy Real-Time VR Engine | Author: MasterLaplace | License: GPL-3.0 | Source Code
This project is a marathon, not a sprint.