Skip to content

Lessons Learned

MasterLaplace edited this page Jul 14, 2026 · 4 revisions

Lessons Learned

Performance

  • Measure before optimizing: 62.55µs is excellent, don't over-optimize without reason
  • Simplicity wins: static array > complex slab allocator for the ring buffer
  • Zero-copy >> everything: eliminating copies is more effective than any "smart" allocator
  • SoA >> AoS: memory coalescence is measurable, not theoretical

Linux Kernel Development

  • GFP_ATOMIC is mandatory in Netfilter hooks (network interrupt context)
  • SpinLock = contention: RCU is preferable for frequent reads in the kernel
  • Critical error handling: cascade of goto on initialization failure (standard kernel pattern)
  • NF_DROP vs NF_ACCEPT: conscious design — intercepted packets don't go up the network stack

Architecture

  • One SpinLock per chunk, not a global lock: distributes contention
  • Backward iteration for swap-and-pop: bug fixed, memorable lesson
  • Double buffers cost 2× in hot memory only: acceptable compromise
  • Transparent PinnedAllocator: the #ifdef __CUDACC__ makes g++ compilation possible without modifications

Methodology

  • README ≠ contract: the roadmap is a guide, not an absolute obligation
  • Question the premises: "why this optimization now?"
  • Validate empirically: performance data before any added complexity
  • Document decisions: every architectural choice must be justified with technical reasons

Networking

  • Unify server and client networking: a single Network class eliminates code duplication and protocol drift
  • Header-only = simpler builds: removing the separate .cpp/.o eliminates link order issues
  • Socket fallback is essential: developing/testing without the kernel module (WSL, CI) greatly speeds up iteration
  • Batch TX via kernel thread: one ioctl wake for N outbound packets is far cheaper than N sendto() syscalls
  • Identify clients by src_ip/src_port in RxPacket: the kernel module can extract sender info from the IP header

Modularization

  • Flat module architecture scales: 20 independent lpl-<name> libraries with explicit deps is cleaner than one monolithic engine/
  • xmake > Make for C++ projects: native dependency resolution, package management, and build modes eliminate Makefile complexity
  • Archive legacy, don't delete: the previous-generation prototype was kept (out of the build tree, with git history) for reference and feature-parity checks rather than deleted outright
  • Code quality requires consistency: Doxygen style (/**@), .inl files, guard clauses — must be enforced from the start, not retrofitted

← Scientific Contributions | Next: References →

Clone this wiki locally