Skip to content

Performance Notes

Deepratna Awale edited this page Oct 6, 2026 · 3 revisions

Performance notes

What the optimisation work measured and changed. Figures come from the optimisation plans and the perf: commits; they were measured on Apple silicon (an M4 for the 4K runs) and will differ on other Macs.

Where the time goes

  • On a 4K display, particles cost 0.2–2.8 ms of GPU a frame in every tested scene; the render thread 0.5–2.5 ms; scripts 0.1–0.4 ms.
  • The lag in heavy scenes is WE effects at the layers' texture size: e.g. 245–501 megapixels of effect passes a frame on scenes with several very large layers. Without effects every tested scene drew in under 4 ms.
  • The heaviest particle wallpaper is fill-bound (20 000 refracting rain sprites), not simulation-bound.

What changed

Shaders and pipelines

  • In-process glslang and SPIRV-Cross instead of process spawns: about 90× faster first loads.
  • The MTLBinaryArchive skips the GPU backend compile of pipelines seen before, written once per compile burst.
  • Variants keyed only on the combos their shaders name; precompiled regular expressions; each source analysed once.

Effects and render targets

  • A chain's static prefix is kept; ping-pong targets are made on first use; a chain is reused while its live-bound constants hold.
  • Scene Detail: Match Display runs effects at on-screen size. GPU ms at a 3840×2160 display, before → full / match / match+desktop: 27.7 → 23.3 / 19.7 / 5.1 on one heavy scene, 30.2 → 30.2 / 10.0 / 4.0 on another.
  • A scaled-up composition layer keeps its own buffer size (one 4.1× ring shaded 67 MPix a frame before).
  • Render targets are pooled, leased, and evicted by an LRU byte budget (256 MB) and idle time; spare effect targets are kept by bytes, not count.
  • Scene-reading effects read the paused scene target; the reflection snapshot shares the mip-mapped buffer (−122 MB at 5K on one scene).
  • The scene pass stores samples and depth only when resumed or read; the reflection's depth is memoryless.

Particles

  • GPU stages run for every system at once: a small system's fixed cost fell from about 48 µs to 3 µs; library simulations take 0.03–0.27 ms of GPU (was up to 0.43 ms).
  • Pipeline keys, uniform members, time of day and sprite axes cached; strips indexed; compaction writes survivors directly; boids read their slice's list; pre-simulation steps skip unused work.
  • The particle budget thins only scenes above it, counting what emitters keep alive, not only maxcount.

Textures

  • R8, RG88 and RGBA8888 .tex images upload in their own channels without intermediate copies.
  • Layers, particle systems and effect assets loading the same image share one GPU texture.
  • Texture Resolution: High Performance loads from the second mipmap.

Scripts and timelines

  • JavaScriptCore JIT; no per-frame closures or arrays per bound script; unchanged material writes send nothing.
  • Timelines under 4 µs a frame for the whole library.

Multi-display

  • One instance per wallpaper: two 1080p displays of one scene cost 0.38–0.56 of two renderers' CPU time and mostly 0.25–0.74 of their GPU time (SceneSharedInstanceBenchmarkTests).

Shadows

  • A frame that would redraw exactly what the shadow atlas holds keeps it; a caster draws only into the views that hold it; casters are prepared once a frame.

UI

  • Switching to Installed no longer rescans the library a dozen times; GIF animations pause while the app is inactive.
  • Per-frame settings apply without rebuilding the content.

Measuring

Tool Use
SceneFrameBenchmarkTests (OWE_SCENE_BENCH, OWE_SCENE_BENCH_PASSES) Frame GPU/CPU time per scene and per effect pass
SceneDetailEquivalenceTests Match Display vs Full, cost and image difference
ParticleLibraryBenchmarkTests, LightingFrameBenchmarkTests, SceneTextureLoadBenchmarkTests, SceneSharedInstanceBenchmarkTests Area benchmarks
OWESignpost Signposts are always emitted and cost nothing unless Instruments is recording
Log Level Verbose Turns on per-frame metrics (OWEFrameMetrics)

User guide: Settings › Performance

Clone this wiki locally