Repository navigation
Performance Notes
Deepratna Awale edited this page Oct 6, 2026
·
3 revisions
What the optimisation work measured and changed. Figures come from the optimisation plans and the perf: commits; they were measured on Apple silicon (an M4 for the 4K runs) and will differ on other Macs.
- On a 4K display, particles cost 0.2–2.8 ms of GPU a frame in every tested scene; the render thread 0.5–2.5 ms; scripts 0.1–0.4 ms.
- The lag in heavy scenes is WE effects at the layers' texture size: e.g. 245–501 megapixels of effect passes a frame on scenes with several very large layers. Without effects every tested scene drew in under 4 ms.
- The heaviest particle wallpaper is fill-bound (20 000 refracting rain sprites), not simulation-bound.
- In-process glslang and SPIRV-Cross instead of process spawns: about 90× faster first loads.
- The MTLBinaryArchive skips the GPU backend compile of pipelines seen before, written once per compile burst.
- Variants keyed only on the combos their shaders name; precompiled regular expressions; each source analysed once.
- A chain's static prefix is kept; ping-pong targets are made on first use; a chain is reused while its live-bound constants hold.
- Scene Detail: Match Display runs effects at on-screen size. GPU ms at a 3840×2160 display, before → full / match / match+desktop: 27.7 → 23.3 / 19.7 / 5.1 on one heavy scene, 30.2 → 30.2 / 10.0 / 4.0 on another.
- A scaled-up composition layer keeps its own buffer size (one 4.1× ring shaded 67 MPix a frame before).
- Render targets are pooled, leased, and evicted by an LRU byte budget (256 MB) and idle time; spare effect targets are kept by bytes, not count.
- Scene-reading effects read the paused scene target; the reflection snapshot shares the mip-mapped buffer (−122 MB at 5K on one scene).
- The scene pass stores samples and depth only when resumed or read; the reflection's depth is memoryless.
- GPU stages run for every system at once: a small system's fixed cost fell from about 48 µs to 3 µs; library simulations take 0.03–0.27 ms of GPU (was up to 0.43 ms).
- Pipeline keys, uniform members, time of day and sprite axes cached; strips indexed; compaction writes survivors directly; boids read their slice's list; pre-simulation steps skip unused work.
- The particle budget thins only scenes above it, counting what emitters keep alive, not only
maxcount.
- R8, RG88 and RGBA8888
.teximages upload in their own channels without intermediate copies. - Layers, particle systems and effect assets loading the same image share one GPU texture.
- Texture Resolution: High Performance loads from the second mipmap.
- JavaScriptCore JIT; no per-frame closures or arrays per bound script; unchanged material writes send nothing.
- Timelines under 4 µs a frame for the whole library.
-
One instance per wallpaper: two 1080p displays of one scene cost 0.38–0.56 of two renderers' CPU time and mostly 0.25–0.74 of their GPU time (
SceneSharedInstanceBenchmarkTests).
- A frame that would redraw exactly what the shadow atlas holds keeps it; a caster draws only into the views that hold it; casters are prepared once a frame.
- Switching to Installed no longer rescans the library a dozen times; GIF animations pause while the app is inactive.
- Per-frame settings apply without rebuilding the content.
| Tool | Use |
|---|---|
SceneFrameBenchmarkTests (OWE_SCENE_BENCH, OWE_SCENE_BENCH_PASSES) |
Frame GPU/CPU time per scene and per effect pass |
SceneDetailEquivalenceTests |
Match Display vs Full, cost and image difference |
ParticleLibraryBenchmarkTests, LightingFrameBenchmarkTests, SceneTextureLoadBenchmarkTests, SceneSharedInstanceBenchmarkTests
|
Area benchmarks |
OWESignpost |
Signposts are always emitted and cost nothing unless Instruments is recording |
| Log Level Verbose | Turns on per-frame metrics (OWEFrameMetrics) |
User guide: Settings › Performance
Open Wallpaper Engine · GPL-3.0 · Released by Deepratna Awale · Based on Open Wallpaper Engine by Haren Chen and MrWindDog · Not affiliated with Wallpaper Engine or Valve · Home · User Guide · Developer Guide