Repository navigation
Hades II (PPSA36082) #1965
Raghav314
started this conversation in
Game bring-up
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Tested
Title: Hades II, PPSA36082, v01.006.000
Host:
Linux (CachyOS, kernel 7.2.9), Ryzen 7 7800X3D, 32 GB RAM.
RTX 4090 24 GB on the NVIDIA proprietary driver 615.71.09.
Relinked for Linux with the relinker from #1249, built with GCC 16.2.1.
Last update: 2026-10-11 (evening)
Status
Playable, with warm shader caches. On main 961b25f with the PRs below, it boots, loads a game and plays, with no skipped draws. Standing in the hub it now sits at the 60 fps cap (p95 frame 16.7 ms) with pipelined draws from #1201 on (
APS5_PIPELINED_DRAWS=1). That's the original executable with its 512 MiB upload pool, no patches.The level-load crash from yesterday is gone on main: #2647 gives guest threads their host stack room back. The one catch left is that it only boots when the shader caches are warm. More on that below.
Some things worth knowing about how the game runs:
libc.prx,libSceJobManager.prx,libSceNpCppWebApi.prx,libSceFontGsm.prx,libScePfs.prxand FMOD (libfmod.prx,libfmodstudio.prx). All of them load as guest modules fromsce_module.3).0x32): a 512 MiB upload pool, a 6 GiB heap and another 128 MiB.You need an import budget of at least 6784 MiB, so either #290 or
APS5_HOST_IMPORT_MIB=8192. I run with the env var, since #290 doesn't merge cleanly with current main.Everything it needs right now
All open unless marked otherwise.
APS5_HOST_IMPORT_MIB=8192(#861, #1249, #1367, #1905, #1959, #1960 and #1961 are merged)feat/agc-msaa-sample-locations, which is written on #419). #419, #2121 and #2132 need a rebase on current mainSpeed in the hub
Standing still in the hub, 20 s of MangoHud each, warm caches. The first four rows are without draw profiling, the rest with
APS5_PROFILE_DRAWon (which costs a few fps), each against a run of the row before in the same session:APS5_PIPELINED_DRAWS=1)#3231 lets draws share a render pass (84% of them now continue one), and #3237 lets them reuse their descriptor sets (91% resource cache hits). Both had been blocked because nearly all of this game's shaders carry a fault buffer for guarded descriptor slots.
#494's data-only hits reused almost no stages here, because each sprite changes a user word its vertex stage never reads. Comparing only the words a stage reads took the reused stages from ~7k to ~200k per 10 s (numbers in my comment on #494).
With #1201 the draw recording moves to its own thread: the packet worker is at about 16 us per draw and the recording thread at about 8 us, so the hub is capped by vsync. One thing to know when combining #1201 with #419: the CB resolve pass ran ahead of queued draws, which made the characters flicker or go dark. Draining the pipeline before the pass fixes it (details in my comment on #1201).
Speed in menus and gameplay
The hub is capped at 60 now, but the item and boon pickup menus and busy fights still dip (~1000+ draws a frame there instead of ~700). In a pickup menu, profiling on, warm caches:
The GPU is only about 37% busy here, so the limit is still the packet worker translating draws (about 13.6 us per draw). Next thing I'm trying is moving draw preparation onto helper threads (Michon's
APS5_PARALLEL_DRAWSfrom his gta-v branch); my local port of it crashes at startup for now.Cold cache vs warm cache
This is the main thing left for booting.
Why: on a cold start, the graphics queue hits draws whose pipelines aren't compiled yet and sits in
vkCreateGraphicsPipelines: about 65 ms for the worst one, and three new pipelines show up back to back. The loader's fence is behind those draws on the same queue, so it lands too late and the loader keeps staging until the pool runs out.Also, the pipeline cache is only written when the game exits cleanly. A run that crashes never saves it, so retrying cold doesn't warm anything up. A run that gets past startup and exits normally does.
The fix is being done in #2242: compile the shader stages at registration, so a draw only links them. The maintainers went with
VK_EXT_graphics_pipeline_library; phase A (#2636, linking from per-stage libraries), phase B1 (#2645, dynamic rendering), #2661 (fixed push slots) and #2713 (link-time optimization in the background) are merged. A cold start still overflows: one vertex shader compile at the draw (86-90 ms), which phase B moves to registration. Measurements are in #2242.What's still broken
What stopped it on main, and what fixed it
sleep,sceKernelGetAvailableCpumask,sceVideoOutAllowOutputResolutionWqhdDetection,sceAudioInAsyncOpen,sceAgcSetSemaphoreMemorysleep, #801, #802, #910, #1019 (merged)libSceNpCppWebApi's constructor. The bundled libc rejects mspace bases outside the PS5 address window, so its heap ends up null0x11) fails to pin, and later allocations on the device fail0x28fba9) when the first flip replaces the deviceAPS5_HOST_IMPORT_MIB=8192sceAjmBatchJobSetResampleParametersdisplay buffer reads mixed DCC keysprepared shader artifact is missing for the requested static ABI: every draw without a pixel program misses its prepared variant (#1901)sceAgcCreateInterpolantMapping→ shader preparation, where computing post-dominators took most of the timeunsupported color depth, dimension, resource level or metadata modefor a 512x1 1D color targetPA_SC_AA_CONFIGat 0 and puts the count in the target)0x444cc4cc)feat/agc-msaa-sample-locationsdescribePages) are repeated for the same pages every frameAJM: decoding codec <garbage> is not implemented, AJM state freed while FMOD still runs0xFFFFFE74, and AJM copied ~4 GB per jobNotes
libfmod.prxandlibfmodstudio.prxfrom/app0, so in my run folder those are links to the relinked copies insce_module.All reactions