Repository navigation
STAR WARS: Galactic Racer (PPSA28416) #3086
CnRJay
started this conversation in
Game bring-up
Replies: 1 comment
|
Nice work. The (found with nidhunt, also in zecoxao/sce_symbols#23). I opened #3105 to name it as a throwing stub so the import resolves; the real behaviour is yours to keep in the bring-up. Keep up the good work! |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Tested
--windows --registry. Built with WinLibs GCC 15.2.0 r7, Release,ANYPS5_ENABLE_SPIRV_TOOLS=OFF(with SPIR-V validation on, the per-shader cost makes UE's hang watchdog fire).Testing the branch
Everything below, merged, open and local only, is combined on
CnRJay/AnyPS5port/ppsa28416. It is a test branch, not for review: it carries work-in-progress commits, one shim (ZVZ++68nNI8) and a debug probe ([probe-indirect]lines in the log).Every stop in the table below is handled on the branch, and everything under "Local only" is on it:
2f9cee4d6: storage-image writes to a depth surface's planes are copied into its depth image);704dec9f8), so geometry is not clipped by them.Relink the title with
--windows --registryand append-nothreadtimeout -HangDuration=1200toapp0/uecommandline.txt(see Status). Expect 0.04-0.14 FPS and the stop described in blocker A.Status
Not in game. On a branch of
mainwith the pull requests below and the local changes, the title:ZVZ++68nNI8, is a local shim);It runs at 0.04-0.14 FPS, and stops after a few minutes on garbage indirect dispatch arguments (blocker A).
Test-only setting:
-nothreadtimeout -HangDuration=1200is appended toapp0/uecommandline.txt. First-time pipeline compiles of very large shaders (100k-2M SPIR-V words) and the slow frames otherwise make UE'sFThreadHeartBeatcrash the title on purpose.Open blockers
609717307x1x1,66977792x1x1,66978560x4286578689x1. With y and z in range the 67M-group dispatch runs for minutes and looks like a GPU hang (recorded batch ... still running). The read does not happen withAPS5_TRACE_DISPATCH_IO, which only slows the CPU, so it looks like a missing sync between a GPU write of the arguments and the CPU read. The writer is not identified yet.[gpu] deferred write-back failed: BDA access failed ... reason=1at0x8000000000000000+...(pc 0x674, 0xbfc, 0xc74, 0xe48 of the same program): stores through a pointer with bit 63 set. Not fatal.guest snapshot differs from registered memory(compute shader 0x131d2a0000) andbuffer descriptor uses an unsupported type(0x7ff710ad5600). Seen once each, after blocker A had already produced garbage; possibly follow-on damage.PA_CL_VS_OUT_CNTL= 0x0040000f); locally they are accepted and not applied. Covered by #423 (draft, conflicting).Stops on
main, in the order they are hitscePsmlMfsrCreateContext1300,scePsmlMfsrGetDispatchMfsrPacket1300party-ps5-shipping.prx:sceAudioOut2EnableChatZVZ++68nNI8(libSceNpManager, name unknown). OnlineSubsystemPS5 callsf(char* buf)with a 0x50-byte stack buffer and reads it as a stringCB metadata pass over DCC keys that are mixed: comp-to-single DCC keys (0x10) in a DCC decompressVOP1 SDWA destination selector 4 ... of opcode 0x50(v_cvt_f16_u16with a byte source)CFG shared merge splitting exceeded budgetvkAllocateMemory -2: the sparse 4 GiB descriptor heap copied wholestorage image access to depth/stencil surface: UE writes scene depth (R32F) and stencil (R8) from computeHeader: invalid packet header(one-dword NOP)...RegistersIndirectGetSize: reserved bits (0xf5f5039c): the count argument holds leftover register contentssceAgcBranchPatchSetThenTarget_0300 not implementeddispatch modifiers 0x49(ORDERED_APPEND_ENBL)null or misaligned addressin a prepared shader's SRT walk (s_loadfrom a null pointer, skipped on hardware bys_cbranch_execz)clip distances ... (0x207 = 0x0040000f)workgroup count 16926492x1116079131x3241227624: the compute-queue look-ahead read indirect arguments before an earlier dispatch of the same run wrote themcomputed/data-dependent s_swappc_b64 call target ... is not statically resolvable(ordered-append compute shaders, 0x131f420000)ds_ordered_count ... not modeledguest memory is not readable at 0xfe0040000storage image access to the depth plane ... as a 1x1 image(a depth surface's memory reused for a smaller or differently sized image)shader buffer exceeds descriptor range limit(a V# of0x11e1a30160+0x3fdfffffc0, stride 64)buffer atomic on a V# whose base is not DWORD aligned; alsoBDA access failed ... reason=7(unaligned) at pc 0xa08Local only (pull requests to follow)
s_swappc_b64targets. The ordered-append shaders call a hook through a table at the fixed address 0xfe0040000, only when the table entry is non-null. Such a call now compiles; if it runs, it records a fault and the dispatch throws.ds_ordered_count. Implemented as one GDS add per wave of the first active lane's value; M0[31:16] is the address and M0[15:0] the ordered wave ID (LLVMSIISelLowering.cpp,IntrinsicsAMDGPU.td). Waves get their ranges in the order their atomics run, not in launch order. Swap, more than one dword, andwave_donewithoutwave_releasethrow. ORDERED_APPEND_ENBL in the dispatch initiator is accepted.maxStorageBufferRangeis clamped to the device limit, ending at its last committed byte; the uncommitted parts read as zero.Notes
libSceAjm.native.prx, libSceJson vslibSceJson2.prx, libScePsml vslibScePsml_debug.prx, EOSSDK-PS5-Shipping vs.debug_prx, ...). Those imports are resolved by searching every library and load fine.import_audit.pylists Party.prx and libPlayFabMultiplayer.prx as missing, because the bundled files are namedparty-ps5-shipping.prxandlibplayfabmultiplayer-ps5-shipping.prx.its s_swappc_b64 ... calls a computed target) and are prepared when first used.All reactions