v1.3.5 — full AMDGPU operand decode + differential gate
Highlights
v1.3.5 closes the entire v1.3.x AMDGPU docket. Every byte hexray emits is now byte-checked against llvm-objdump for eight SCALE-built kernels spanning GFX9 (gfx900), RDNA1 (gfx1010), RDNA2 (gfx1030), and RDNA3 (gfx1100/1101/1102) — including a multi-kernel fixture. RDNA3-specific opcode renumbering for SOP2/VOPC/VOP2/SOP1 is fully handled, and operand rendering is wired for VOP3 / SMEM / FLAT / MUBUF / DS classes including VOP3 NEG+ABS modifiers. Three pre-existing cross-cutting Linux decompiler bugs surfaced during release validation also got fixed: PLT-stub @plt symbol synthesis, DWARF CFA→fp frame correction, and var_NN → DWARF parameter-name override.
What's new (AMDGPU)
- Differential gate (
amdgpu_differential_gate.rs) — two CI tests asserting (a) zeroxxx.op0x...placeholder mnemonics across the corpus and (b) per-mnemonic frequency parity (±1) against the llvm-objdump sidecar within each<kernel$local>:block. - Operand rendering for VOP3 / SMEM / FLAT / MUBUF / DS classes. Width-aware register pairs (
v[0:1],s[2:3]), VOP3B SDST, VOPC-as-VOP3 with implicit EXEC dst, FLAT SADDR null sentinel, MUBUF resource-descriptor pair (s[N*4:N*4+3]), DS offset0/offset1. - VOP3 NEG + ABS modifier rendering. Renders
|src|for ABS,-|src|for NEG+ABS to match llvm-objdump. Suppressed for VOP3B forms. - SIMM16 sub-decoding for
s_clause,s_waitcnt,s_delay_alu— full sub-field decode likes_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_1) | instid1(VALU_DEP_1). - HIP host-binary fatbin extraction — parses the
__CLANG_OFFLOAD_BUNDLE__schema. 14 new tests on synthetic bundles. LZ4-compressed bundles remain a documented limitation. - SCALE
.AMDGPU.kinfoparser — reverse-engineers SCALE-free's private kernarg-layout section format sohexray cmpcan show arg counts and per-arg sizes. - GFX9 opcode tables —
SOP2_GFX9/VOPC_GFX9/VOP3_GFX9/FLAT_GFX9carved out from the shared tables.VOP2_GFX9extended with the carry/no-carry pairs and corrected mnemonics —0x19isv_add_co_u32_e32(with carry), notv_add_u32_e32. - RDNA3 SOP2 audit — new
SOP2_GFX11table. Onlys_add_u32..s_subb_u32(0x00..=0x05) survive from GFX10;s_min_i32moved 0x06→0x12,s_and_b320x0e→0x16,s_xor_b320x12→0x1a,s_lshl_b32to 0x08,s_mul_i320x26→0x2c. Cross-checked against LLVMSOPInstructions.td. - RDNA1+ coverage —
SOP1_GFX10s_and_saveexec_b32at 0x3c,VOP3_GFX10/VOP3_GFX11v_cndmask_b32_e64at 0x101,VOPC_GFX11(RDNA3 packs i32 comparators into 0x40..=0x4e).VOP2_GFX11integer min/max e32 family at 0x11..=0x14. - cargo-mutants sweep — full pass over
amdgpu/**andelf/amdgpu/**(499 mutants) drove ~80 new unit tests. 162 of 170 missed mutations now caught (the remaining 8 are semantically equivalent — everyOperationin theVopcandSmemtables matches the class default, so table-vs-default branches are unobservable).
What's new (decompiler — Linux fixes surfaced during release validation)
- PLT-stub
@pltsymbol synthesis in the ELF parser. Calls through the GOT (call qsort@plt) target an address inside.plt(or.plt.secon CET-aware glibc). The new code walks.rela.pltin order, looks each entry's symbol up in.dynsym, and adds aSymbol { name: "<dynsym>@plt", ... }at the matching stub address. Calls now decompile asqsort(...)instead ofsub_NNN(...). - DWARF CFA → fp frame correction. When
DW_AT_frame_baseisDW_OP_call_frame_cfa(theclang -O0 -gdefault), DWARF emits operand offsets relative to the CFA — atfp + 16after the standard prologue. The variable-name map now rebases by +16 so DWARF parameter names actually surface in decompiled output. var_NN→ DWARF override in signature emission. SignatureRecovery's stack-spill heuristic was shadowing DWARF parameter names withvar_NNslot offsets. The emitter now post-processes signature param names:var_<hex>→ DWARF name when available.
Coverage
amdgpu/encoding.rs 99% lines • amdgpu/mod.rs 88% • amdgpu/opcodes.rs 100% • amdgpu/registers.rs 100% • elf/amdgpu/descriptor.rs 100% • elf/amdgpu/scale_kinfo.rs 91% • elf/amdgpu/msgpack.rs 61% (the weakest — lots of error-path branches).
What's deferred
- CDNA MFMA / WMMA / VOP3P / DPP / SDWA opcodes — SCALE-free doesn't ship gfx906/908/90a/940 targets. Tracked alongside a future CDNA corpus.
- End-to-end fixture validation for MUBUF / DS / VOP3 ABS —
scale-freebuilds don't exercise these classes; the new code is byte-validated against synthetic encodings only. - LZ4 fatbin decompression (CCOB magic) — needs a real hipcc-built binary.
- gfx1200 (RDNA4) validation — requires commercial SCALE.
- Linux-snapshot decompiler tests for
test_decompile_callback_*— the existing snapshot tests are macOS-locked; Linux output now decompiles correctly but matches a different shape.