Skip to content

v1.3.5 — full AMDGPU operand decode + differential gate

Choose a tag to compare

@zoratu zoratu released this 27 Apr 05:08
· 532 commits to main since this release

Highlights

v1.3.5 closes the entire v1.3.x AMDGPU docket. Every byte hexray emits is now byte-checked against llvm-objdump for eight SCALE-built kernels spanning GFX9 (gfx900), RDNA1 (gfx1010), RDNA2 (gfx1030), and RDNA3 (gfx1100/1101/1102) — including a multi-kernel fixture. RDNA3-specific opcode renumbering for SOP2/VOPC/VOP2/SOP1 is fully handled, and operand rendering is wired for VOP3 / SMEM / FLAT / MUBUF / DS classes including VOP3 NEG+ABS modifiers. Three pre-existing cross-cutting Linux decompiler bugs surfaced during release validation also got fixed: PLT-stub @plt symbol synthesis, DWARF CFA→fp frame correction, and var_NN → DWARF parameter-name override.

What's new (AMDGPU)

  • Differential gate (amdgpu_differential_gate.rs) — two CI tests asserting (a) zero xxx.op0x... placeholder mnemonics across the corpus and (b) per-mnemonic frequency parity (±1) against the llvm-objdump sidecar within each <kernel$local>: block.
  • Operand rendering for VOP3 / SMEM / FLAT / MUBUF / DS classes. Width-aware register pairs (v[0:1], s[2:3]), VOP3B SDST, VOPC-as-VOP3 with implicit EXEC dst, FLAT SADDR null sentinel, MUBUF resource-descriptor pair (s[N*4:N*4+3]), DS offset0/offset1.
  • VOP3 NEG + ABS modifier rendering. Renders |src| for ABS, -|src| for NEG+ABS to match llvm-objdump. Suppressed for VOP3B forms.
  • SIMM16 sub-decoding for s_clause, s_waitcnt, s_delay_alu — full sub-field decode like s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_1) | instid1(VALU_DEP_1).
  • HIP host-binary fatbin extraction — parses the __CLANG_OFFLOAD_BUNDLE__ schema. 14 new tests on synthetic bundles. LZ4-compressed bundles remain a documented limitation.
  • SCALE .AMDGPU.kinfo parser — reverse-engineers SCALE-free's private kernarg-layout section format so hexray cmp can show arg counts and per-arg sizes.
  • GFX9 opcode tables — SOP2_GFX9 / VOPC_GFX9 / VOP3_GFX9 / FLAT_GFX9 carved out from the shared tables. VOP2_GFX9 extended with the carry/no-carry pairs and corrected mnemonics — 0x19 is v_add_co_u32_e32 (with carry), not v_add_u32_e32.
  • RDNA3 SOP2 audit — new SOP2_GFX11 table. Only s_add_u32..s_subb_u32 (0x00..=0x05) survive from GFX10; s_min_i32 moved 0x06→0x12, s_and_b32 0x0e→0x16, s_xor_b32 0x12→0x1a, s_lshl_b32 to 0x08, s_mul_i32 0x26→0x2c. Cross-checked against LLVM SOPInstructions.td.
  • RDNA1+ coverage — SOP1_GFX10 s_and_saveexec_b32 at 0x3c, VOP3_GFX10/VOP3_GFX11 v_cndmask_b32_e64 at 0x101, VOPC_GFX11 (RDNA3 packs i32 comparators into 0x40..=0x4e). VOP2_GFX11 integer min/max e32 family at 0x11..=0x14.
  • cargo-mutants sweep — full pass over amdgpu/** and elf/amdgpu/** (499 mutants) drove ~80 new unit tests. 162 of 170 missed mutations now caught (the remaining 8 are semantically equivalent — every Operation in the Vopc and Smem tables matches the class default, so table-vs-default branches are unobservable).

What's new (decompiler — Linux fixes surfaced during release validation)

  • PLT-stub @plt symbol synthesis in the ELF parser. Calls through the GOT (call qsort@plt) target an address inside .plt (or .plt.sec on CET-aware glibc). The new code walks .rela.plt in order, looks each entry's symbol up in .dynsym, and adds a Symbol { name: "<dynsym>@plt", ... } at the matching stub address. Calls now decompile as qsort(...) instead of sub_NNN(...).
  • DWARF CFA → fp frame correction. When DW_AT_frame_base is DW_OP_call_frame_cfa (the clang -O0 -g default), DWARF emits operand offsets relative to the CFA — at fp + 16 after the standard prologue. The variable-name map now rebases by +16 so DWARF parameter names actually surface in decompiled output.
  • var_NN → DWARF override in signature emission. SignatureRecovery's stack-spill heuristic was shadowing DWARF parameter names with var_NN slot offsets. The emitter now post-processes signature param names: var_<hex> → DWARF name when available.

Coverage

amdgpu/encoding.rs 99% lines • amdgpu/mod.rs 88% • amdgpu/opcodes.rs 100% • amdgpu/registers.rs 100% • elf/amdgpu/descriptor.rs 100% • elf/amdgpu/scale_kinfo.rs 91% • elf/amdgpu/msgpack.rs 61% (the weakest — lots of error-path branches).

What's deferred

  • CDNA MFMA / WMMA / VOP3P / DPP / SDWA opcodes — SCALE-free doesn't ship gfx906/908/90a/940 targets. Tracked alongside a future CDNA corpus.
  • End-to-end fixture validation for MUBUF / DS / VOP3 ABS — scale-free builds don't exercise these classes; the new code is byte-validated against synthetic encodings only.
  • LZ4 fatbin decompression (CCOB magic) — needs a real hipcc-built binary.
  • gfx1200 (RDNA4) validation — requires commercial SCALE.
  • Linux-snapshot decompiler tests for test_decompile_callback_* — the existing snapshot tests are macOS-locked; Linux output now decompiles correctly but matches a different shape.

Full changelog

CHANGELOG.md