Skip to content

dae v0.1.8: 50x faster decompilation, a 20-command progressive CLI

Choose a tag to compare

@github-actions github-actions released this 27 Sep 21:52
· 16 commits to main since this release
v0.1.8
3b4cc18

The last release listed "decompiler wall-clock is not improved, and no speedup is claimed" as its top open item. That item is closed: --decompile is ~50× faster, and the CLI grows from 12 subcommands to 20.

Speed

Interleaved A/B against v0.1.7 on the same machine, two rounds each (the host carries unrelated load, so single runs are not evidence):

workload v0.1.7 v0.1.8
material_3_demo (15,082 functions) --decompile 106.8 / 108.7 s 1.91 / 2.35 s ~50×
real Android arm64 app (compressed pointers) --decompile 175.2 / 178.1 s 2.74 / 2.75 s ~64×
material_3_demo, export only 1.49 / 1.43 s 0.54 / 0.50 s ~2.8×
peak RSS 206–232 MB 179–211 MB −27…−47 MB

None of this came from a faster algorithm; it came from stopping three rebuilds of run-invariant data:

  • lift rebuilt the whole object-pool map once per function — 122,064 entries on a real app, rebuilt 15,082 times on the sample above — and Structurer owned its Roles by value, so emit_function deep-cloned that same map again per function. Both now borrow. This is the ~33× and it also explains why export got 2.8× faster: field recovery runs through the same lift path.
  • mask_regs rebuilt its register-alias table on every instruction — clone the profile aliases, append pp/thr and nine hardcoded pairs, stable-sort by descending key length, then one replace_word per pair, each allocating a String and rescanning the whole text. ~1M instructions × 20 pairs ≈ 20M allocations. The table is now built once; lift went 1.95 s → 0.63 s.
  • sort_by_key does not cache its key. callgraph sorted edges with sort_by_key(|a| (…, a.to_text.clone())), so it allocated a String per comparison: 96,904 edges × O(log N) ≈ 1.6M clones. Plus an O(n²) name lookup in the stubs exporter.

A single-pass alias lookup is not equivalent to sequential replacement in general, because the replacements cascade: the arm64 profile maps x29 → fp and x30 → lr (lowercase) and the hardcoded tail maps fp → FP, lr → LR, so x29 reaches the output as FP. The new table resolves each key by simulating the chain in the original pair order — graph reachability would be wrong here, since it also fires on a value that equals an earlier key, which sequential replacement never re-applies.

Two changes are kept but credited with nothing: capstone's .detail(true) → .detail(false) (no detail API is used anywhere, so it is strictly less work, but the A/B showed no measurable gain), and the callgraph fix's effect on wall time (it bought memory and removed 1.6M allocations; sorting 96k edges was already fast).

CLI: 20 subcommands on a declarative tree

get oriented   info · libs · classes · functions · largest
find things    strings · fields · members · findrefs · callers · callees
object layer   pp · objs · stubs
decompile      getclass · getmethod · getlib · decompile
low level      disasm
full export    export        (and `dae <binary> <out_dir>` stays as an equivalent shortcut)

Parsing moved to clap. The reason is not fashion — the hand-rolled parser was written twice over, and that shape is where the sibling tool's CLI bugs come from: an -o flag that was dead code because an earlier loop bailed on unknown options first, a subcommand that took a flag's value as its input path, and one that parsed --dex then dropped it. With flags declared once and positionals parsed independently, those three cannot occur. Shared options live in one struct flattened into every command; seven commands share one <binary> [pattern> shape and five share <binary> <name>, so their semantics cannot drift apart.

New in this release:

  • findrefs <bin> string TEXT — every code site that loads a given string literal from the object pool, and findrefs <bin> kind NAME for an object kind (the /* TypeArguments */ word that appears in dart/).
  • members — methods and fields in one search; callees — the other direction of the call graph, with columns identical to callers so the two read side by side.
  • pp / objs / stubs — the object layer becomes queryable. Each renders through the same function that writes text/pp.txt, text/objs.txt and text/stubs.txt, so "what the artifact shows" and "what the query returns" cannot disagree.
  • decompile <bin> — decompile without writing any other artifact, to stdout by default. This is the only way to pipe a whole app's pseudocode; a full export requires an out_dir.
  • Scope filters — --exclude-lib PATTERN, --no-sdk (URLs starting with dart:), --app (also package:flutter). Decided by the library's original URL, not by guessing from the mangled name: dart:core mangles to dart_core, and a package named dart_core_extra would look the same. On a real Flutter app the three levels measure 505 libraries / 15,796 functions → 489 / 11,016 (--no-sdk) → 56 / 765 (--app).

Function names now also accept the all-dots lib.Class.method form that callers/callees/findrefs/call_edges.txt print, so one command's output feeds straight into the next. That was broken, not merely missing: piping findrefs's from column into dae disasm reported "nothing matched".

Query latency on a 15,796-function app: info 53 ms, objs 43 ms, pp 59 ms, stubs 62 ms, members 68 ms, getclass 82 ms, decompile --app 185 ms, findrefs 209 ms, callees 311 ms.

Correctness fixes

  • pp.txt no longer prints a fabricated constant. Its first line was hardcoded pool heap offset: 0x10f000080, carried over from the Python reference implementation — which also hardcodes it. blutter computes that value as pool_addr − heap_base; dae parses the snapshot stream and has neither the pool's image address nor a heap base. The same constant printed for macOS and Android, compressed and uncompressed pointers, so it was wrong for at least some of them. It now reads unavailable with the reason, and a gate fails if a number ever comes back. This is the one place v0.1.8's output differs from v0.1.7's.
  • dae info exits 1 on parse drift instead of printing the drift and returning 0, and its drift message no longer prints a wrapped u64 class id (18446744073709551553 for −63).
  • Usage errors exit 2, matching the documented convention (0 ok / 1 runtime error / 2 usage error). They used to exit 1, conflating the two. The message is clap's precise diagnosis plus one bilingual line pointing at dae help — clap has no i18n, so that is the honest trade-off.

Output compatibility

Across all 11 commits since v0.1.7, the whole artifact tree is byte-identical to v0.1.7 except that one pp.txt line. Verified with diff -rq on both a Mach-O arm64 Flutter app (1,011 files including all 489 under asm/ and 505 under dart/) and a real Android compressed-pointer build, and the export summary — every count, including the decompiler's blocks/statements/structured/unstructured/unmapped figures — matches item for item. dae export <bin> <out> and the dae <bin> <out> shortcut produce identical trees.

Known limitation: the superclass chain is wrong

Documented now rather than newly broken. Analyzer::parent_of does not resolve superclasses correctly. Measured against an app's own source, which is on disk:

source says dae resolves
App extends StatefulWidget (cid 2285) SceneBuilder (cid 1142)
BrightnessButton extends StatelessWidget (cid 2115) ParagraphBuilder (cid 1057)
_AppState extends State<App> _MixinApplication163&…

The class names and libraries in the same output are correct, so cid → name is fine; only the super hop is wrong. Probing all 13 Class-cluster refs found that only positions 9 and 11 resolve through type_cids at all, and 11 is also wrong — so either the superclass is not among those refs, or the Type cluster's (flags >> 4) & cid_tag_mask decode is. The second is the likelier suspect, because it would explain getting a valid cid that is systematically the wrong one.

Two shipped artifacts consume this chain and are therefore unreliable, both now marked at their source sites: frida.js's sid field (measured: {id:2316,name:"App",…,sid:1142} where the real super is 2285) and the ancestor grouping inside text/objs.txt (the field values themselves are readable; the inheritance grouping is not). Neither was silently changed — altering frida.js bytes means re-validating 25 regression archives, which is its own piece of work.

A hierarchy command was built and then withdrawn: it produced the wrong chain, and a wrong inheritance chain is worse than none. The full diagnosis is left in src/cli.rs so nobody has to redo it.

Also deliberately not provided: findrefs field (compiled code carries no symbolic field reference, only a bare displacement, so matching on displacement would report unrelated [x, #0x18] as hits — guessing, not querying), findrefs type (a pool entry's description is Kind: content, and that prefix is the object category, not a type name), and a manifest-derived --app package name (a Dart snapshot has no manifest).

Still open

  • The superclass chain, above — fixing it needs the Dart SDK's Class::Serialize / AbstractType::Serialize for ref order and flags encoding, then re-baselining frida.js and the regression archives.
  • Object-pool names are still absent from the IDA/r2 scripts (blutter emits ~52,700 pp.* flags). The data exists — text/pp.txt carries 122,064 resolved entries on a real app, and findrefs now queries it — it is just not projected into the tool scripts.
  • text/pp.txt does not escape values, so a pool string containing a real newline spills across two lines and breaks the file's own one-entry-per-line format (text/strings.txt escapes; pp.txt does not). findrefs escapes correctly and is the more correct of the two. Fixing pp.txt changes artifact bytes and therefore the archives.
  • Dart 2.18.1 remains unusable (registered in KNOWN_COLLAPSED).

Verified

54 tests under DAE_REQUIRE_GATES=1, plus every ignored release gate · dart_valid full scorecard: 26 corpus artifacts, 0 dart analyze errors · the new decompile verb's output is byte-identical to export --decompile's dart/, and both the full (505 files) and --app (56 files) outputs analyze at 0 errors · app_truth against real demo source: classes 98.8% / 100%, literals 95.4% / 97.4%, source-file→library 18/18 and 21/23 — unchanged from the recorded baseline · regress_all 25/25 byte-identical · check_profiles 47/47 · clippy 0.

The new gate tests/cli_query.rs cross-checks every findrefs hit two ways that share no code with the thing being tested: the value column must equal text/pp.txt's line verbatim (written by a different exporter), and at+offset must be literally visible in dae disasm output (a third path). Sharing the decompiler's own offset resolver would prove only self-consistency, so it is not the criterion. Measured: 679 hits, 678 pass both, 1 is the pp.txt newline wart above and is verified by address alone. Negative-tested: adding 8 to a reported offset fails all 679. The scope filters are asserted with negative controls — an "no dart_* files after --no-sdk" check passes vacuously if the flag was never wired up, so the gate also asserts the files are there without it.

Install

cargo install dae-rs          # crate name is dae-rs; the binary is `dae`
brew install ejfkdev/tap/dae
scoop bucket add ejfkdev https://github.com/ejfkdev/scoop-bucket; scoop install dae

Or grab a binary for Linux / macOS / Windows (x64 and arm64) below.