Skip to content

v0.1.7

Choose a tag to compare

@github-actions github-actions released this 27 Sep 16:14
· 34 commits to main since this release
v0.1.7
158f703

The export side was silently dropping 86% of the function names it could have emitted. This release fixes that, and cuts peak memory on the export path by 9%.

Recovered names: 6× more

entry_for rejected any function whose instructions-table index fell below first_entry_with_code, on the theory that those entries were dispatch stubs with no function body. That reading was wrong. The SDK is explicit about the field — it is "the first Instructions object which is going to have Code object associated with it", recorded so the runtime can "reduce the binary search space when searching specifically for the code object". Entries below it belong to discarded Code objects: under dwarf_stack_traces_mode Dart drops the Code wrapper to save space, but the machine code stays in the image — the program has to run — and the serializer merely omits their payload_info.

Every real shipping app builds with dwarf stack traces on, so this hit all of them:

artifact table entries named functions libraries classes
Reqable (android) 57,960 2,164 → 13,371 496 → 1,734 1,141 → 3,567
Lark 飞书 79,327 5,262 → 25,183 1,418 → 3,194 2,868 → 6,306
Reqable.app (macOS) 70,996 2,358 → 17,319 564 → 1,955 1,285 → 4,262
ChatGLM 30,782 → 27,517 1,211 4,603
CHSI 学信网 19,752 → 17,438 875 3,256
Weibo 微博 22,623 → 19,807 750 3,671

All at 0 warnings. The IDA and radare2 scripts are what actually felt this: for Android Reqable, addNames.py goes 1,766 → 11,296 lines and addNames.r2 7,602 → 37,344. Class and library counts now land at or above aotopsy's on the same artifacts (CHSI 3,256 vs 3,819 classes; Weibo 3,671 vs 4,232).

Address validity was checked independently, not assumed. Of the 9,530 newly reachable entries on Android Reqable, 95.5% begin with stp x29, x30, [x7, #-16]! (Dart's arm64 prologue) against 82.7% for the already-validated set, 98.2% of their first instructions also occur in the validated set, and none is out of bounds. Decompiler output is unchanged apart from improving: Lark still emits exactly 3,517 function blocks / 20,335 blocks / 110,881 statements / 3,374 structured / 1 unmapped line, with named direct calls up 6,974 → 7,130.

No regression risk on the existing corpus: every hello sample builds with no-dwarf_stack_traces_mode, so their first_entry_with_code is 0 and the code path is untouched — regress_all stays 25/25 byte-identical and ground_truth stays 6,963/6,963 addresses at 89.4% naming.

Memory: export peak −9%

Every text/ file and both tool scripts were built as one String and then handed to fs::write, with capacity guessed as rows × bytes-per-row. Guessing high wastes resident memory; guessing low reallocs and memmoves repeatedly. They are now streamed through a BufWriter, so peak holds one 8 KB buffer per file instead of the whole file — and the per-row byte-count constants are gone rather than maintained.

Measured on Reqable.app (26 MB), old and new binaries back to back:

plain export peak RSS   163 -> 148 MB, and again 161 -> 146 MB     (-15 MB, -9%)
whole output tree       byte-identical (diff -rq over every file)
--decompile output      byte-identical

The --decompile peak did not measurably move: base ranged 172–196 MB and streamed 184–191 MB over four runs, so the spread exceeds the effect. That path's peak is dominated by the decompiler's own structures, not these buffers. An earlier "172 vs 191" reading was noise, not a regression — said plainly because the honest result is one path improved, the other did not.

Still open, and why

  • Decompiler wall-clock is not improved in this release, and no speedup is claimed. Profiling a debug-symbol build attributes it to format_inner, Formatter::pad/pad_integral, RawVecInner::finish_grow and _platform_memmove — i.e. roughly one format!("... // {addr:#x}") per statement, ~1M of them. Two allocation reductions aimed at that (replace_word slice copies, pre-sized buffers) were verified byte-identical but produced no measurable gain; the fix that would is converting render_op's call sites from format! to write! into reused buffers, which is a real refactor rather than a tweak. Timing on the verification host was also unusable — load average 16, the same binary varying 35% between runs.
  • Object-pool names are still absent from the IDA/r2 scripts (blutter emits ~52,700 pp.* flags). dae has the data — text/pp.txt carries 122,064 resolved entries for Reqable.app — it just is not projected into the tool scripts yet. This is the remaining export-side gap.
  • Dart 2.18.1 remains unusable (registered in KNOWN_COLLAPSED with the measurement showing both candidate layouts are unhealthy).

Verified

48 tests under DAE_REQUIRE_GATES=1 · full scorecard 26 artifacts / 0 dart analyze errors · regress_all 25/25 byte-identical · check_profiles 47/47 · app_truth 98.8%/100% class recovery against real demo source · clippy 0.

Install

cargo install dae-rs          # crate name is dae-rs; the binary is `dae`
brew install ejfkdev/tap/dae
scoop bucket add ejfkdev https://github.com/ejfkdev/scoop-bucket; scoop install dae

Or grab a binary for Linux / macOS / Windows (x64 and arm64) below.