Skip to content

AVBTool MOD v1.2.8

Latest

Choose a tag to compare

@ELF-RC ELF-RC released this 01 Oct 06:35
· 2 commits to main since this release
6e2e83f

English

Highlights

Large Images No Longer OOM — mmap-backed FEC

The fec tool now memory-maps ordinary input files instead of loading
the whole image into a private heap copy. A 4 GiB image that used to
be a guaranteed out-of-memory kill now completes in ~45 s with an RSS
peak around a third of the input size, on a machine whose available
memory is smaller than the image. Non-ordinary inputs (pipes, block
devices) fall back to the previous malloc + fread path; output is
byte-identical either way.

  • Before (v1.2.0): fec read the entire image into heap memory.
    An 8 GiB image on an 8 GiB machine was OOM-killed ~100 s in.
  • After (v1.2.8): fec maps the file read-only. The process
    peak RSS is ~33 % of the input (page-cache-backed, reclaimable),
    not 1:1 with the image. Cold start (never-read input) behaves the
    same; the kernel reclaims cold pages under pressure instead of
    killing the process.

OpenMP FEC Encoding — up to ~4x Faster

The interleaved RS-8 column loop is now compiled with OpenMP
(-fopenmp). Columns are independent, so the encode is
schedule(static) and the output is byte-identical regardless of
thread count
(verified sha256 across 1 vs 4 threads).

  • 8 GiB cold image: serial ~176 s → 4 threads ~45 s (~3.9x, on a 4-core
    i5).
  • Thread count is controlled by the standard OMP_NUM_THREADS
    environment variable; no CLI flag needed — by default the fec
    tool automatically uses every logical core.

--threads N for Level-0 Hashtree Hashing

add_hashtree_footer accepts --threads N (default 1, serial, matching
upstream). For N > 1, level-0 block digests are computed in
multiprocessing workers and stitched back by block offset, bypassing
the GIL on large images. Output is byte-identical to the serial path
independent of N or scheduling.

  • 128 MiB image: serial ~3.3 s → --threads 4 ~2.7 s end-to-end
    (the hashing stage itself runs ~3.9x faster; the overall speedup is
    smaller because RSA signing and the FEC write are serial and
    un-parallelisable).

Bug Fixes

  • --threads N>1 no longer crashes. The multiprocessing branch
    passed image.name, but the image is an ImageHandler whose
    attribute is .filename. Any --threads N>1 run crashed with
    'ImageHandler' object has no attribute 'name'. Fixed. (The
    serial path was unaffected.)
  • --parallel renamed to --threads for the CLI flag (and the
    internal parameter). --parallel was ambiguous next to the OpenMP
    thread count used by the fec tool; --threads names it directly.

60-Byte libfec Footer

The fec tool now writes exactly sizeof(struct fec_header) (60 bytes)
after the raw parity — not a 4 KiB block. avbtool reads the last
struct.calcsize(FEC_FOOTER_FORMAT) = 60 bytes, so the footer now
sits at the correct byte offset and verification passes.

Nuitka-Only CI

The build matrix is now Nuitka-only (amd64 + arm64, 2 jobs).
PyInstaller was dropped. CI triggers on push to both main and
edge.

Compatibility Notes

  • --threads N (was --parallel N) is the only intentional CLI
    deviation from upstream avbtool 1.2.0. Default N=1 preserves
    upstream serial behaviour; N > 1 output is byte-identical.
  • fec is automatically multi-threaded via OpenMP; set
    OMP_NUM_THREADS=1 to force a single thread.
  • Operating requirements: unpack the tarball and put its bin/ on
    PATH; the three binaries (avbtool, fec, openssl) are
    static/self-contained with no runtime .so dependencies.

中文

主要更新

大镜像不再 OOM — mmap 化 FEC

fec 工具现在对普通输入文件使用内存映射(mmap),不再把整个镜像
读进私有堆内存。一个原本必然触发 OOM kill 的 4 GiB 镜像,现在在
"可用内存小于镜像"的机器上也能跑完(~45 秒,进程 RSS 峰值约为输入
的 1/3)。非普通输入(管道、块设备)自动退回原来的 malloc + fread
路径,两种路径输出逐字节一致。

  • 之前(v1.2.0):fec 把整个镜像读入堆内存。8 GiB 镜像在
    8 GiB 机器上会在 ~100 秒时被 OOM kill。
  • 现在(v1.2.8):fec 只读映射文件。进程峰值 RSS 约输入的
    33%(页缓存支撑、可被内核回收),不再与镜像大小 1:1。冷启动
    (从未读过的文件)表现一致,内存紧张时内核回收冷页,而不是杀进程。

OpenMP FEC 编码 — 最高约 4 倍速

交织的 RS-8 列循环现在用 OpenMP(-fopenmp)编译。各列彼此独立,
采用 schedule(static),输出与线程数无关、逐字节一致(1 vs 4
线程已用 sha256 验证)。

  • 8 GiB 冷启动镜像:串行 ~176 秒 → 4 线程 ~45 秒(~3.9x,4 核 i5)。
  • 线程数由标准 OMP_NUM_THREADS 环境变量控制,无需 CLI 参数 —
    默认 fec 工具自动使用全部逻辑核心。

--threads N 加速 Level-0 Hashtree 哈希

add_hashtree_footer 接受 --threads N(默认 1,串行,与上游一致)。
N > 1 时用 multiprocessing worker 计算 level-0 块摘要,再按块偏移
拼回,绕过 GIL 处理大镜像。输出与串行路径逐字节一致,与 N 和调度
顺序无关。

  • 128 MiB 镜像:串行 ~3.3 秒 → --threads 4 ~2.7 秒(全流程;哈希段
    本身约 3.9 倍快,因 RSA 签名和 FEC 写盘不可并行,全流程加速较小)。

缺陷修复

  • --threads N>1 不再崩溃。 multiprocessing 分支原来传了
    image.name,但 image 是 ImageHandler,属性叫 .filename。任何
    --threads N>1 的运行都会崩
    ('ImageHandler' object has no attribute 'name')。已修复
    (串行路径本来就没问题)。
  • --parallel 改名为 --threads(CLI 参数及内部参数同步)。
    --parallel 与 fec 工具的 OpenMP 线程数语义混淆,--threads
    更准确。

60 字节 libfec Footer

fec 工具现在在原始 parity 后精确写入 sizeof(struct fec_header)
(60 字节)—— 不再是 4 KiB 块。avbtool 读取最后
struct.calcsize(FEC_FOOTER_FORMAT) = 60 字节,footer 现在位于正确
字节偏移,校验通过。

Nuitka-only CI

构建矩阵改为 Nuitka-only(amd64 + arm64,2 个 job),移除
PyInstaller。CI 在 main 和 edge 两个分支的 push 上都触发。

兼容性说明

  • --threads N(原 --parallel N)是 CLI 相对上游
    avbtool 1.2.0 的唯一有意偏离。默认 N=1 保持上游串行行为;
    N > 1 输出逐字节一致。
  • fec 默认通过 OpenMP 自动多线程;设 OMP_NUM_THREADS=1 可强制
    单线程。
  • 运行要求:解包 tarball 并把其 bin/ 加到 PATH;三个二进制
    (avbtool、fec、openssl)静态自包含,无运行时 .so 依赖。