Repository navigation
English
Highlights
Large Images No Longer OOM — mmap-backed FEC
The fec tool now memory-maps ordinary input files instead of loading
the whole image into a private heap copy. A 4 GiB image that used to
be a guaranteed out-of-memory kill now completes in ~45 s with an RSS
peak around a third of the input size, on a machine whose available
memory is smaller than the image. Non-ordinary inputs (pipes, block
devices) fall back to the previous malloc + fread path; output is
byte-identical either way.
- Before (v1.2.0):
fecread the entire image into heap memory.
An 8 GiB image on an 8 GiB machine was OOM-killed ~100 s in. - After (v1.2.8):
fecmaps the file read-only. The process
peak RSS is ~33 % of the input (page-cache-backed, reclaimable),
not 1:1 with the image. Cold start (never-read input) behaves the
same; the kernel reclaims cold pages under pressure instead of
killing the process.
OpenMP FEC Encoding — up to ~4x Faster
The interleaved RS-8 column loop is now compiled with OpenMP
(-fopenmp). Columns are independent, so the encode is
schedule(static) and the output is byte-identical regardless of
thread count (verified sha256 across 1 vs 4 threads).
- 8 GiB cold image: serial ~176 s → 4 threads ~45 s (~3.9x, on a 4-core
i5). - Thread count is controlled by the standard
OMP_NUM_THREADS
environment variable; no CLI flag needed — by default the fec
tool automatically uses every logical core.
--threads N for Level-0 Hashtree Hashing
add_hashtree_footer accepts --threads N (default 1, serial, matching
upstream). For N > 1, level-0 block digests are computed in
multiprocessing workers and stitched back by block offset, bypassing
the GIL on large images. Output is byte-identical to the serial path
independent of N or scheduling.
- 128 MiB image: serial ~3.3 s →
--threads 4~2.7 s end-to-end
(the hashing stage itself runs ~3.9x faster; the overall speedup is
smaller because RSA signing and the FEC write are serial and
un-parallelisable).
Bug Fixes
--threads N>1no longer crashes. The multiprocessing branch
passedimage.name, but the image is anImageHandlerwhose
attribute is.filename. Any--threads N>1run crashed with
'ImageHandler' object has no attribute 'name'. Fixed. (The
serial path was unaffected.)--parallelrenamed to--threadsfor the CLI flag (and the
internal parameter).--parallelwas ambiguous next to the OpenMP
thread count used by the fec tool;--threadsnames it directly.
60-Byte libfec Footer
The fec tool now writes exactly sizeof(struct fec_header) (60 bytes)
after the raw parity — not a 4 KiB block. avbtool reads the last
struct.calcsize(FEC_FOOTER_FORMAT) = 60 bytes, so the footer now
sits at the correct byte offset and verification passes.
Nuitka-Only CI
The build matrix is now Nuitka-only (amd64 + arm64, 2 jobs).
PyInstaller was dropped. CI triggers on push to both main and
edge.
Compatibility Notes
--threads N(was--parallel N) is the only intentional CLI
deviation from upstreamavbtool 1.2.0. DefaultN=1preserves
upstream serial behaviour;N > 1output is byte-identical.fecis automatically multi-threaded via OpenMP; set
OMP_NUM_THREADS=1to force a single thread.- Operating requirements: unpack the tarball and put its
bin/on
PATH; the three binaries (avbtool,fec,openssl) are
static/self-contained with no runtime.sodependencies.
中文
主要更新
大镜像不再 OOM — mmap 化 FEC
fec 工具现在对普通输入文件使用内存映射(mmap),不再把整个镜像
读进私有堆内存。一个原本必然触发 OOM kill 的 4 GiB 镜像,现在在
"可用内存小于镜像"的机器上也能跑完(~45 秒,进程 RSS 峰值约为输入
的 1/3)。非普通输入(管道、块设备)自动退回原来的 malloc + fread
路径,两种路径输出逐字节一致。
- 之前(v1.2.0):
fec把整个镜像读入堆内存。8 GiB 镜像在
8 GiB 机器上会在 ~100 秒时被 OOM kill。 - 现在(v1.2.8):
fec只读映射文件。进程峰值 RSS 约输入的
33%(页缓存支撑、可被内核回收),不再与镜像大小 1:1。冷启动
(从未读过的文件)表现一致,内存紧张时内核回收冷页,而不是杀进程。
OpenMP FEC 编码 — 最高约 4 倍速
交织的 RS-8 列循环现在用 OpenMP(-fopenmp)编译。各列彼此独立,
采用 schedule(static),输出与线程数无关、逐字节一致(1 vs 4
线程已用 sha256 验证)。
- 8 GiB 冷启动镜像:串行 ~176 秒 → 4 线程 ~45 秒(~3.9x,4 核 i5)。
- 线程数由标准
OMP_NUM_THREADS环境变量控制,无需 CLI 参数 —
默认 fec 工具自动使用全部逻辑核心。
--threads N 加速 Level-0 Hashtree 哈希
add_hashtree_footer 接受 --threads N(默认 1,串行,与上游一致)。
N > 1 时用 multiprocessing worker 计算 level-0 块摘要,再按块偏移
拼回,绕过 GIL 处理大镜像。输出与串行路径逐字节一致,与 N 和调度
顺序无关。
- 128 MiB 镜像:串行 ~3.3 秒 →
--threads 4~2.7 秒(全流程;哈希段
本身约 3.9 倍快,因 RSA 签名和 FEC 写盘不可并行,全流程加速较小)。
缺陷修复
--threads N>1不再崩溃。 multiprocessing 分支原来传了
image.name,但 image 是ImageHandler,属性叫.filename。任何
--threads N>1的运行都会崩
('ImageHandler' object has no attribute 'name')。已修复
(串行路径本来就没问题)。--parallel改名为--threads(CLI 参数及内部参数同步)。
--parallel与 fec 工具的 OpenMP 线程数语义混淆,--threads
更准确。
60 字节 libfec Footer
fec 工具现在在原始 parity 后精确写入 sizeof(struct fec_header)
(60 字节)—— 不再是 4 KiB 块。avbtool 读取最后
struct.calcsize(FEC_FOOTER_FORMAT) = 60 字节,footer 现在位于正确
字节偏移,校验通过。
Nuitka-only CI
构建矩阵改为 Nuitka-only(amd64 + arm64,2 个 job),移除
PyInstaller。CI 在 main 和 edge 两个分支的 push 上都触发。
兼容性说明
--threads N(原--parallel N)是 CLI 相对上游
avbtool 1.2.0的唯一有意偏离。默认N=1保持上游串行行为;
N > 1输出逐字节一致。fec默认通过 OpenMP 自动多线程;设OMP_NUM_THREADS=1可强制
单线程。- 运行要求:解包 tarball 并把其
bin/加到PATH;三个二进制
(avbtool、fec、openssl)静态自包含,无运行时.so依赖。