Skip to content

How It Works

CodingJeffRoblox edited this page Sep 23, 2026 · 1 revision

How It Works

A technical look at the recovery engine's architecture, for anyone reading the code or evaluating what ByteRescue's results actually mean. See Project Structure for the file layout this describes.

GUI vs. engine

byterescue/recovery/ contains the entire recovery engine — carving, validation, scanning, filesystem parsing — as plain Python importable without Tkinter or a display. byterescue/app.py is GUI plumbing only: it wires user actions to that engine and renders results. This split is what makes most of the test suite (Testing) runnable headless.

Reading the source: mmap vs. chunked

byterescue/recovery/scanner.py uses two different strategies depending on the source, for a specific reason each:

  • mmap (_scan_mmap) — used for an ordinary file or disk-image file. The whole file is one addressable range, so there's no chunk-boundary problem to solve at all. Simplest and fastest option when available.
  • ChunkedReader (_scan_chunked) — used for Deep / Raw Scan mode, meant to also work against a raw physical-drive path. Windows device paths generally don't report a usable size via stat(), so mmap isn't reliable there. Instead, fixed-size chunks (DEFAULT_CHUNK_SIZE = 8 MB) are read with a 1 MB overlap buffer (DEFAULT_OVERLAP), so a signature split across a chunk boundary is still complete within at least one window. A format whose true end lies beyond the current window (rare — a huge embedded video/archive) is reported as boundary-truncated rather than silently guessed at.

Two families of signatures

byterescue/recovery/signatures.py's SIGNATURES table splits formats into two families, handled differently by carve():

  • Weak/short magics (ZIP, RIFF, SQLite, BMP, PE, ELF, MP4) get a real structural check. If that check fails, the candidate is rejected outright — not emitted at all. A failed structural check on a weak magic is good evidence it was a coincidental byte match, not a damaged real file.
  • Strong/long magics (JPEG, PNG, PDF, GZIP, 7Z, FLAC, GIF, TIFF, RAR, OLE, MKV, ID3) are accepted even without a confirmed end, since the magic itself is distinctive enough to be worth keeping — but they're tagged unverified/capped rather than claimed complete.

This is why Recovery Signatures marks some formats "Verified when accepted" (weak-magic family — rejection already filtered out the noise) versus "Verified"/"Heuristic"/"Unverified" for the strong-magic family (accepted on sight, but end-offset confidence varies).

Text recovery

Plain text has no magic header, so text_recovery.py doesn't carve by signature at all — it scans for long, plausible runs of readable ASCII/UTF-8/UTF-16 LE/UTF-16 BE and scores each run's confidence, purely from pattern shape.

Filesystem-aware recovery

filesystem.py is the one mode that reads real filesystem metadata instead of raw bytes: it parses a FAT12/16/32 boot sector/BPB directly, walks directories recursively (short 8.3 and VFAT long-filename entries), and reads deleted directory entries whose data hasn't been overwritten. See Supported File Systems for why multi-cluster deleted files are lower confidence here.

Destination safety

Before writing any recovered file, the scanner checks whether the destination could overwrite the very data being recovered: same path, same parent folder, or (for a raw device-path source) the same physical drive as the destination. It warns rather than silently blocking — you may have a deliberate reason, like recovering to a different partition on the same physical disk.

Device paths are never passed through pathlib.Path()

A historical bug (fixed in 0.3.0 — see Changelog) is guarded against directly in code now: pathlib.Path() parses a Windows device-namespace path like \\.\PhysicalDrive2 as a UNC path and re-serializes it with a trailing backslash, which Windows then refuses to open. Device paths are kept as plain strings everywhere in the scanner instead.

Structural validation, separately from carving

Every carved candidate that gets far enough to be reported also goes through a separate validation pass (validate()) using a real parser for its format — Pillow for images, zipfile for ZIP, sqlite3's PRAGMA integrity_check for SQLite, the wave module for WAV, decompression for GZIP — and is labeled Passed / Partial / Failed / Unknown. A signature match and a validation pass answer different questions: carving asks "does this look like a <format> at this offset," validation asks "can a real parser actually open what we carved."

Related pages

Recovery Center · Recovery Signatures · Adding a New Signature · Testing

Clone this wiki locally