Skip to content

Static Translation Pipeline

M T edited this page Oct 4, 2026 · 1 revision

Static Translation Pipeline

Master Chef does not emulate Halo at run time. It statically translates the user's own retail halo.exe (Halo: Combat Evolved PC 1.10) into C: every reachable x86 function becomes a C function that manipulates an explicit guest CPU state (EngineCPU) and a flat 32-bit guest address space. Those C files are generated locally on the user's machine, are never committed, and are compiled together with the hand-written host (native/EngineHost) into the macOS diagnostic runner or the visionOS app.

This page follows the whole path end to end: executable hash check, PE analysis, function entry lists, the closure walk, decoding, lifting, the generated file layout, the import table export, and how the build systems consume the output. The decoder and lifter internals are on XWA Decoder and Lifter; the C runtime the generated code targets is on EngineReuse Runtime; floating point is on x87 Floating Point; the checked-in entry lists are on Function Address Lists.

Source files

File Role
tools/generate_engine_reuse.py The generator: hash check, closure walk, jump-table harvesting, lifting, direct-call rewriting, chunking, dispatch table, generation.json receipt
tools/export_engine_imports.py Emits engine_imports.c: PE entry point, image base, size, TLS directory, section table, import (IAT) table
tools/engine_reuse/upstream.py Loads the pinned third_party/xwa tools under a private package name; defines HALO_EXE lookup and the required SHA-256
tools/engine_reuse/decode.py EngineDisassembler: bounded, fail-closed recursive decoding of one function
tools/engine_reuse/lift.py EngineLifter: strict x86 to C lowering, composes the mixins below
tools/engine_reuse/flags.py, fpu.py, extra.py EFLAGS, x87 and remaining integer/string/system instruction lowering
third_party/xwa/tools/ Vendored MIT toolkit (PE analysis via pefile, Capstone disassembly, base lifter)
decompilation/c9acf0c46954/ Checked-in function entry lists for the supported executable
native/EngineReuse/ Headers every generated file includes (engine_cpu.h, engine_flags.h, engine_registers.h, engine_hooks.h)
native/EngineHost/Makefile macOS host build; can invoke both generators; compiles chunks
tools/setup_halo.py Setup wizard step that runs both generators when inputs change
tools/build_engine_vision.py visionOS direct build; validates the generated inventory and compiles it
native/EngineVision/project.yml XcodeGen project; adds the generated chunks to the app target
docs/BUILDING.md User-facing generation instructions

End-to-end flow

flowchart TD
    EXE["User-supplied halo.exe<br/>(game/halo.exe or HALO_EXE)"] --> HASH{"SHA-256 ==<br/>c9acf0c4...9545 ?"}
    HASH -- no --> ABORT["AssertionError, nothing written"]
    HASH -- yes --> PE["load_pe (pefile)<br/>sections, image base, IAT map"]
    LISTS["decompilation/c9acf0c46954/<br/>function-addresses.txt<br/>extra-function-entries.txt"] --> ENTRIES["Entry queue + known-function set"]
    PE --> ENTRIES
    ENTRIES --> DEC["EngineDisassembler.disassemble_function<br/>(Capstone, bounded, fail-closed)"]
    DEC --> SW["Jump-table harvesting<br/>(up to 4 re-decode rounds)"]
    SW --> LIFT["EngineLifter.lift_function<br/>(flags, x87, extra mixins)"]
    LIFT --> REC["generation.json record"]
    LIFT --> QUEUE["Enqueue direct call and tail-call targets"]
    QUEUE --> DEC
    QUEUE -.->|"--discover"| IMM["Scan immediates for code addresses<br/>(fixed point)"]
    IMM --> DEC
    LIFT --> DIRECT["direct_calls rewrite<br/>(ENGINE_DIRECT, entry fast path)"]
    DIRECT --> FILES["sub_XXXXXXXX.c per function"]
    FILES --> CHUNKS["chunk_000.c .. chunk_031.c<br/>engine_functions.h"]
    FILES --> BUNDLE["engine_bundle.c<br/>(sorted entry table + engine_dispatch)"]
    PE --> IMPORTS["export_engine_imports.py<br/>engine_imports.c"]
    CHUNKS --> CC["clang (-O2, ENGINE_FLAT_MEMORY=1,<br/>HALO_ARM64_FENV_FAST=1, ...)"]
    BUNDLE --> CC
    IMPORTS --> CC
    HOST["native/EngineHost/*.c, *.m<br/>(shims, Metal, audio, input)"] --> CC
    CC --> APP["halo-host (macOS) or<br/>HaloVision.app (visionOS)"]
Loading

The canonical invocation (identical in BUILDING.md, the Makefile and setup_halo.py):

python tools/generate_engine_reuse.py \
  @decompilation/c9acf0c46954/function-addresses.txt \
  @decompilation/c9acf0c46954/extra-function-entries.txt \
  --label whole-exe --max-functions 10000 --trap-unsupported --discover --chunks 32
python tools/export_engine_imports.py

Stage 1: locating and verifying the executable

upstream.py resolves the input path:

PE = Path(os.environ.get('HALO_EXE', str(VISION / 'game/halo.exe'))).expanduser().resolve()
PE_SHA256 = 'c9acf0c469543283cfed6d7dc04ade976dbdfc7cb4532cf070386de169c19545'
  • VISION is the repository root (two directories above tools/engine_reuse).
  • game/ is git-ignored; the user places their own executable there, or sets HALO_EXE to an absolute path. The setup wizard sets HALO_EXE to its staged copy (setup_halo.py:362).
  • generate_engine_reuse.py:79 asserts hashlib.sha256(PE.read_bytes()).hexdigest()==PE_SHA256 before anything is decoded or written. A different build (Custom Edition, a patched or cracked binary, another patch level) fails here with a bare AssertionError.
  • export_engine_imports.py does not repeat the hash check; it relies on the same PE path and stamps PE_SHA256[:12] into its output comment. Run it only after the generator has accepted the file.

The same digest is hard-coded in tools/prepare_engine_vision_device.py:35 and EngineAssets.swift:12, and its first 12 hex digits name the decompilation/c9acf0c46954 directory. Every address in this pipeline is meaningful only for that exact file.

Stage 2: PE analysis

load_pe is the vendored translator.load_pe, which calls pe_analyze.analyze_pe (built on pefile) and build_iat_map. It returns:

Value Contents
info (PEInfo) file name, image base, entry RVA, timestamp, linker version, sections (Section name/VA/virtual size/raw offset/raw size/characteristics), imports (ImportEntry dll/name/ordinal/IAT RVA), code and data VA ranges
data the raw file bytes
iat dict IAT slot VA -> (dll, name); unnamed imports become ordinal_N

A section counts as code when characteristics & 0x20000020 is non-zero (CODE or EXECUTE), excluding .bind. The disassembler only decodes inside code sections (disasm.py:132-142); read_bytes can also read data sections, which the generator uses to read jump tables and selector bytes.

load_pe prints a short summary (image base, code and data ranges, IAT entry count). export_engine_imports.py suppresses that output with contextlib.redirect_stdout.

Stage 3: entry points and the known-function set

Lines 81-87 build two collections:

  1. known – addresses treated as separate functions when decoding. If a directory decompilation/c9acf0c46954/functions/ containing *.c files exists, known is the set of their file stems parsed as hex; otherwise it is every address in function-addresses.txt. The public repository ships only the text index, so the fallback is what normally runs. All command-line entries are then added (known.update(entries)), so extra-function-entries.txt also becomes "known" when passed with @.
  2. entries – the initial work queue, from positional arguments:
    • @path reads VISION/path (relative to the repository root, not the current directory) and parses every whitespace-separated token with int(x, 16).
    • any other argument is parsed as one hex address (5590a0, 0x5590a0 both work).
    • with no positional arguments, the defaults are 5590a0 55cfd0 55efd0 with the default label player-movement, the first bounded closure the project translated.

known matters because the decoder uses it to decide where one function stops: a branch or fallthrough into another known entry is a tail call, not part of the current body. See Function Address Lists for the list formats.

Stage 4: the closure walk

The main loop (lines 88-180) is a breadth-first worklist:

pending = deque(entries); seen = {}
loop:
  while pending and len(seen) < --max-functions:
      address = pending.popleft()
      skip if seen or not in a code section
      decode -> harvest jump tables -> lift -> record
      pending += (calls_to | tail_calls) - seen
  if not --discover: stop
  new = immediate operands of every decoded instruction that are code
        addresses, not yet seen, not already an instruction start
  stop if new is empty; else known += new; pending += sorted(new)

Hand-proven switch tables

Three entries carry jump-table targets that the generic harvester cannot prove, so the generator asserts the original bytes and adds the targets as extra block leaders (lines 96-113):

Entry Evidence asserted Extra leaders
0x005590A0 guard at 559B62 bounds EAX to 0..7; selector bytes at 0x559E30 must equal 00 02 00 00 01 01 01 01 (max 2) three dwords at 0x559E24
0x005061C0 compare with 8 at 50628F; selector bytes at 0x506434 must equal 00 01 02 02 02 02 01 01 01 three dwords at 0x506428
0x005067B0 table at 0x506F9C must equal four code pointers followed by CC CC CC CC four dwords at 0x506F9C

If the bytes differ, the assertion stops generation, a second guard against a wrong executable.

Decoding

Each address goes to EngineDisassembler.disassemble_function(address, iat, extra_entries=extra, known_functions=known) (decode.py:27-130). The decoder follows explicit edges only, records indirect calls and jumps as indirect_edges instead of guessing targets, treats branches into other known functions as tail_calls, makes every call return site a block leader, rejects overlapping decodes, and fails if one function exceeds 50,000 instructions. Any exception becomes a record {"entry": ..., "undecodable": "<message>"} and the function is skipped. Details: XWA Decoder and Lifter.

Jump-table harvesting

After the first decode, up to four rounds look at every indirect jump of the exact shape jmp dword ptr [index*4 + disp] (no base register, scale 4, positive displacement) (lines 118-156):

  1. switch_case_count walks back at most 16 linear predecessors (via an end_address -> instruction map, stopping at calls, returns and unconditional jumps) looking for the MSVC guard shapes:
    • cmp idx, N ; ja default ; jmp [idx*4+T] gives N+1 entries;
    • cmp sel, N ; ja default ; movzx idx, byte ptr [sel+B] ; jmp [idx*4+T] gives max(B[0..N]) + 1 entries. N must be below 4096. Any other write to the index register between the compare and the jump hides the guard and the count is 0 (unproven).
  2. The plausible target range is [function entry, next known entry after max(entry, function.end-1)), or entry + 0x10000 when there is no later known entry. The code comment explains why the bound uses the decoded extent: alternate entries into one body are listed as separate functions (the script interpreter 424C20 has 424C4E).
  3. Table entries are read as little-endian dwords. With a proven count, implausible pointers and pointers to other known functions are skipped and reading continues; without a proven count, reading stops at the first implausible pointer (at most 4096 entries).
  4. New targets are added as extra_entries and the function is re-decoded. Rounds stop when nothing new appears or a re-decode fails (the previous decode is kept).

Harvested tables are listed in the record's jumpTables. Targets that were not harvested still work at run time: the lifted indirect jump falls back to engine_dispatch (see Stage 6).

Lifting and the record

Lines 157-170 lift with EngineLifter(iat_map=iat, trap_unsupported=args.trap_unsupported, static_functions=args.chunks<=1) and store a per-function record:

Record key Meaning
entry hex entry address
instructions, blocks decoded instruction and basic-block counts
calls, tailCalls direct call targets and out-of-function branch targets (hex)
indirectEdges [address, "call" or "jump", operand text] for every indirect transfer
jumpTables table base addresses harvested for this function
generatedSHA256 SHA-256 of the lifted C (before the direct-call rewrite)
instructionTraps with --trap-unsupported, {address, reason} for every instruction replaced by a runtime trap
unsupported without --trap-unsupported, the first lowering error; the function is not generated
undecodable decode error; the function is not generated

Then every direct call target and tail-call target not yet seen is appended to pending.

--discover

Without --discover, the walk only follows direct call/jmp/jcc targets. Function pointers stored as instruction immediates (callbacks passed to push imm32, mov [mem], imm32, vtables built in code) would be missed. With --discover (lines 171-180), once the queue drains, every immediate operand of every decoded instruction is checked: if it is a code-section address, not yet visited, and not already the start of a decoded instruction, it is added to known and the queue. This repeats until no new address appears. Immediates that are not really code can produce undecodable records; that is expected noise.

Pointers that live only in data sections (for example vtables in .rdata) are not scanned. Those functions must come from the entry lists, which is one reason extra-function-entries.txt exists.

--max-functions

len(seen) counts every address taken off the queue that was in a code section (including ones that later fail to decode). When it reaches the cap, remaining queued addresses are written to frontier in the receipt instead of being translated. The value must satisfy 0 < N <= 20000 (line 77); the whole-executable build uses 10000.

Stage 5: output layout

Output goes to native/build/engine-reuse/<label>/ (line 89). The label must be alphanumeric apart from hyphens (line 78). native/build/ is git-ignored.

File Written by Contents
sub_XXXXXXXX.c generator, one per translated function the lifted function after direct_calls rewriting (upper-case hex, 8 digits)
chunk_000.c .. chunk_NNN.c generator, only when --chunks N > 1 #includes of engine_cpu.h, engine_flags.h, engine_functions.h, engine_registers.h, then about ceil(functions / N) consecutive sub_*.c files in address order
engine_functions.h generator, only when --chunks > 1 prototypes of engine_dispatch, engine_dispatch_external, engine_dispatch_override, engine_record, every sub_*, then #include "engine_hooks.h"
engine_bundle.c generator, always sorted engine_entries[], parallel engine_fns[], ENGINE_FN_COUNT, engine_dispatch, weak default hooks, engine_reuse_entry; in single-unit mode it also #includes every sub_*.c
generation.json generator the receipt (see below)
engine_imports.c export_engine_imports.py PE constants, section table and import table for the host loader
host-obj/ native/EngineHost/Makefile macOS host object files (not a generator output)

--chunks splits the generated functions across N translation units (its help text: "Split generated functions across N translation units (functions become non-static)"); the build systems compile each chunk as a separate object, which lets them build in parallel. In chunked mode generated functions are non-static (static_functions=False) because they call each other across units; in single-unit mode (--chunks 1, the default) everything, including engine_dispatch, is static inside engine_bundle.c.

Before writing chunks, the generator deletes any chunk_*.c not in the new set (lines 185-189), because the build systems glob chunk_*.c and a stale chunk would define functions twice. Stale sub_*.c files are not deleted; they are harmless because only chunk files and the bundle include them.

Single-unit vs chunked engine_bundle.c

flowchart LR
    subgraph ONE ["--chunks 1 (default)"]
      B1["engine_bundle.c"] --> I1["include engine_cpu.h, engine_flags.h"]
      B1 --> S1["static engine_dispatch prototype"]
      B1 --> H1["include engine_hooks.h, engine_registers.h"]
      B1 --> P1["static sub_* prototypes"]
      B1 --> F1["include every sub_*.c"]
      B1 --> T1["entry table + static engine_dispatch"]
    end
    subgraph MANY ["--chunks 32 (whole-exe)"]
      C["chunk_000.c .. chunk_031.c"] --> FH["engine_functions.h"]
      C --> SUB["include ~1/32 of sub_*.c"]
      B2["engine_bundle.c"] --> FH
      B2 --> T2["entry table + non-static engine_dispatch"]
    end
Loading

The dispatch function

Both modes end engine_bundle.c with the same body (lines 220-231), shown here reformatted:

void engine_dispatch(EngineCPU *cpu, uint32_t address) {
    if (address >= 0xFE000000u) {                       /* host import / magic handle */
        if (engine_dispatch_external(cpu, address)) return;
        cpu->pc = address; engine_fail(cpu, "unresolved import boundary");
    }
    if (engine_dispatch_override(cpu, address)) return; /* host hooks and overrides */
    cpu->pc = address;
    engine_record(address);                             /* crash-dump call ring */
    int lo = 0, hi = ENGINE_FN_COUNT - 1, idx = -1;     /* greatest entry <= address */
    while (lo <= hi) { int mid = (lo + hi) >> 1;
        if (engine_entries[mid] <= address) { idx = mid; lo = mid + 1; } else hi = mid - 1; }
    if (idx >= 0) { engine_fns[idx](cpu); return; }
    if (engine_dispatch_external(cpu, address)) return;
    engine_fail(cpu, "unresolved original engine boundary");
}
__attribute__((weak)) int engine_dispatch_external(EngineCPU*, uint32_t) { return 0; }
__attribute__((weak)) int engine_dispatch_override(EngineCPU*, uint32_t) { return 0; }
__attribute__((weak)) void engine_record(uint32_t) {}
void engine_reuse_entry(EngineCPU *cpu, uint32_t entry) { engine_dispatch(cpu, entry); }

The search returns the function whose entry is the greatest one at or below the target, not only exact matches. That is deliberate: a dispatch to an address inside a function (a call return site after a non-local return, an unharvested switch case) enters that function's L_ENTRY switch, which jumps to the matching block or fails with dispatch to non-leader address. The weak defaults let the bundle link alone (tests, the sandboxed runtime); the host supplies strong definitions. See EngineReuse Runtime for the full dispatch model.

The receipt: generation.json

Lines 233-244:

Key Meaning
sourceExecutableSHA256 the verified digest
seconds wall time
functionsVisited len(seen)
functionsGenerated functions with C output
generatedInstructions instructions in generated functions
instructionTraps total instructions replaced by runtime traps
frontier queued but never visited (cap reached)
functions every per-function record
scope "Static generated code, unresolved functions trap explicitly. Not execution or gameplay evidence."

On stdout the generator prints the receipt without functions/frontier, then Unresolved function entries, an Unsupported histogram keyed by the error text before the first :, the number of undecodable entries, and the number discovered via immediates. BUILDING.md tells users to inspect this file for generation errors.

Stage 6: generated C shape

Every function has the same skeleton. This example was produced by running EngineLifter and direct_calls on a small hand-assembled function (it is illustrative, not code from Halo):

void sub_00401000(EngineCPU *cpu) {
if (cpu->pc == 0x00401000u) goto L_00401000;          /* chunked mode: entry fast path */
L_ENTRY: switch (cpu->pc) {                             /* resume at any block leader */
case 0x00401000u: goto L_00401000;
case 0x00401008u: goto L_00401008;
case 0x0040100Eu: goto L_0040100E;
/* ... */
default: engine_fail(cpu, "dispatch to non-leader address"); }
L_00401000:;
engine_step(cpu,0x00401000u); /* mov eax, dword ptr [esp + 4] */
eax = engine_read_u32(cpu, esp + 0x4); /* 0x00401000: mov eax, dword ptr [esp + 4] */
engine_step(cpu,0x00401004u); /* test eax, eax */
(void)engine_flags_logic(&eflags, (eax) & (eax), 32u);
engine_step(cpu,0x00401006u); /* je 0x40101e */
if (((eflags & ENGINE_EFLAGS_ZF) != 0u)) { goto L_0040101E; }
goto L_00401008;
L_00401008:;
engine_step(cpu,0x00401008u); /* push eax */
engine_push(cpu, eax, 4);
engine_step(cpu,0x00401009u); /* call 0x401100 */
{ engine_push(cpu, 0x0040100Eu, 4);
ENGINE_DIRECT(cpu, 0x00401100u, sub_00401100);
if (cpu->pc != 0x0040100Eu) { if (cpu->pc >= 0x00401000u && cpu->pc < 0x0040101Fu) goto L_ENTRY; else return; } }
goto L_0040100E;
L_0040100E:;
engine_step(cpu,0x0040100Eu); /* add esp, 4 */
esp = engine_flags_add(&eflags, esp, 4, 0u, 32u);
engine_step(cpu,0x00401011u); /* call dword ptr [0x402000] */
{ uint32_t target = engine_read_u32(cpu, 0x402000);
engine_push(cpu, 0x00401017u, 4);
engine_dispatch(cpu, target);
if (cpu->pc != 0x00401017u) { if (cpu->pc >= 0x00401000u && cpu->pc < 0x0040101Fu) goto L_ENTRY; else return; } }
/* ... */
L_0040101E:;
engine_step(cpu,0x0040101Eu); /* ret  */
cpu->pc = engine_pop(cpu, 4);
esp += 0u;
return;
}

Points to notice:

  • Register names (eax, esp, eflags) are macros over cpu->gpr[]/cpu->flags from engine_registers.h.
  • Every instruction is preceded by engine_step(cpu, pc), which in release builds only stores cpu->pc (so faults, shims and backtraces know where the guest is).
  • The guest stack is real guest memory: calls push a real return address and ret pops it into cpu->pc. The caller then checks cpu->pc against its expected return site. See EngineReuse Runtime.
  • Direct calls to translated entries become ENGINE_DIRECT (see below); calls through the IAT or registers go through engine_dispatch.
  • Indirect jmp becomes a switch over every block leader of the current function with default: engine_dispatch(cpu, target); return;.

direct_calls: the post-pass

direct_calls rewrites the lifted text with regular expressions, using the set of functions actually generated:

  1. { uint32_t target = 0xXXXXXXXXu;\nengine_push(...);\nengine_dispatch(cpu, target); becomes { engine_push(...);\nENGINE_DIRECT(cpu, 0xXXXXXXXXu, sub_XXXXXXXX); when the constant target is a generated entry.
  2. engine_dispatch(cpu, 0xXXXXXXXXu); return; (tail calls and out-of-function branches) becomes ENGINE_DIRECT(cpu, 0x...u, sub_...); return; under the same condition.
  3. If the function starts with the non-static header and its own entry is a leader, an if (cpu->pc == entry) goto L_entry; line is inserted before the L_ENTRY switch. Because the header match is on void sub_..., this fast path only appears in chunked builds; single-unit builds emit static void.

The rationale is recorded in the docstring and in engine_hooks.h: on build 30, engine_dispatch lookup cost about a tenth of the engine thread, and 25,835 of 32,373 translated call sites name a fixed entry. ENGINE_DIRECT still routes hooked addresses through engine_dispatch, so host overrides keep working. A consequence is that direct calls skip engine_record, so the crash dump's call ring only holds indirect calls.

Stage 7: import and loader table export

export_engine_imports.py reads the PE header fields directly from the file bytes:

  • e_lfanew at offset 0x3C; optional header at e_lfanew + 24;
  • AddressOfEntryPoint at optional header +16, SizeOfImage at +56;
  • the TLS data directory (index 9) at +96 + 9*8.

It writes engine_imports.c:

/* Generated from halo.exe c9acf0c46954 by tools/export_engine_imports.py */
#include "host.h"
const uint32_t engine_pe_entry_point = 0x...u;
const uint32_t engine_pe_image_base = 0x...u;
const uint32_t engine_pe_size_of_image = 0x...u;
const uint32_t engine_pe_tls_directory = 0x...u;   /* 0 when the TLS directory is empty */
const EngineSection engine_pe_sections[] = { {".text", va, vsize, raw_offset, raw_size, characteristics}, ... };
const size_t engine_pe_section_count = N;
const EngineImport engine_imports[] = { {iat_va, "KERNEL32.dll", "GetTickCount"}, ... };  /* sorted by IAT address */
const size_t engine_import_count = M;

Ordinal-only imports are named #<ordinal>. The structure types are declared in host.h:12-16. The only option is --output-dir (default native/build/engine-reuse/whole-exe; a relative path is resolved against the repository root).

How the host uses it

At startup (host.c host_run), the host maps a 4 GiB flat guest arena, then:

  1. load_image copies the first 0x1000 bytes (headers) to engine_pe_image_base and each section's min(raw_size, virtual_size) bytes to its VA. The original x86 code bytes are loaded too: generated code reads data (and jump tables) from guest memory, but never executes those bytes.
  2. setup_imports writes into every IAT slot the value host_proc_address(dll, name), a magic handle 0xFF000000 | index into the host's procedure table, whose entry points at a shim when one exists.
  3. The TLS directory constant is used by setup_thread_teb to build each thread's TLS block.
  4. engine_reuse_entry(cpu, engine_pe_entry_point) starts the game, with the magic return address 0xFF000000 pushed so that the CRT entry returning is detectable.

Guest code then calls imports exactly as the original did, call dword ptr [IAT]. The lifted call reads the slot, gets 0xFF0000xx, and engine_dispatch routes any address >= 0xFE000000 to the host's engine_dispatch_external, which runs the shim. See Win32 Compatibility Layer and EngineHost Overview.

Stage 8: compiling and linking

Consumer What it compiles Flags relevant to generated code
native/EngineHost/Makefile (macOS halo-host) $(GEN)/chunk_*.c, engine_bundle.c, engine_imports.c, host sources, MojoShader -O2 -w -std=c11 -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 -I. -I native/EngineReuse -I $(GEN)
tools/build_engine_vision.py (--direct) sorted chunk_*.c, engine_bundle.c, engine_imports.c plus app sources -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 -frounding-math -ffp-contract=off, -O2 for Release, -O0 otherwise
native/EngineVision/project.yml (Xcode) ../build/engine-reuse/whole-exe with chunk_*.c, engine_bundle.c, engine_imports.c OTHER_CFLAGS with the same four engine flags (line 105)

Notes:

  • The Makefile's $(GEN)/engine_bundle.c rule has no prerequisites, so make only generates when the bundle is missing. CHUNKS is a $(wildcard) evaluated when the Makefile is parsed, so on a clean tree the generator must run before the chunk objects are known (the setup wizard and BUILDING.md run it explicitly first).
  • Makefile line 12 uses sed to make each chunk object depend on the sub_*.c files it includes. Chunk objects also depend on engine_hooks.h, engine_cpu.h, engine_arm64_fenv_prototype.h and engine_flags.h (line 30), so editing the hook list recompiles all translated code.
  • The macOS Makefile does not pass -frounding-math -ffp-contract=off; the visionOS builds and the source checks that compare against generated code do. engine_cpu.h itself has #pragma STDC FENV_ACCESS ON and #pragma STDC FP_CONTRACT OFF. See x87 Floating Point.
  • build_engine_vision.py refuses to build unless there are exactly 32 chunk_*.c files and engine_bundle.c, engine_imports.c and engine_functions.h exist. It reads ENGINE_FN_COUNT from the bundle and records the function count, total generated C bytes and SHA-256 of the bundle and imports file in its build report (lines 37-52).

Regeneration in the setup wizard

setup_halo.py regenerates only when needed. It hashes the expected executable digest, the requirements file hash, and the path and bytes of every tools/**/*.py, every .py/.h file under third_party/xwa, and every decompilation/**/*.txt. If the stored fingerprint (generation.sha256 in the wizard's private state directory) differs, or the output is incomplete (fewer than 32 chunks or a missing engine_bundle.c, engine_imports.c, engine_functions.h or generation.json), both generators run with HALO_EXE pointing at the staged executable. Before that it checks that capstone, pefile, numpy, PIL and unicorn import, installing requirements-development.txt (pinned capstone==5.0.6, pefile==2024.8.26, unicorn==2.1.4, ...) if not.

Generated vs checked in

Checked in (public source) Generated locally, never committed
Translator: tools/generate_engine_reuse.py, tools/export_engine_imports.py, tools/engine_reuse/*.py native/build/engine-reuse/<label>/sub_*.c, chunk_*.c, engine_functions.h, engine_bundle.c
Vendored toolkit: third_party/xwa/** (MIT, see THIRD_PARTY.md) engine_imports.c
Runtime headers and sandboxed runtime: native/EngineReuse/** generation.json
Numeric entry lists: decompilation/c9acf0c46954/*.txt test receipts such as native/build/engine-reuse-fpu-validation.json
Host: native/EngineHost/** object files in host-obj/, halo-host, app bundles

ARCHITECTURE.md states the policy: the repository includes the translator and numeric entry lists; whole-executable C, original executable data and compiled engine objects are produced locally and excluded from distribution. .gitignore excludes /native/build/. The checked-in list files are pinned by SHA-256 in SOURCE_MANIFEST.json.

CLI reference

tools/generate_engine_reuse.py

Argument Default Effect
entries (positional, any number) 5590a0 55cfd0 55efd0 hex entry addresses, or @path to read whitespace-separated hex from a file relative to the repository root
--max-functions N 250 stop visiting after N functions; must be 1..20000
--label NAME player-movement output directory native/build/engine-reuse/NAME; letters, digits and - only
--trap-unsupported off instead of dropping a function at its first unliftable instruction, emit engine_fail(cpu, "unsupported original instruction"); at that instruction and keep the rest
--chunks N 1 1: one engine_bundle.c with static functions; N > 1: N chunk files plus engine_functions.h, non-static functions
--discover off after the call graph drains, add code addresses found as immediate operands and repeat to a fixed point

Environment: HALO_EXE (input path). There is no option to skip the hash check.

tools/export_engine_imports.py

Argument Default Effect
--output-dir DIR native/build/engine-reuse/whole-exe where to write engine_imports.c; relative paths are under the repository root

Upstream entry points (not used by the build)

The vendored toolkit has its own CLIs, documented on XWA Decoder and Lifter: translator.py (--output, --split, --all, --analyze-only, --functions-json, --stubs, --pe-json), generate.py (positional exe/output/split) and pe_analyze.py (--json). None of the project's build paths call them.

Failure modes

Symptom Cause Where
AssertionError at startup, no output executable digest mismatch, invalid label, or --max-functions out of range lines 77-79
AssertionError mentioning selector or table bytes a hand-proven switch table differs, wrong executable lines 96-113
undecodable records entry outside code, budget exceeded, fallthrough outside the function, overlapping decode, empty block decode.py DecodeError
unsupported records (no --trap-unsupported) an instruction with no exact lowering; whole function dropped lift.py / mixins
instructionTraps instruction replaced by engine_fail at run time lift.py:149-153
Run time unsupported original instruction a trapped instruction actually executed generated code
Run time dispatch to non-leader address dispatch into the middle of a function at a non-leader, including into the gap after an ungenerated function generated L_ENTRY switch
Run time unresolved original engine boundary dispatch below the first generated entry with no host handler engine_dispatch
Run time unresolved import boundary dispatch to >= 0xFE000000 that the host does not own engine_dispatch

Testing

Test What it covers How it runs
tools/engine_reuse/test_decode.py conditional fallthrough into a separately listed epilogue becomes a tail call (0x4D0580/0x4D05CA), and stays internal when 0x4D05CA is not listed needs the real halo.exe; manual
tools/engine_reuse/test_flags.py EFLAGS helpers vs Unicorn, conditions, rejection rules manual (python3 tools/engine_reuse/test_flags.py)
tools/engine_reuse/test_fpu.py lifted x87 snippets vs Unicorn and an x86_64 Rosetta probe manual, macOS with Rosetta
native/EngineHost/tests/test_engine_direct_calls.c ENGINE_DIRECT behavior tools/run_source_checks.py
native/EngineHost/tests/test_dispatch_interest.py every hooked address is listed in engine_hooks.h tools/run_source_checks.py
native/Tests/EngineRuntimeContract.c sandboxed runtime API contract tools/run_source_checks.py
tools/benchmark_x87_rotating_stack.py four real generated leaves under two x87 layouts run_source_checks.py --check-only --allow-missing (SKIP without generated code)

The CI workflow (source-checks.yml) runs run_source_checks.py --portable-only with and without sanitizers on macOS 15. It has no halo.exe, so every test that needs generated code prints SKIP there. Details per test are on Testing and Source Checks and on the pages for each component.

Related pages

Clone this wiki locally