-
Notifications
You must be signed in to change notification settings - Fork 5
x87 Floating Point
Halo PC 1.10 does its floating-point math on the x87 FPU: an eight-register stack of 80-bit values with its own control word (rounding and precision), status word (condition codes, TOP) and tag word. Apple Silicon has no x87 and no 80-bit format, so the translator lowers every x87 instruction to a helper call in engine_cpu.h that computes in IEEE binary64 (double) and keeps the architectural stack, tags and status bits. This page covers the data model, rounding and precision, the ARM64 floating-point environment path, the round-to-nearest fast path, the lifter mapping, the move from a shifting to a rotating register stack, and the tests that pin all of it down. It is part of the EngineReuse Runtime.
| File | Role |
|---|---|
| native/EngineReuse/engine_cpu.h | All x87 helpers: register file, push/pop, compare, arithmetic, rounding, integer and 80-bit conversion, transcendental operations, environment save/restore |
| native/EngineReuse/engine_arm64_fenv_prototype.h | arm64 inline assembly that performs one operation under the guest rounding mode while preserving host FPCR/FPSR |
| tools/engine_reuse/fpu.py |
FPUMixin: x87 instruction to helper call lowering |
| tools/engine_reuse/test_fpu.py | Lifted snippets vs Unicorn and x86_64 under Rosetta |
| native/EngineHost/tests/test_x87_rotating_stack.c | Random operation streams: rotating layout vs the frozen shifting layout |
| native/EngineHost/tests/x87_shifting_engine_cpu.h | The pre-rotation engine_cpu.h, verbatim, as a test reference |
| x87_stack_ops.h, x87_stack_ops.inc, x87_stack_shifting.c | One operation set compiled once per layout |
| x87_translated.h, x87_translated_impl.inc, x87_translated_rotating.c, x87_translated_shifting.c, x87_translated_test.c | Real generated leaves compiled against both layouts; equivalence and timing |
| tools/benchmark_x87_rotating_stack.py | Driver for the translated-leaf comparison |
| native/EngineHost/tests/test_x87_nearest_fast_path.c | Nearest-mode fast path is bit-identical to the FPCR path |
| native/EngineHost/tests/test_arm64_fenv_prototype.c | Inline assembly vs libc fenv differential and benchmark |
| Aspect | x87 hardware | Master Chef |
|---|---|---|
| Register format | 80-bit extended (64-bit significand) |
double (53-bit significand) |
| Precision control (CW bits 8-9) | 24/53/64-bit significand rounding | recorded, not enforced; everything is binary64 |
| Rounding control (CW bits 10-11) | nearest/down/up/zero | honored for arithmetic, float stores, integer rounding, FSQRT
|
| Exception masks (CW bits 0-5) | masked results or #MF | recorded, not enforced; masked IEEE results (inf/NaN) are produced "as hardware would" |
| Stack overflow/underflow | invalid-operation, masked result is the indefinite NaN |
engine_fail ("x87 stack overflow", "empty x87 stack register") |
| Condition codes | C0..C3 | C0, C2, C3 for compares/FXAM/FPREM; C1 only for FXAM sign and FPREM; otherwise unmodeled |
| Status exception flags | IE, DE, ZE, OE, UE, PE, SF, ES | only IE (bit 0) is set, by out-of-range integer stores; FNCLEX clears 0x80FF
|
Halo runs with control word 0x027F (all exceptions masked, 53-bit precision, round to nearest), which the tests label "Halo's" (test_x87_rotating_stack.c:50). With precision control at 53 bits, add, subtract, multiply, divide and square root on finite values in the normal range round exactly like binary64, which is why a double model is sufficient; the scope line of test_fpu.py's receipt is "PC53 finite arithmetic". Values outside binary64's exponent range, which the x87's wider exponent could hold in a register, are not modeled. The fpu.py docstring: "Arithmetic runs in double regardless of the guest precision-control field. Masked exceptions produce IEEE results (inf/NaN) as hardware would; only an instruction the lifter cannot express fails translation."
logical ST(i) -> physical fp_reg[(fp_top + i) & 7]
fp_valid bit i -> ST(i) is not empty (logical, not physical)
| Helper | Behavior |
|---|---|
engine_fp_physical(c, i) |
(fp_top + i) & 7 |
engine_fp_read(c, i) |
fails "empty x87 stack register" if i >= 8 or ST(i) is empty |
engine_fp_write(c, i, v) |
stores and marks ST(i) valid; i >= 8 fails "invalid x87 stack register"
|
engine_fp_push(c, v) |
fails "x87 stack overflow" if ST(7) is valid; fp_top = (fp_top - 1) & 7; writes the new ST(0); fp_valid = (fp_valid << 1) | 1
|
engine_fp_pop(c) |
reads ST(0) (so popping an empty stack fails); copies physical ST(7) into the vacated slot; fp_valid >>= 1; fp_top = (fp_top + 1) & 7
|
engine_fp_exchange(c, i) |
FXCH: both registers must be valid |
engine_fp_free(c, i) |
FFREE: clears logical bit i |
engine_fp_set_top(c, top) |
changes TOP without changing logical contents by rotating all eight physical slots; used only by engine_fp_init and engine_fp_load_environment
|
engine_fp_status(c) |
fp_status with bits 11-13 replaced by fp_top
|
engine_fp_tag_word(c) |
builds the architectural 16-bit tag word per physical register: 3 empty, 1 zero, 2 special (NaN, infinity, subnormal), 0 valid |
flowchart LR
subgraph BEFORE ["before push (fp_top = 6)"]
P6["fp_reg[6] = ST(0)"]
P7["fp_reg[7] = ST(1)"]
end
subgraph AFTER ["after push v (fp_top = 5)"]
Q5["fp_reg[5] = ST(0) = v"]
Q6["fp_reg[6] = ST(1)"]
Q7["fp_reg[7] = ST(2)"]
end
P6 --> Q6
P7 --> Q7
Push and pop touch one slot instead of moving all eight doubles. The odd details (pop copying the old ST(7), set_top rotating contents) exist so that the rotating layout is observably identical to the shifting layout it replaced, including the contents of empty registers, which a later FLDENV with an edited tag word can make valid again. Comments at lines 161-164 and 200-201 state this.
- The upstream XWA lifter first gave each generated function its own
double _st[8]. The upstream header records the consequence: "any function returning a value in st(0) silently returned nothing ... That is what made the CRT __ftol helper return 0 for every float->int conversion in the game" (recomp_types.h:41-44). It moved to a global array shifted on every push/pop. -
EngineCPUkept the x87 state in the CPU struct, initially asdouble fp[8]with ST(i) always infp[i]and every push/pop shifting eight doubles. That header survives verbatim as x87_shifting_engine_cpu.h ("Test reference only ... exactly as it was before the x87 stack became a rotating register file"). - The current rotating
fp_reg[8]is required by tests to match the shifting version bit for bit on every observable (see Tests).
Anything that exports x87 state uses logical order. HaloEngineState.fp_values[8] in engine_runtime.h is "ST(0)..ST(7) in stack order, whatever fp_top is", and FNSAVE's register area is "in stack order, ST(0) first". Callers must use the helpers rather than indexing fp_reg (struct comment).
flowchart TD
A["engine_fp_arithmetic(cpu, op, a, b)"] --> N{"RC == 0 (nearest)?<br/>engine_fp_nearest"}
N -- yes --> P["plain C: a+b, a-b, a*b, a/b"]
N -- no --> F{"HALO_ARM64_FENV_FAST<br/>and __aarch64__?"}
F -- yes --> ASM["halo_arm64_fp_arithmetic(RC, op, a, b)<br/>(FPCR rounding swap in inline asm)"]
F -- no --> FE["feholdexcept, fesetround(mode),<br/>volatile op, fesetenv"]
The same structure applies to engine_write_f32 (double to float store). engine_write_f64 is a plain store because the value already is binary64.
From engine_cpu.h:132-141:
Round-to-nearest is the guest's rounding mode almost all the time, and the host's own: the engine thread's FPCR is left at nearest by every path that changes it. In that mode an x87 operation is exactly the plain binary64 one. The FPCR/FPSR round trip below exists for the other modes, and every guest add, subtract, multiply, divide and float store paid for it anyway: two serialising FPSR writes each, about 23 ns an operation against about 1 ns, at 31k call sites that the game tick and every bearing pass run through. The host's floating-point exception flags are observed by neither the engine nor the host (the guest's own status word is fp_status), so leaving them sticky loses nothing.
So in nearest mode the host's cumulative FPSR exception bits may be set by guest arithmetic; nothing reads them.
engine_fp_round_integer never touches the host rounding mode: truncate (FISTTP) uses trunc; otherwise RC 0/1/2/3 use __builtin_roundeven/floor/ceil/trunc (on arm64, the frintn/frintm/frintp/frintz instructions). The comment notes this "costs a tenth of switching the host mode around nearbyint, which CRT floor (FRNDINT under round-down) and _ftol (FISTP under chop) did on every call."
FSQRT in a non-nearest mode uses fegetround/fesetround around sqrt (not the assembly helper). Transcendental functions (sin, cos, tan, atan2, log2, exp2, ldexp, fmod, remainder) always use the host libm in the host mode.
engine_arm64_fenv_prototype.h is compiled in when HALO_ARM64_FENV_PROTOTYPE or HALO_ARM64_FENV_FAST is non-zero on __aarch64__. engine_cpu.h includes it when HALO_ARM64_FENV_FAST is set (lines 9-14), which every game build does. (The header's first line still says "Isolated prototype. Not included by engine_cpu.h; default OFF."; that comment predates the include.)
Each call runs one instruction (fadd/fsub/fmul/fdiv on doubles, or fcvt double to single) inside this sequence:
mrs old, fpcr ; save host control
mrs status, fpsr ; save host status
temp = (old & ~0xC09F00) | round(mode)
(cmp temp, old; b.eq skip) ; HALO_ARM64_FENV_SKIP_UNCHANGED (default 1)
msr fpcr, temp
msr fpsr, status & ~0x9F ; clear cumulative exception flags
<operation>
(cmp temp, old; b.eq skip)
msr fpcr, old ; restore host control
msr fpsr, status ; restore host status
| x87 RC | Meaning | ARM FPCR.RMode (bits 22-23) |
|---|---|---|
| 0 | nearest even | 0 (RN) |
| 1 | toward minus infinity | 2 (RM) |
| 2 | toward plus infinity | 1 (RP) |
| 3 | toward zero | 3 (RZ) |
The mask ~0xC09F00 clears RMode and the trap-enable bits (8-12 and 15) and keeps everything else (for example FZ and DN), "matching feholdexcept + fesetround". The host's FPCR and FPSR are bit-identical afterwards. test_arm64_fenv_prototype.c checks this after every call.
engine_fp_compare(c, a, b, integer_flags):
| Result | Bits |
FCOM-family: status word |
FCOMI-family: EFLAGS |
|---|---|---|---|
| unordered (either NaN) | 0x45 |
C3, C2, C0 set | ZF, PF, CF set |
a < b |
0x01 |
C0 | CF |
a == b |
0x40 |
C3 | ZF |
a > b |
0x00 |
none | none |
For the status word, C0/C2/C3 (0x4500) are replaced and C1 is untouched. For EFLAGS, 0x8D5 (CF, PF, AF, ZF, SF, OF) is cleared first, so OF, SF and AF end up 0 as the architecture specifies. FCOM and FUCOM are treated identically (no invalid-operation distinction for quiet NaNs). Guest code then usually does fnstsw ax; test ah, 0x41; jcc or sahf; jcc, which the integer EFLAGS model handles exactly (see XWA Decoder and Lifter).
| Helper | Behavior |
|---|---|
engine_read_f32 / engine_read_f64
|
load and widen to double
|
engine_write_f32(c, a, v) |
narrow with guest rounding (fast path in nearest) |
engine_write_f64 |
plain store |
engine_read_f80 / engine_write_f80
|
via engine_f80_to_double / engine_double_to_f80 (lines 339-356): zero, infinity and NaN are mapped explicitly; finite values use ldexp(mantissa, exponent - 16383 - 63) and frexp. Stores write a quiet NaN 0xC000000000000000 for any NaN. The extra 11 significand bits are lost |
engine_fp_load_integer(c, a, 2/4/8) |
FILD; sign-extends and converts. For 64-bit integers "x87 holds every int64 exactly; this double-backed stack does not"; inexact magnitudes round as a 53-bit significand would. The integer indefinite 0x8000000000000000 is exact |
engine_fp_integer_operand(c, a, 2/4) |
the integer operand of FIADD/FIMUL/FISUB(R)/FIDIV(R)/FICOM |
engine_fp_store_integer(c, a, size, truncate) |
FIST/FISTP/FISTTP: round per RC (or truncate), then if not finite or outside [-2^(n-1), 2^(n-1)) write the integer indefinite (0x8000, 0x80000000, 0x8000000000000000) and set IE (status bit 0) |
engine_fp_transcendental and engine_fp_unary:
| Instruction | Result | Status |
|---|---|---|
FCHS, FABS
|
negate, fabs
|
|
FSQRT |
sqrt with guest rounding |
|
FRNDINT |
engine_fp_round_integer |
|
FSIN, FCOS
|
sin, cos of ST(0) |
C2 cleared (no out-of-range case) |
FSINCOS |
ST(0) = sin, then push cos | C2 cleared |
FPTAN |
ST(0) = tan, then push 1.0 | C2 cleared |
FPATAN |
ST(1) = atan2(ST(1), ST(0)), pop | |
FYL2X |
ST(1) = ST(1) * log2(ST(0)), pop | |
FYL2XP1 |
ST(1) = ST(1) * log2(1 + ST(0)), pop | |
F2XM1 |
exp2(x) - 1 |
|
FSCALE |
ldexp(ST(0), trunc(ST(1))), exponent clamped to +-20000 |
|
FPREM, FPREM1
|
fmod, remainder
|
C0/C1/C3 from the low three quotient bits, C2 cleared (reduction always complete) |
FXAM |
C3/C2/C0 class (empty 0x4100, NaN 0x0100, infinity 0x0500, zero 0x4000, subnormal 0x4400, normal 0x0400), C1 = sign |
|
FXTRACT |
ST(0) = exponent, then push significand |
Note on FPREM: Intel documents C0 = Q2, C3 = Q1, C1 = Q0. The helper sets C1 from quotient bit 0, C0 from bit 1 and C3 from bit 2 (lines 383-385), which swaps C0 and C3 relative to that description. The rotating-stack test only compares the helper against its own previous version, so it does not detect this; the code does not establish whether Halo reads those bits.
| Instruction | Helper / code | Notes |
|---|---|---|
FNINIT/FINIT
|
engine_fp_init |
TOP 0 (contents rotated, not cleared), all empty, CW 0x037F, SW 0 |
FLDCW |
engine_fp_control |
stores the word; "Exception masks and precision control are recorded, not enforced" |
FNSTCW |
write cpu->fp_control
|
|
FNSTSW (memory or AX) |
write engine_fp_status(cpu)
|
TOP merged in |
FNCLEX |
fp_status &= ~0x80FF |
clears exception flags, SF, ES and B |
FNSTENV |
engine_fp_store_environment |
28-byte protected-mode layout: CW, SW, tag word as dwords at +0/+4/+8, instruction and operand pointers written as zero |
FLDENV |
engine_fp_load_environment |
CW; SW & 0x47FF; TOP from SW bits 11-13 via set_top; valid bits from the tag word (physical to logical) |
FNSAVE |
engine_fp_save |
environment, then eight 10-byte registers in ST order (empty ones written as 0.0), then engine_fp_init
|
FRSTOR |
engine_fp_restore |
environment, then the eight registers |
FWAIT |
comment only | "Masked x87 exceptions are checked by each operation" |
The non-N forms (FSTENV, FSAVE, FSTSW, FSTCW, FCLEX, FINIT) lower identically.
FPUMixin.lift_fpu_instruction handles any mnemonic starting with f, plus wait/fwait:
| x86 | Generated C |
|---|---|
fldz fld1 fldpi fldl2e fldl2t fldlg2 fldln2 |
engine_fp_push(cpu, <constant>); (decimal constants in _CONSTANTS) |
fld m32/m64/m80 |
engine_fp_push(cpu, engine_read_f32/f64/f80(cpu, addr)); |
fld st(i) |
engine_fp_push(cpu, engine_fp_read(cpu, i)); |
fild m16/m32/m64 |
engine_fp_load_integer(cpu, addr, size); |
fst/fstp st(i) |
engine_fp_write(cpu, i, engine_fp_read(cpu, 0)); [+ engine_fp_pop(cpu);] |
fst/fstp m32/m64/m80 |
engine_write_f32/f64/f80(cpu, addr, engine_fp_read(cpu, 0)); [+ pop] |
fist/fistp/fisttp m16/m32/m64 |
engine_fp_store_integer(cpu, addr, size, <1 for fisttp>); [+ pop for fistp/fisttp] |
fadd fsub fsubr fmul fdiv fdivr with one memory or ST operand |
engine_fp_write(cpu, 0, engine_fp_arithmetic(cpu, ENGINE_FP_<OP>, ST(0), src)); (operands swapped for the r forms) |
same, two register operands st(i), st(j)
|
destination is the first operand |
faddp ... fdivrp [st(i)] |
destination ST(i) (default ST(1)), then pop |
fiadd fimul fisub fisubr fidiv fidivr m16/m32 |
as arithmetic with engine_fp_integer_operand
|
fcom fcomp fcompp fucom fucomp fucompp [src] |
engine_fp_compare(cpu, ST(0), src, 0); + 0/1/2 pops |
fcomi fcomip fucomi fucomip (and Capstone's fcompi/fucompi spellings) |
engine_fp_compare(cpu, ST(0), src, 1); + 0/1 pops |
ficom ficomp m16/m32 |
compare with integer operand, + pop for ficomp
|
ftst |
engine_fp_compare(cpu, ST(0), 0.0, 0); |
fnstsw/fstsw |
<dst> = engine_fp_status(cpu); (ax becomes SET_LO16(eax, ...)) |
fnstcw/fstcw, fldcw
|
<dst> = cpu->fp_control;, engine_fp_control(cpu, <src>);
|
fninit/finit, fnclex/fclex
|
engine_fp_init(cpu);, cpu->fp_status &= (uint16_t)~0x80ffu;
|
fxch [st(i)] |
engine_fp_exchange(cpu, i); (bare form uses ST(1)) |
ffree st(i) |
engine_fp_free(cpu, i); |
fchs fabs fsqrt frndint |
engine_fp_unary(cpu, ENGINE_FP_<NAME>); |
fsin fcos fsincos fptan fpatan fyl2x fyl2xp1 f2xm1 fscale fprem fprem1 fxam fxtract |
engine_fp_transcendental(cpu, ENGINE_FP_<NAME>); |
fcmovb/e/be/u/nb/ne/nbe/nu st(i) |
if (<EFLAGS condition b/e/be/p/ae/ne/a/np>) { engine_fp_write(cpu, 0, engine_fp_read(cpu, i)); } |
fnstenv fstenv fldenv fnsave fsave frstor m |
environment helpers above |
wait/fwait
|
comment |
Anything else beginning with f, or an unsupported operand form, raises ValueError (for example fbld/fbstp, fdecstp/fincstp, fnop, fxsave). With --trap-unsupported, that instruction becomes a runtime engine_fail.
engine_cpu.h sets #pragma STDC FENV_ACCESS ON and #pragma STDC FP_CONTRACT OFF, and the fallback paths use volatile results, so the compiler should neither move operations across rounding-mode changes nor fuse multiply-add pairs (which would change results). The builds also pass flags:
| Build | Flags for generated code |
|---|---|
visionOS (build_engine_vision.py, project.yml) |
-frounding-math -ffp-contract=off -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 |
macOS host (native/EngineHost/Makefile) |
-O2 -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 (no -frounding-math/-ffp-contract=off; relies on the pragmas) |
Source checks that run translated code (x87_rotating_stack, visible_surfaces, native_leaves_headset_flags, ...) |
-frounding-math -ffp-contract=off, often with -DHALO_ARM64_FENV_FAST=1, "built as the headset builds the engine" |
test_fpu.py assembles 25 snippets from raw opcode bytes, each loading operands from memory, running one x87 form and storing or reporting a result:
- popping arithmetic
faddp fsubp fsubrp fmulp fdivp fdivrp; - register forms
fsub/fsubr/fdiv/fdivr/fadd st(1)andfsubp st(2); -
fxch st(0),fxch st(1),fstp st(1); -
fcompp/fucomppfollowed byfnstsw ax(status),fcompi/fucompi(EFLAGS); -
fistpm16/m32/m64,fisttpm32,fstpm32,frndint.
It lifts them with FPUMixin over the upstream Lifter into one C program, compiles it at -O0, -O3 and -O1 -fsanitize=address,undefined (all with -frounding-math -ffp-contract=off), and runs each snippet for four control words (0x027F, 0x067F, 0x0A7F, 0x0E7F: 53-bit precision, each rounding mode) over 13 fixed inputs (including signed zeros, half-integers, values just beyond the int16/int32 limits, 2^53 and +-2^63) plus 32 seeded random triples, and NaN/infinity pairs for compares. The reference is the same bytes executed in Unicorn. Because "Unicorn's FISTTP overflow and FCOMI flag clearing differ from x86", those two are also checked against an x86_64 helper built with clang -arch x86_64 and run under Rosetta (arch -x86_64). Comparisons: stored bytes; for status cases EAX under mask 0xFFFF7D00 (upper half, C0, C2, C3 and TOP; C1 and B excluded); for EFLAGS cases ZF/PF/CF against Unicorn and all of 0x8D5 against Rosetta. It writes native/build/engine-reuse-fpu-validation.json with scope "PC53 finite arithmetic; actual C0/C2/C3 and EFLAGS comparisons; not full x87 exception/C1 fidelity or gameplay." Requires macOS with Rosetta, unicorn and capstone; not part of run_source_checks.py.
test_x87_rotating_stack.c links x87_stack_ops.inc twice: once against the current engine_cpu.h (prefix rotating) and once, through x87_stack_shifting.c, against x87_shifting_engine_cpu.h (prefix shifting). Each side has its own 4 GiB MAP_NORESERVE guest arena. It runs 12,000 sequences (1,500 under AddressSanitizer) of 300 operations from 29 kinds, one per helper sequence the lifter emits (loads from every source, stores to every destination, integer loads/stores, arithmetic in register/memory/integer forms, compares with their pops, FNSTSW/FNSTCW/FLDCW/FNINIT/FNCLEX, FXCH, FFREE, unary, all 13 transcendental operations, FCMOV, FNSTENV, arbitrary environment edits through FLDENV, FNSAVE/FRSTOR, raw memory pokes, reads of possibly empty registers). Values mix special constants, arbitrary bit patterns and scaled integers; control words are mostly 0x027F, sometimes other rounding modes, sometimes random. Push weights are steered by depth so the stack reaches every depth and occasionally overflows or underflows.
After every operation it requires identical bits of all eight logical slots (including empty ones), GPRs, EFLAGS, pc, control, status, status word with TOP, tag word, valid mask, TOP, the engine_fail reason, and every byte of the 1 KiB guest window. Coverage assertions at the end require every operation kind to have run and succeeded, all 13 transcendental and 4 unary and 4 arithmetic selectors, all 4 rounding modes, every depth 0-8, every TOP 0-7, and some failures. run_source_checks.py builds it twice, with HALO_ARM64_FENV_FAST=0 and =1, using -frounding-math -ffp-contract=off (lines 52-54).
x87_translated_test.c compiles four real generated functions, sub_0050D5B0, sub_004CC0D0, sub_00554260 and sub_00553380, twice without modifying them: x87_translated_impl.inc renames the symbols per layout with macros and #includes <sub_XXXXXXXX.c> from the generated directory, once under the shifting header and once under the rotating one. Each guest arena is 4 GiB PROT_NONE except eight 16 KiB-aligned regions (stack, a constant page at 0x670000, and per-call object/box/matrix/output/plane/byte records). For two input corpora (finite values, and one with NaN, infinity, subnormal and signed-zero floats mixed in) it calls each function for 16 records x 5 control words (0x027F, 0x037F, 0x067F, 0x0A7F, 0x0E7F) x 8 TOP values x entry stack depths 0/2/6/8, and for sub_004CC0D0 four pointer-aliasing patterns. It sets a different host rounding mode and a raised FE_DIVBYZERO before each call, and requires identical CPU state, x87 state (including empty slots), host fegetround() and fetestexcept() results, failure reasons and every mapped guest byte after each call, plus normal return when the entry stack is empty. With a non-zero iteration count it then times both layouts in alternating order and prints medians.
benchmark_x87_rotating_stack.py finds the generated directory (--generated, else HALO_GENERATED_ENGINE or HALO_ENGINE_GEN, else native/build/engine-reuse/whole-exe in this checkout or the main worktree), prints SHA-256 of the four sources and both headers, builds in a temporary directory with -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 -frounding-math -ffp-contract=off and runs it.
| Option | Default | Effect |
|---|---|---|
--generated DIR |
auto | generated sources directory |
--check-only |
off | equivalence only, no timing (iterations 0) |
--allow-missing |
off | print SKIP instead of an error when no generated tree exists |
--sanitize |
off |
-O1 -g -fsanitize=address,undefined; disables timing |
--iterations N |
20000 | calls per timing sample |
--samples N |
9 | alternating-order timing pairs; must be odd |
env CC
|
clang |
compiler |
run_source_checks.py calls it with --check-only --allow-missing (and --sanitize in the sanitizer run), so CI without a generated engine prints SKIP. "This is host evidence, not headset frame timing."
Requires arm64 and HALO_ARM64_FENV_FAST=1 (prints SKIP otherwise); part of run_source_checks.py. Checks (source):
- 2,000,000 random or special operand pairs: nearest-mode
engine_fp_arithmeticfor all four operations is bit-identical tohalo_arm64_fp_arithmetic(0, ...); so arefloatstores, integer rounding (vsnearbyint) andFSQRT(vssqrt); - for RC 1-3, 200,000 pairs each still match the assembly path exactly;
- integer rounding in each guest mode equals
nearbyintunder that mode, whatever the host mode is (400,000 values per mode); - the host rounding mode remains nearest throughout;
- timing of a 20,000,000-step dependent multiply chain on both paths.
A manual differential and benchmark for the assembly helper (not wired into run_source_checks.py). Built with HALO_ARM64_FENV_PROTOTYPE=1; adding -DHALO_ARM64_FENV_ACTUAL_HEADER routes the "fast" side through engine_cpu.h (engine_fp_arithmetic, engine_write_f32) instead of the raw helper and also checks that an invalid operation code fails without disturbing FPCR/FPSR. For 4 host rounding modes x 8 host variants (FZ, DN and attempted trap enables varied independently, with sticky status bits set) it runs 4 guest modes x 4 operations x all pairs of 18 special bit patterns, plus 10,000 random pairs per variant, against a feholdexcept/fesetround/fesetenv reference, requiring bit-identical results (double and float) and unchanged FPCR and FPSR after every call. It then prints libc-vs-assembly timings and the hardware's trap-enable readback.
Documents master-chef at commit 9f915af (v1.0.3). Unofficial project, not affiliated with Microsoft, Bungie, Gearbox or Apple. Original code is MIT licensed; game content is not included.
Overview
- Architecture Overview
- Repository Layout
- Glossary
- Environment Variables
- Contributing Guide
- Open Questions
Translation
- Static Translation Pipeline
- XWA Decoder and Lifter
- Function Address Lists
- EngineReuse Runtime
- x87 Floating Point
Host runtime
- EngineHost Overview
- Win32 Compatibility Layer
- Threading and Synchronization
- Guest Memory and Heap
- Engine Overrides and Hooks
- Runtime Settings
Graphics
- Direct3D9 Bridge
- Metal Renderer
- Shader Translation
- Textures and Texture Packs
- Geometry Fast Paths
- Radial Fog
Panorama and presentation
- Panorama System
- Panorama Budget and LOD
- Frame Pacing
- visionOS App
- Immersive Presenter
- Layer Alignment
Audio and input
Tooling and process