Skip to content

x87 Floating Point

M T edited this page Oct 4, 2026 · 1 revision

x87 Floating Point

Halo PC 1.10 does its floating-point math on the x87 FPU: an eight-register stack of 80-bit values with its own control word (rounding and precision), status word (condition codes, TOP) and tag word. Apple Silicon has no x87 and no 80-bit format, so the translator lowers every x87 instruction to a helper call in engine_cpu.h that computes in IEEE binary64 (double) and keeps the architectural stack, tags and status bits. This page covers the data model, rounding and precision, the ARM64 floating-point environment path, the round-to-nearest fast path, the lifter mapping, the move from a shifting to a rotating register stack, and the tests that pin all of it down. It is part of the EngineReuse Runtime.

Source files

File Role
native/EngineReuse/engine_cpu.h All x87 helpers: register file, push/pop, compare, arithmetic, rounding, integer and 80-bit conversion, transcendental operations, environment save/restore
native/EngineReuse/engine_arm64_fenv_prototype.h arm64 inline assembly that performs one operation under the guest rounding mode while preserving host FPCR/FPSR
tools/engine_reuse/fpu.py FPUMixin: x87 instruction to helper call lowering
tools/engine_reuse/test_fpu.py Lifted snippets vs Unicorn and x86_64 under Rosetta
native/EngineHost/tests/test_x87_rotating_stack.c Random operation streams: rotating layout vs the frozen shifting layout
native/EngineHost/tests/x87_shifting_engine_cpu.h The pre-rotation engine_cpu.h, verbatim, as a test reference
x87_stack_ops.h, x87_stack_ops.inc, x87_stack_shifting.c One operation set compiled once per layout
x87_translated.h, x87_translated_impl.inc, x87_translated_rotating.c, x87_translated_shifting.c, x87_translated_test.c Real generated leaves compiled against both layouts; equivalence and timing
tools/benchmark_x87_rotating_stack.py Driver for the translated-leaf comparison
native/EngineHost/tests/test_x87_nearest_fast_path.c Nearest-mode fast path is bit-identical to the FPCR path
native/EngineHost/tests/test_arm64_fenv_prototype.c Inline assembly vs libc fenv differential and benchmark

Model summary

Aspect x87 hardware Master Chef
Register format 80-bit extended (64-bit significand) double (53-bit significand)
Precision control (CW bits 8-9) 24/53/64-bit significand rounding recorded, not enforced; everything is binary64
Rounding control (CW bits 10-11) nearest/down/up/zero honored for arithmetic, float stores, integer rounding, FSQRT
Exception masks (CW bits 0-5) masked results or #MF recorded, not enforced; masked IEEE results (inf/NaN) are produced "as hardware would"
Stack overflow/underflow invalid-operation, masked result is the indefinite NaN engine_fail ("x87 stack overflow", "empty x87 stack register")
Condition codes C0..C3 C0, C2, C3 for compares/FXAM/FPREM; C1 only for FXAM sign and FPREM; otherwise unmodeled
Status exception flags IE, DE, ZE, OE, UE, PE, SF, ES only IE (bit 0) is set, by out-of-range integer stores; FNCLEX clears 0x80FF

Halo runs with control word 0x027F (all exceptions masked, 53-bit precision, round to nearest), which the tests label "Halo's" (test_x87_rotating_stack.c:50). With precision control at 53 bits, add, subtract, multiply, divide and square root on finite values in the normal range round exactly like binary64, which is why a double model is sufficient; the scope line of test_fpu.py's receipt is "PC53 finite arithmetic". Values outside binary64's exponent range, which the x87's wider exponent could hold in a register, are not modeled. The fpu.py docstring: "Arithmetic runs in double regardless of the guest precision-control field. Masked exceptions produce IEEE results (inf/NaN) as hardware would; only an instruction the lifter cannot express fails translation."

The register stack

Rotating register file

logical ST(i)  ->  physical fp_reg[(fp_top + i) & 7]
fp_valid bit i ->  ST(i) is not empty           (logical, not physical)

engine_cpu.h:158-208:

Helper Behavior
engine_fp_physical(c, i) (fp_top + i) & 7
engine_fp_read(c, i) fails "empty x87 stack register" if i >= 8 or ST(i) is empty
engine_fp_write(c, i, v) stores and marks ST(i) valid; i >= 8 fails "invalid x87 stack register"
engine_fp_push(c, v) fails "x87 stack overflow" if ST(7) is valid; fp_top = (fp_top - 1) & 7; writes the new ST(0); fp_valid = (fp_valid << 1) | 1
engine_fp_pop(c) reads ST(0) (so popping an empty stack fails); copies physical ST(7) into the vacated slot; fp_valid >>= 1; fp_top = (fp_top + 1) & 7
engine_fp_exchange(c, i) FXCH: both registers must be valid
engine_fp_free(c, i) FFREE: clears logical bit i
engine_fp_set_top(c, top) changes TOP without changing logical contents by rotating all eight physical slots; used only by engine_fp_init and engine_fp_load_environment
engine_fp_status(c) fp_status with bits 11-13 replaced by fp_top
engine_fp_tag_word(c) builds the architectural 16-bit tag word per physical register: 3 empty, 1 zero, 2 special (NaN, infinity, subnormal), 0 valid
flowchart LR
    subgraph BEFORE ["before push (fp_top = 6)"]
      P6["fp_reg[6] = ST(0)"]
      P7["fp_reg[7] = ST(1)"]
    end
    subgraph AFTER ["after push v (fp_top = 5)"]
      Q5["fp_reg[5] = ST(0) = v"]
      Q6["fp_reg[6] = ST(1)"]
      Q7["fp_reg[7] = ST(2)"]
    end
    P6 --> Q6
    P7 --> Q7
Loading

Push and pop touch one slot instead of moving all eight doubles. The odd details (pop copying the old ST(7), set_top rotating contents) exist so that the rotating layout is observably identical to the shifting layout it replaced, including the contents of empty registers, which a later FLDENV with an edited tag word can make valid again. Comments at lines 161-164 and 200-201 state this.

History: from per-function arrays to a rotating file

  1. The upstream XWA lifter first gave each generated function its own double _st[8]. The upstream header records the consequence: "any function returning a value in st(0) silently returned nothing ... That is what made the CRT __ftol helper return 0 for every float->int conversion in the game" (recomp_types.h:41-44). It moved to a global array shifted on every push/pop.
  2. EngineCPU kept the x87 state in the CPU struct, initially as double fp[8] with ST(i) always in fp[i] and every push/pop shifting eight doubles. That header survives verbatim as x87_shifting_engine_cpu.h ("Test reference only ... exactly as it was before the x87 stack became a rotating register file").
  3. The current rotating fp_reg[8] is required by tests to match the shifting version bit for bit on every observable (see Tests).

The logical ABI

Anything that exports x87 state uses logical order. HaloEngineState.fp_values[8] in engine_runtime.h is "ST(0)..ST(7) in stack order, whatever fp_top is", and FNSAVE's register area is "in stack order, ST(0) first". Callers must use the helpers rather than indexing fp_reg (struct comment).

Control word, rounding and the fast path

Decision flow for one arithmetic operation

engine_fp_arithmetic:

flowchart TD
    A["engine_fp_arithmetic(cpu, op, a, b)"] --> N{"RC == 0 (nearest)?<br/>engine_fp_nearest"}
    N -- yes --> P["plain C: a+b, a-b, a*b, a/b"]
    N -- no --> F{"HALO_ARM64_FENV_FAST<br/>and __aarch64__?"}
    F -- yes --> ASM["halo_arm64_fp_arithmetic(RC, op, a, b)<br/>(FPCR rounding swap in inline asm)"]
    F -- no --> FE["feholdexcept, fesetround(mode),<br/>volatile op, fesetenv"]
Loading

The same structure applies to engine_write_f32 (double to float store). engine_write_f64 is a plain store because the value already is binary64.

Why round-to-nearest is special-cased

From engine_cpu.h:132-141:

Round-to-nearest is the guest's rounding mode almost all the time, and the host's own: the engine thread's FPCR is left at nearest by every path that changes it. In that mode an x87 operation is exactly the plain binary64 one. The FPCR/FPSR round trip below exists for the other modes, and every guest add, subtract, multiply, divide and float store paid for it anyway: two serialising FPSR writes each, about 23 ns an operation against about 1 ns, at 31k call sites that the game tick and every bearing pass run through. The host's floating-point exception flags are observed by neither the engine nor the host (the guest's own status word is fp_status), so leaving them sticky loses nothing.

So in nearest mode the host's cumulative FPSR exception bits may be set by guest arithmetic; nothing reads them.

Rounding to integers

engine_fp_round_integer never touches the host rounding mode: truncate (FISTTP) uses trunc; otherwise RC 0/1/2/3 use __builtin_roundeven/floor/ceil/trunc (on arm64, the frintn/frintm/frintp/frintz instructions). The comment notes this "costs a tenth of switching the host mode around nearbyint, which CRT floor (FRNDINT under round-down) and _ftol (FISTP under chop) did on every call."

FSQRT in a non-nearest mode uses fegetround/fesetround around sqrt (not the assembly helper). Transcendental functions (sin, cos, tan, atan2, log2, exp2, ldexp, fmod, remainder) always use the host libm in the host mode.

ARM64 floating-point environment helper

engine_arm64_fenv_prototype.h is compiled in when HALO_ARM64_FENV_PROTOTYPE or HALO_ARM64_FENV_FAST is non-zero on __aarch64__. engine_cpu.h includes it when HALO_ARM64_FENV_FAST is set (lines 9-14), which every game build does. (The header's first line still says "Isolated prototype. Not included by engine_cpu.h; default OFF."; that comment predates the include.)

Each call runs one instruction (fadd/fsub/fmul/fdiv on doubles, or fcvt double to single) inside this sequence:

mrs old, fpcr                 ; save host control
mrs status, fpsr              ; save host status
temp = (old & ~0xC09F00) | round(mode)
  (cmp temp, old; b.eq skip)  ; HALO_ARM64_FENV_SKIP_UNCHANGED (default 1)
msr fpcr, temp
msr fpsr, status & ~0x9F      ; clear cumulative exception flags
<operation>
  (cmp temp, old; b.eq skip)
msr fpcr, old                 ; restore host control
msr fpsr, status              ; restore host status
x87 RC Meaning ARM FPCR.RMode (bits 22-23)
0 nearest even 0 (RN)
1 toward minus infinity 2 (RM)
2 toward plus infinity 1 (RP)
3 toward zero 3 (RZ)

The mask ~0xC09F00 clears RMode and the trap-enable bits (8-12 and 15) and keeps everything else (for example FZ and DN), "matching feholdexcept + fesetround". The host's FPCR and FPSR are bit-identical afterwards. test_arm64_fenv_prototype.c checks this after every call.

Comparisons and condition codes

engine_fp_compare(c, a, b, integer_flags):

Result Bits FCOM-family: status word FCOMI-family: EFLAGS
unordered (either NaN) 0x45 C3, C2, C0 set ZF, PF, CF set
a < b 0x01 C0 CF
a == b 0x40 C3 ZF
a > b 0x00 none none

For the status word, C0/C2/C3 (0x4500) are replaced and C1 is untouched. For EFLAGS, 0x8D5 (CF, PF, AF, ZF, SF, OF) is cleared first, so OF, SF and AF end up 0 as the architecture specifies. FCOM and FUCOM are treated identically (no invalid-operation distinction for quiet NaNs). Guest code then usually does fnstsw ax; test ah, 0x41; jcc or sahf; jcc, which the integer EFLAGS model handles exactly (see XWA Decoder and Lifter).

Integer and memory conversions

Helper Behavior
engine_read_f32 / engine_read_f64 load and widen to double
engine_write_f32(c, a, v) narrow with guest rounding (fast path in nearest)
engine_write_f64 plain store
engine_read_f80 / engine_write_f80 via engine_f80_to_double / engine_double_to_f80 (lines 339-356): zero, infinity and NaN are mapped explicitly; finite values use ldexp(mantissa, exponent - 16383 - 63) and frexp. Stores write a quiet NaN 0xC000000000000000 for any NaN. The extra 11 significand bits are lost
engine_fp_load_integer(c, a, 2/4/8) FILD; sign-extends and converts. For 64-bit integers "x87 holds every int64 exactly; this double-backed stack does not"; inexact magnitudes round as a 53-bit significand would. The integer indefinite 0x8000000000000000 is exact
engine_fp_integer_operand(c, a, 2/4) the integer operand of FIADD/FIMUL/FISUB(R)/FIDIV(R)/FICOM
engine_fp_store_integer(c, a, size, truncate) FIST/FISTP/FISTTP: round per RC (or truncate), then if not finite or outside [-2^(n-1), 2^(n-1)) write the integer indefinite (0x8000, 0x80000000, 0x8000000000000000) and set IE (status bit 0)

Transcendental and miscellaneous operations

engine_fp_transcendental and engine_fp_unary:

Instruction Result Status
FCHS, FABS negate, fabs
FSQRT sqrt with guest rounding
FRNDINT engine_fp_round_integer
FSIN, FCOS sin, cos of ST(0) C2 cleared (no out-of-range case)
FSINCOS ST(0) = sin, then push cos C2 cleared
FPTAN ST(0) = tan, then push 1.0 C2 cleared
FPATAN ST(1) = atan2(ST(1), ST(0)), pop
FYL2X ST(1) = ST(1) * log2(ST(0)), pop
FYL2XP1 ST(1) = ST(1) * log2(1 + ST(0)), pop
F2XM1 exp2(x) - 1
FSCALE ldexp(ST(0), trunc(ST(1))), exponent clamped to +-20000
FPREM, FPREM1 fmod, remainder C0/C1/C3 from the low three quotient bits, C2 cleared (reduction always complete)
FXAM C3/C2/C0 class (empty 0x4100, NaN 0x0100, infinity 0x0500, zero 0x4000, subnormal 0x4400, normal 0x0400), C1 = sign
FXTRACT ST(0) = exponent, then push significand

Note on FPREM: Intel documents C0 = Q2, C3 = Q1, C1 = Q0. The helper sets C1 from quotient bit 0, C0 from bit 1 and C3 from bit 2 (lines 383-385), which swaps C0 and C3 relative to that description. The rotating-stack test only compares the helper against its own previous version, so it does not detect this; the code does not establish whether Halo reads those bits.

Control, status and environment

Instruction Helper / code Notes
FNINIT/FINIT engine_fp_init TOP 0 (contents rotated, not cleared), all empty, CW 0x037F, SW 0
FLDCW engine_fp_control stores the word; "Exception masks and precision control are recorded, not enforced"
FNSTCW write cpu->fp_control
FNSTSW (memory or AX) write engine_fp_status(cpu) TOP merged in
FNCLEX fp_status &= ~0x80FF clears exception flags, SF, ES and B
FNSTENV engine_fp_store_environment 28-byte protected-mode layout: CW, SW, tag word as dwords at +0/+4/+8, instruction and operand pointers written as zero
FLDENV engine_fp_load_environment CW; SW & 0x47FF; TOP from SW bits 11-13 via set_top; valid bits from the tag word (physical to logical)
FNSAVE engine_fp_save environment, then eight 10-byte registers in ST order (empty ones written as 0.0), then engine_fp_init
FRSTOR engine_fp_restore environment, then the eight registers
FWAIT comment only "Masked x87 exceptions are checked by each operation"

The non-N forms (FSTENV, FSAVE, FSTSW, FSTCW, FCLEX, FINIT) lower identically.

Lifter mapping

FPUMixin.lift_fpu_instruction handles any mnemonic starting with f, plus wait/fwait:

x86 Generated C
fldz fld1 fldpi fldl2e fldl2t fldlg2 fldln2 engine_fp_push(cpu, <constant>); (decimal constants in _CONSTANTS)
fld m32/m64/m80 engine_fp_push(cpu, engine_read_f32/f64/f80(cpu, addr));
fld st(i) engine_fp_push(cpu, engine_fp_read(cpu, i));
fild m16/m32/m64 engine_fp_load_integer(cpu, addr, size);
fst/fstp st(i) engine_fp_write(cpu, i, engine_fp_read(cpu, 0)); [+ engine_fp_pop(cpu);]
fst/fstp m32/m64/m80 engine_write_f32/f64/f80(cpu, addr, engine_fp_read(cpu, 0)); [+ pop]
fist/fistp/fisttp m16/m32/m64 engine_fp_store_integer(cpu, addr, size, <1 for fisttp>); [+ pop for fistp/fisttp]
fadd fsub fsubr fmul fdiv fdivr with one memory or ST operand engine_fp_write(cpu, 0, engine_fp_arithmetic(cpu, ENGINE_FP_<OP>, ST(0), src)); (operands swapped for the r forms)
same, two register operands st(i), st(j) destination is the first operand
faddp ... fdivrp [st(i)] destination ST(i) (default ST(1)), then pop
fiadd fimul fisub fisubr fidiv fidivr m16/m32 as arithmetic with engine_fp_integer_operand
fcom fcomp fcompp fucom fucomp fucompp [src] engine_fp_compare(cpu, ST(0), src, 0); + 0/1/2 pops
fcomi fcomip fucomi fucomip (and Capstone's fcompi/fucompi spellings) engine_fp_compare(cpu, ST(0), src, 1); + 0/1 pops
ficom ficomp m16/m32 compare with integer operand, + pop for ficomp
ftst engine_fp_compare(cpu, ST(0), 0.0, 0);
fnstsw/fstsw <dst> = engine_fp_status(cpu); (ax becomes SET_LO16(eax, ...))
fnstcw/fstcw, fldcw <dst> = cpu->fp_control;, engine_fp_control(cpu, <src>);
fninit/finit, fnclex/fclex engine_fp_init(cpu);, cpu->fp_status &= (uint16_t)~0x80ffu;
fxch [st(i)] engine_fp_exchange(cpu, i); (bare form uses ST(1))
ffree st(i) engine_fp_free(cpu, i);
fchs fabs fsqrt frndint engine_fp_unary(cpu, ENGINE_FP_<NAME>);
fsin fcos fsincos fptan fpatan fyl2x fyl2xp1 f2xm1 fscale fprem fprem1 fxam fxtract engine_fp_transcendental(cpu, ENGINE_FP_<NAME>);
fcmovb/e/be/u/nb/ne/nbe/nu st(i) if (<EFLAGS condition b/e/be/p/ae/ne/a/np>) { engine_fp_write(cpu, 0, engine_fp_read(cpu, i)); }
fnstenv fstenv fldenv fnsave fsave frstor m environment helpers above
wait/fwait comment

Anything else beginning with f, or an unsupported operand form, raises ValueError (for example fbld/fbstp, fdecstp/fincstp, fnop, fxsave). With --trap-unsupported, that instruction becomes a runtime engine_fail.

Build flags

engine_cpu.h sets #pragma STDC FENV_ACCESS ON and #pragma STDC FP_CONTRACT OFF, and the fallback paths use volatile results, so the compiler should neither move operations across rounding-mode changes nor fuse multiply-add pairs (which would change results). The builds also pass flags:

Build Flags for generated code
visionOS (build_engine_vision.py, project.yml) -frounding-math -ffp-contract=off -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1
macOS host (native/EngineHost/Makefile) -O2 -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 (no -frounding-math/-ffp-contract=off; relies on the pragmas)
Source checks that run translated code (x87_rotating_stack, visible_surfaces, native_leaves_headset_flags, ...) -frounding-math -ffp-contract=off, often with -DHALO_ARM64_FENV_FAST=1, "built as the headset builds the engine"

Tests

test_fpu.py: Unicorn and Rosetta differential

test_fpu.py assembles 25 snippets from raw opcode bytes, each loading operands from memory, running one x87 form and storing or reporting a result:

  • popping arithmetic faddp fsubp fsubrp fmulp fdivp fdivrp;
  • register forms fsub/fsubr/fdiv/fdivr/fadd st(1) and fsubp st(2);
  • fxch st(0), fxch st(1), fstp st(1);
  • fcompp/fucompp followed by fnstsw ax (status), fcompi/fucompi (EFLAGS);
  • fistp m16/m32/m64, fisttp m32, fstp m32, frndint.

It lifts them with FPUMixin over the upstream Lifter into one C program, compiles it at -O0, -O3 and -O1 -fsanitize=address,undefined (all with -frounding-math -ffp-contract=off), and runs each snippet for four control words (0x027F, 0x067F, 0x0A7F, 0x0E7F: 53-bit precision, each rounding mode) over 13 fixed inputs (including signed zeros, half-integers, values just beyond the int16/int32 limits, 2^53 and +-2^63) plus 32 seeded random triples, and NaN/infinity pairs for compares. The reference is the same bytes executed in Unicorn. Because "Unicorn's FISTTP overflow and FCOMI flag clearing differ from x86", those two are also checked against an x86_64 helper built with clang -arch x86_64 and run under Rosetta (arch -x86_64). Comparisons: stored bytes; for status cases EAX under mask 0xFFFF7D00 (upper half, C0, C2, C3 and TOP; C1 and B excluded); for EFLAGS cases ZF/PF/CF against Unicorn and all of 0x8D5 against Rosetta. It writes native/build/engine-reuse-fpu-validation.json with scope "PC53 finite arithmetic; actual C0/C2/C3 and EFLAGS comparisons; not full x87 exception/C1 fidelity or gameplay." Requires macOS with Rosetta, unicorn and capstone; not part of run_source_checks.py.

test_x87_rotating_stack.c: random streams vs the shifting layout

test_x87_rotating_stack.c links x87_stack_ops.inc twice: once against the current engine_cpu.h (prefix rotating) and once, through x87_stack_shifting.c, against x87_shifting_engine_cpu.h (prefix shifting). Each side has its own 4 GiB MAP_NORESERVE guest arena. It runs 12,000 sequences (1,500 under AddressSanitizer) of 300 operations from 29 kinds, one per helper sequence the lifter emits (loads from every source, stores to every destination, integer loads/stores, arithmetic in register/memory/integer forms, compares with their pops, FNSTSW/FNSTCW/FLDCW/FNINIT/FNCLEX, FXCH, FFREE, unary, all 13 transcendental operations, FCMOV, FNSTENV, arbitrary environment edits through FLDENV, FNSAVE/FRSTOR, raw memory pokes, reads of possibly empty registers). Values mix special constants, arbitrary bit patterns and scaled integers; control words are mostly 0x027F, sometimes other rounding modes, sometimes random. Push weights are steered by depth so the stack reaches every depth and occasionally overflows or underflows.

After every operation it requires identical bits of all eight logical slots (including empty ones), GPRs, EFLAGS, pc, control, status, status word with TOP, tag word, valid mask, TOP, the engine_fail reason, and every byte of the 1 KiB guest window. Coverage assertions at the end require every operation kind to have run and succeeded, all 13 transcendental and 4 unary and 4 arithmetic selectors, all 4 rounding modes, every depth 0-8, every TOP 0-7, and some failures. run_source_checks.py builds it twice, with HALO_ARM64_FENV_FAST=0 and =1, using -frounding-math -ffp-contract=off (lines 52-54).

Translated leaves: benchmark_x87_rotating_stack.py

x87_translated_test.c compiles four real generated functions, sub_0050D5B0, sub_004CC0D0, sub_00554260 and sub_00553380, twice without modifying them: x87_translated_impl.inc renames the symbols per layout with macros and #includes <sub_XXXXXXXX.c> from the generated directory, once under the shifting header and once under the rotating one. Each guest arena is 4 GiB PROT_NONE except eight 16 KiB-aligned regions (stack, a constant page at 0x670000, and per-call object/box/matrix/output/plane/byte records). For two input corpora (finite values, and one with NaN, infinity, subnormal and signed-zero floats mixed in) it calls each function for 16 records x 5 control words (0x027F, 0x037F, 0x067F, 0x0A7F, 0x0E7F) x 8 TOP values x entry stack depths 0/2/6/8, and for sub_004CC0D0 four pointer-aliasing patterns. It sets a different host rounding mode and a raised FE_DIVBYZERO before each call, and requires identical CPU state, x87 state (including empty slots), host fegetround() and fetestexcept() results, failure reasons and every mapped guest byte after each call, plus normal return when the entry stack is empty. With a non-zero iteration count it then times both layouts in alternating order and prints medians.

benchmark_x87_rotating_stack.py finds the generated directory (--generated, else HALO_GENERATED_ENGINE or HALO_ENGINE_GEN, else native/build/engine-reuse/whole-exe in this checkout or the main worktree), prints SHA-256 of the four sources and both headers, builds in a temporary directory with -DENGINE_FLAT_MEMORY=1 -DHALO_ARM64_FENV_FAST=1 -frounding-math -ffp-contract=off and runs it.

Option Default Effect
--generated DIR auto generated sources directory
--check-only off equivalence only, no timing (iterations 0)
--allow-missing off print SKIP instead of an error when no generated tree exists
--sanitize off -O1 -g -fsanitize=address,undefined; disables timing
--iterations N 20000 calls per timing sample
--samples N 9 alternating-order timing pairs; must be odd
env CC clang compiler

run_source_checks.py calls it with --check-only --allow-missing (and --sanitize in the sanitizer run), so CI without a generated engine prints SKIP. "This is host evidence, not headset frame timing."

test_x87_nearest_fast_path.c

Requires arm64 and HALO_ARM64_FENV_FAST=1 (prints SKIP otherwise); part of run_source_checks.py. Checks (source):

  • 2,000,000 random or special operand pairs: nearest-mode engine_fp_arithmetic for all four operations is bit-identical to halo_arm64_fp_arithmetic(0, ...); so are float stores, integer rounding (vs nearbyint) and FSQRT (vs sqrt);
  • for RC 1-3, 200,000 pairs each still match the assembly path exactly;
  • integer rounding in each guest mode equals nearbyint under that mode, whatever the host mode is (400,000 values per mode);
  • the host rounding mode remains nearest throughout;
  • timing of a 20,000,000-step dependent multiply chain on both paths.

test_arm64_fenv_prototype.c

A manual differential and benchmark for the assembly helper (not wired into run_source_checks.py). Built with HALO_ARM64_FENV_PROTOTYPE=1; adding -DHALO_ARM64_FENV_ACTUAL_HEADER routes the "fast" side through engine_cpu.h (engine_fp_arithmetic, engine_write_f32) instead of the raw helper and also checks that an invalid operation code fails without disturbing FPCR/FPSR. For 4 host rounding modes x 8 host variants (FZ, DN and attempted trap enables varied independently, with sticky status bits set) it runs 4 guest modes x 4 operations x all pairs of 18 special bit patterns, plus 10,000 random pairs per variant, against a feholdexcept/fesetround/fesetenv reference, requiring bit-identical results (double and float) and unchanged FPCR and FPSR after every call. It then prints libc-vs-assembly timings and the hardware's trap-enable readback.

Related pages

Clone this wiki locally