Skip to content

Geometry Fast Paths

M T edited this page Oct 4, 2026 · 1 revision

Geometry Fast Paths

This page covers the code that makes Halo's geometry work cheaper on the host without changing what the engine computes: index normalisation, resident GPU copies of static vertex buffers, CPU execution of IDirect3DDevice9::ProcessVertices (with an exact-token skinning fast path and a differential checker), and native C replacements for a handful of hot translated engine routines that walk the BSP (frustum/visible-surface marking, light and shadow surface gathers, bounds tests, a matrix product). The panorama renders the world from up to nine bearings per engine frame (see Panorama System), so every per-draw and per-view cost is multiplied; these paths exist because of that multiplier. They sit at two layers: the Direct3D 9 Bridge and Metal Renderer for draws, and the dispatch override hook of the EngineReuse Runtime for the native engine routines (see Engine Overrides and Hooks).

Source files

File Role
d3d9_render.inc draw_indices, vb_resident, traffic report.
d3d9.c Residency budget and per-object ownership; Lock-flag retirement.
d3d9_process_vertices.inc ProcessVertices front end, skin/differential switches.
vertex_shader_cpu.h Bounded vs_1_1 compiler and interpreter.
process_vertices.h Input decoding, output layout from the destination FVF, atomic commit.
process_vertices_skin.h Exact-token match of Halo's skinning shader and its specialised executor.
process_vertices_differential.h Runs the generic and skin executors side by side and compares bytes.
native_leaves.h Native versions of six leaf routines (matrix product, bounds tests, BSP node bounds, surface list, BSP material walk).
visible_surfaces.h Native 00553920, the per-view visible-triangle marker.
native_gather.h, native_gather_tree.h Native light/shadow BSP surface gathers 005540C0, 00553C40, 00553F10.
overrides.c Wires the native routines into engine_dispatch_override and implements the visibility verify mode.

What is reused, and what is not

The word "reuse" is easy to misread here. The table is the complete list.

Mechanism Reuses across draws/views? What exactly Bound
Resident vertex buffers Yes An immutable GPU copy of a static vertex buffer's bytes, valid for one content_generation, verified against guest bytes on every draw that uses it. 128 MiB of live copies; at most 4 copies per buffer before it is permanently refused; dynamic/streamed buffers never.
Metal texture cache Yes Decoded textures keyed by content (see Textures and Texture Packs). 1024 entries / 512 MiB, LRU.
Pipelines, samplers, depth states Yes Immutable Metal state objects (see Metal Renderer). Per-cache limits.
Encoder state, folded clears Within one encoder / until the next consumer Skips re-sending identical encoder state; turns clears into load actions. Forgotten on every new encoder.
Native visible surfaces, gathers and leaves No Nothing. Each call reads the BSP, the visible-cluster records, the frustum and the current view's triangle bitset afresh and writes exactly what the original writes. The gather headers state it plainly: "no result survives a call", "No geometry or visibility is cached across calls", "Each entry rereads geometry and this bearing's visibility bits; no cross-view cache." The original engine's own caps (0x4000 marked triangles per view, caller-supplied output capacities, 0x7FFF surface-list entries).
ProcessVertices skin fast path No A faster but bit-identical executor for one exact shader; output is recomputed every call. Engine call pattern.

So the BSP/PVS speedups make the same per-view work cheaper; they never carry geometry or visibility from one view, bearing or frame to another.

Index normalisation

draw_indices runs on every indexed draw and every fan. It converts 16- or 32-bit guest indices, adds the base (-MinIndex for DrawIndexedPrimitive, so indices become relative to the first vertex actually passed), expands fans into lists (0, i+1, i+2), and rejects any result below 0, at or above the vertex count, or above 65535. The common case (16-bit indices, list or strip) first finds the minimum and maximum over the whole range in a loop the compiler vectorises, checks the rebased range once, then copies with the base added; it produces exactly what the general loop does and only differs in leaving the scratch untouched on a rejected draw (both callers treat NULL as failure). Scratch buffers grow to the largest draw and are reused, because per-draw malloc/free and page faults were pure overhead in the hottest path in the port.

Resident static vertex buffers

Halo fills every model part's and BSP section's vertex buffer once at map load (sub_00524980: managed pool, no D3DUSAGE_DYNAMIC, Lock/copy/Unlock) and then draws from it in every frame and every bearing. Without residency, each of those draws copied the part's whole vertex range into the frame arena, the largest single cost of the host draw path. With HALO_DRAW_FASTPATH=1, vb_resident takes one copy per content generation and binds it instead.

flowchart TD
    A["Programmable draw: stream s range (start, length)"] --> B{"Eligible? K_VB, not refused, not locked, not DYNAMIC, has data, fast paths on, start 4-aligned, range in buffer"}
    B -->|"no"| X["Arena copy (exact original path)"]
    B -->|"yes"| C{"Copy exists for current content_generation?"}
    C -->|"no"| D{"Copies taken >= 4?"}
    D -->|"yes"| R["Retire: refuse residency forever"] --> X
    D -->|"no"| E["Drop stale copy, budget check (128 MiB live)"]
    E -->|"over budget"| F["Count budget fallback"] --> X
    E -->|"ok"| G["mr_buffer_create(whole buffer), copies++"] --> H["Bind copy at offset"]
    C -->|"yes"| I{"Copy bytes == guest bytes for this range?"}
    I -->|"no"| J["Log, retire"] --> X
    I -->|"yes"| H
Loading

Why the bound bytes are always the bytes the arena path would have copied (from the source comment and code):

  • content_generation moves on Lock, Unlock and ProcessVertices (the host paths that hand out or write buffer memory), and a locked buffer is never copied or bound.
  • Buffers created D3DUSAGE_DYNAMIC (0x200) and buffers ever locked with NOOVERWRITE/DISCARD (vb_resident_retire at d3d9.c:895-896) keep the per-draw copy; so does a buffer rewritten more than four times (VB_RESIDENT_MAX_COPIES).
  • Every reused range is compared with the guest bytes before binding (vb_resident_equal: OR-reduce of XORs, no early exit, vectorisable), which catches writes through a retained Lock pointer or any other path that bypassed the host API. A mismatch retires the buffer before the draw uses the copy. The optimisation therefore replaces arena writes with reads; it does not assume all guest writes were observed.
  • The copy belongs to the D3D object and is released with it (vb_resident_drop in obj_destroy), so a reused object slot or recycled guest pages never meet an old copy. Queued Metal commands retain the buffer until the GPU finishes, because the command buffer uses retained references.
  • Only DrawPrimitive/DrawIndexedPrimitive pass stream zero's buffer; the UP draws take user memory and never use residency. Secondary declaration streams use the same check.
  • The renderer additionally requires a 4-byte-aligned offset inside the buffer (resident_range) and falls back to the arena copy otherwise.

Residency applies only to the programmable backend (mr_draw_program); the fixed-function and RHW backends transform or repack vertices on the CPU anyway.

Budget: VB_RESIDENT_BUDGET_BYTES = 128 MiB of live copies (d3d9.c:139-141), independent of the texture budget "so a large mission cannot turn the copy optimization into memory pressure". When a new copy would exceed it the draw uses the arena and budgetFallbacks is counted; a later draw retries.

Traffic report

Every 600 presented frames draw_traffic_report logs:

[draw-traffic] frame=N residentDraws=.. residentMB=.. arenaVertexMB=.. copies=.. mismatches=.. foldedClears=.. fast=0|1 cacheMB=.. peakCacheMB=.. budgetFallbacks=..

residentMB/arenaVertexMB are the renderer's cumulative bytes bound from copies and copied to the arena; cacheMB/peakCacheMB the live and peak resident bytes.

ProcessVertices on the CPU

Halo uses ProcessVertices with its skinning vertex shader to pre-transform skinned geometry into a vertex buffer that is then drawn. The bridge runs the original vs_1_1 token stream on the CPU.

process_vertices requires a bound vertex shader and declaration, a vertex-buffer destination and a null output declaration, otherwise D3DERR_INVALIDCALL. It increments the destination's content_generation up front (the API writes destination bytes without Lock/Unlock, which also invalidates any resident copy), then:

  1. pv_compile accepts only version 0xFFFE0101, at most 256 instructions, the opcodes mov add sub mad mul rcp rsq dp3 dp4 min max slt sge exp log lrp frc, dcl and def, temporaries r0-r11, inputs v0-v15, constants c0-c255 (relative addressing only on constants), a0, and outputs oPos, oFog, oPts, oD0-1, oT0-7. Destination shifts and source modifiers other than negate are rejected; _sat is supported. Anything else fails compilation with D3DERR_NOTAVAILABLE before any write. The current VS constant file is copied into the program, then def constants are applied.
  2. pv_layout derives the output layout from the destination buffer's FVF. Only untransformed D3DFVF_XYZ destinations are accepted (these are programmable stream-output buffers, not XYZRHW screen vertices), with optional normal, point size, diffuse, specular and up to 8 texture coordinate sets of 1-4 components.
  3. pv_inputs matches every used input register to a declaration element by usage (or by declaration order for old shaders without dcl) and checks, without overflow, that every stream covers start + count vertices.
  4. pv_process_execute copies the destination range into scratch, runs the executor per vertex, writes only the components the shader wrote (colours as clamped, rounded D3DCOLOR bytes), leaves normals and other unwritten fields untouched (D3DPV_DONOTCOPYDATA, the only flag value accepted), and commits the scratch to the destination only if every vertex succeeded. Invalid shaders, declarations, ranges or dynamic addresses never leave a partially processed buffer.

The interpreter (pv_execute) fails a vertex whose relative constant address is non-finite or outside the register file, and floors mov a0 (vs_1_1 semantics). Input types follow pv_decode (float1-4, D3DCOLOR, UBYTE4/N, SHORT2/4 and normalised forms, USHORT2N/4N, FLOAT16).

Skinning fast path (HALO_PV_SKIN_FAST=1)

pv_skin_shader_matches compares the complete token stream against one recorded Halo vs_1_1 skinning shader; any difference (declaration, masks, swizzles, constants, comments, version) uses the generic interpreter. On a match, pv_skin_execute evaluates the same program directly: bone indices from the blend-indices input are scaled by c9.w and offset by c5.w, floored, and used to address two 3-row (4x3) bone matrices starting at c29, blended by the two blend weights; oPos.xyz is the dot product of the position with the blended rows, oPos.w and oT0 come from the texture-coordinate input. It is written with FENV_ACCESS ON and FP_CONTRACT OFF, evaluates all four lanes before masking as the original mul/add do, and preserves the per-read failure order (including zero-weight reads), so it is byte-identical to the generic interpreter in every rounding mode. The visionOS app enables it by default ("exact-token skinning validated against original campaign calls"). It is never used while the differential mode is on.

Differential mode (HALO_PV_DIFFERENTIAL=1)

pv_process_differential, for calls whose shader matches the skin signature, from HALO_PV_DIFFERENTIAL_FROM_FRAME (strict decimal; an invalid value disables comparison) and for at most 128 checks:

  • runs the generic executor into the real destination (its output and success are always authoritative);
  • skips the comparison if the layout or range is invalid, the output exceeds 4 MiB (PV_DIFFERENTIAL_BYTE_CAP), the destination overlaps a source stream (the generic commit could change the candidate's input), or allocation/feholdexcept fails;
  • otherwise runs the skin executor into a copy under a held FP environment (so it cannot leak exception flags, rounding or traps into the engine), compares bytes, records the first differing byte and words, and restores the environment.

Results are logged as [pv-differential] ... authoritative=generic (first 4 checks, every 32nd, any mismatch, any restore failure, up to 32 lines). After the cap, calls stay on the generic path.

Native engine leaves (HALO_NATIVE_LEAVES)

native_leaves.h replaces six translated functions that only compute and are called thousands of times per bearing pass. The design contract (header comment): the native version leaves exactly what the original leaves - the same output bytes, the same dead stack frame (saved registers, locals, reused argument slots), the same eax..edi, the same flags from the last flag-setting instruction, and the same x87 status word and stale register contents that FLDENV could expose. Arithmetic is the translation's own: every x87 value is a binary64 holding a float, each operation is the same binary64 operation in the same operand order, stores round to float, and contraction is off (-ffast-math is a compile error). Because a native version keeps inputs in registers and writes outputs at the end, it first checks that no output range overlaps an input range or its own frame; if one does, or the guest is not rounding to nearest, or the x87 stack has too few free registers, or the direction flag is set, it returns 0 having written nothing and the translated function runs.

Address Function Called for
004CC0D0 leaf_matrix4x3_multiply: out = a * b for Halo's 52-byte real_matrix4x3 (scale, forward, left, up, position), with the original's copy-to-frame when out aliases an operand Every model node in every bearing and both eyes, and every object node each tick
00553380 leaf_bsp_node_bounds: decode a child node's bounds from its parent's and six bytes (byte/254 of the axis, 0xFF = maximum) Every BSP node the light/shadow gathers visit
00554260 leaf_bounds_planes: box against N planes: 0 all corners behind one plane, 2 none behind any, else 1; keeps the original's three float-rounded partial sums Every node and cluster the gathers visit
005541B0 leaf_bounds_overlap: box against box: 0 apart, 2 contained, 1 overlapping Gathers
00552C20 leaf_surface_list: for each set bit in the view's surface bitset, append the surface index and its three vertex indices (popcount/ctz instead of one bit at a time) Once per bearing over every surface of the level
00552DE0 leaf_bsp_walk_dead: the lightmap/material walk when its only callbacks are the bare ret at 0044AD80; computes only the frame and registers the original would leave 0050BFB0 walks the BSP nine times per bearing; the walks at 0050C523/0050C544 pass only that callback

All but the last are tried first in engine_dispatch_override through host_native_leaf_dispatch; 00552DE0 is checked after the A10 trace boundary so that trace still sees every walk. HALO_NATIVE_LEAVES=0 turns all six off (any other value, or unset, leaves them on).

Native visible-surface marking (HALO_NATIVE_VISIBILITY)

host_mark_visible_surfaces replaces 00553920. In the original:

  • 005537C0 floods portals from the camera cluster and leaves one 0x1A0-byte record per visible cluster at 007C3390 (count at 007D0390), each with the cluster index and a frustum narrowed to the portals it was seen through.
  • 00553920 (cdecl, argument: the structure BSP) tests every subcluster of each visible cluster against that frustum with 0050D5B0, and for each survivor sets its triangles' bits in the bitset at 007D0394, counting newly set bits in the int16 at 00850394, up to 0x4000 triangles. In PVS mode (structures_use_pvs_for_vs, 00724A45) or with the camera outside the BSP (007C3348 == -1) it uses the view frustum at 007C3168 instead of the portal frustum.

0050D5B0 is about 670 translated instructions, 400 of them emulated x87, run for a few hundred subclusters per view; before the x87 fast path it was the hottest guest function in the profiles. vs_box_test reproduces it for the flag value 00553920 passes: a six-compare world-bounds rejection (NaN never rejects), then eight corner outcodes against four side planes, each plane summed in the original's own order (((x·nx + y·ny) + nz·z) - d for the first, and so on), returning 2 (all inside), 0 (all outside one plane) or 1.

The native routine writes guest memory in the original's order: its frame, each call's argument and return address, 0050D5B0's frame whenever a box passes the bounds test (saved registers, loop counter, corner copies), the bitset words and the count. It rereads the BSP, records, frustums and counts after each write the original makes, so even a triangle id so far out of range that its bitset write lands in one of those inputs produces the same result. The single documented exception is a frustum pointer bent into 0050D5B0's own stack frame. Registers, flags of the final add esp, the last FCOMP condition bits and stale x87 slots are restored as the original leaves them. Non-nearest rounding or a deep x87 stack falls back to the translated routine. Its other caller (00458E49) always keeps the translated routine.

Modes (ov_visible_surfaces): unset or anything else = native; 0 = translated; verify = run the translated routine, snapshot its results, rewind CPU, bitset/count and both stack frames, run the native routine from the same state, compare registers, flags, x87 state, bitset, count and frames, log up to 16 mismatches (and a summary every 600 calls), and keep the translated result (ov_visible_surfaces_verify).

Native BSP surface gathers (HALO_NATIVE_GATHER)

ov_native_gather replaces the light and shadow surface-gather loops when HALO_NATIVE_GATHER=1 (default off on desktop; the visionOS app sets 1 since Build82, and an explicit 0 keeps the translated comparison path):

Address Native function What it does
005540C0 gather_native_leaf Enumerates one compact BSP leaf: decodes its bounds, tests overlap and planes (unless the caller passed "inside"), then for each surface whose bit is set in the view's visible-triangle bitset (007D0394) and not yet in the caller's visited set, marks it visited and appends it to the output list until the output capacity is reached.
00553C40 gather_native_clusters The same over cluster groups: for each listed cluster, each subcluster box that overlaps and passes the planes, append its visible, unvisited surfaces.
00553F10 gather_native_tree The recursive collision-BSP gather: classify the sphere against the node plane with the original x87 operation sequence, recurse into the front/back children (00553F10 again) or call the leaf gather (005540C0) for leaves, accumulating the count.

These follow the same exactness rules as the leaves, more strictly: they keep the original load/store order including redundant loads, argument slots, saved registers, child return addresses and abandoned stack locals, because output can alias input and because the test compares against the translated engine byte for byte. Only integer flag computations killed by the final add esp are omitted. Children (00553380, 005541B0, 00554260, 005540C0, 00553F10) are called through native implementations when they accept, and through normal translated dispatch when they decline. A call declines untouched when fewer than four x87 slots are free (so an overflow takes the full translated exception path), when bounded diagnostic execution is active (instruction_limit), or when the stack span is invalid.

Diagnostics: host_gather_get_diagnostics reports the effective switch and cumulative native and fallback entries (recursion is not counted separately); the visionOS app shows them through EngineGatherDiagnostics.swift (see Diagnostics and Telemetry).

Dispatch order

flowchart TD
    A["Translated call to address X reaches engine_dispatch_override"] --> B{"X in dispatch_interest bitmap?"}
    B -->|"no"| Z["Run translated function"]
    B -->|"yes"| C["host_native_leaf_dispatch: 004CC0D0, 00554260, 005541B0, 00553380, 00552C20"]
    C -->|"handled"| Y["Return to caller"]
    C -->|"not handled"| D{"005540C0 / 00553C40 / 00553F10 and HALO_NATIVE_GATHER=1?"}
    D -->|"native accepted"| Y
    D -->|"no / declined"| E["Campaign unlock, audio, model capture, panorama, A10 and HSC hooks"]
    E --> F{"00449590 int sort?"}
    F --> G{"00553920 and HALO_NATIVE_VISIBILITY != 0?"}
    G -->|"native or verify"| Y
    G -->|"no"| H{"00552DE0 dead walk and leaves on?"}
    H -->|"accepted"| Y
    H -->|"no"| I["Static override table, else run translated function"]
Loading

Environment variables

Variable Default Effect Read at
HALO_DRAW_FASTPATH off (visionOS 1) Resident vertex copies, folded clears, encoder reuse. metalrenderer.m:197
HALO_PV_SKIN_FAST off (visionOS 1) Exact-token skinning executor for ProcessVertices. d3d9_process_vertices.inc:19
HALO_PV_DIFFERENTIAL off Generic-vs-skin byte comparison (generic authoritative). d3d9_process_vertices.inc:5
HALO_PV_DIFFERENTIAL_FROM_FRAME 0 First frame compared; invalid text disables comparison. d3d9_process_vertices.inc:8
HALO_PROCESS_VERTICES_TRACE off Logs the first 20 calls and 8 successes. d3d9_render.inc:1060
HALO_NATIVE_LEAVES on (0 disables) Six native leaves. native_leaves.h:441
HALO_NATIVE_VISIBILITY native (0 translated, verify compare) Native 00553920. overrides.c:157
HALO_NATIVE_GATHER off (visionOS 1) Native gathers 005540C0, 00553C40, 00553F10. overrides.c:193

Tests

Test Run by run_source_checks.py What it asserts
test_vb_resident.c yes (macOS) Real d3d9.c with a fake renderer and a 1 MiB budget: every resident copy handed to a draw is byte-identical to what the arena path would copy (stream 0 and secondary streams); one copy per generation; Lock/Unlock and successful ProcessVertices invalidate; never while locked, never for dynamic, UP or unaligned streams; released with the buffer; clean slot reuse; churn fallback; fast-path toggle; NOOVERWRITE/DISCARD; state blocks; immediate fallback after an unlocked write.
test_draw_indices.c yes (macOS) The fast path equals a verbatim copy of the previous loop for all primitive types, 16/32-bit and invalid formats, negative and overflowing bases, vertex counts past 65536, and index ranges at the top of guest space.
test_process_vertices.c no Captured Halo vertices match independent bone transforms; UVs, preserved normals, offsets; a failing vertex leaves the whole destination unchanged; range validation cannot wrap.
test_process_vertices_differential.c no Skin executor byte-exact on captured vertices in four rounding modes; any token mutation falls back; FP flags preserved; generic authority, atomicity, range, alias, byte-cap and frame-gate checks.
test_native_leaves.c yes, twice (release flags and headset flags) Each leaf against its translated sub_*.c on random and adversarial inputs (NaN payloads, infinities, subnormals, signed zeros, aliasing, every walk shape) with guard pages around all reachable memory: same memory, registers, flags, return PC and complete x87 state, for all eight initial x87 TOP values; declines leave everything untouched. Needs the generated engine; prints SKIP otherwise.
test_visible_surfaces.c yes Native 00553920 against translated sub_00553920.c/sub_0050D5B0.c on randomized and edge cases (the 0x4000 cap at each check, int16 limits, negative counts, duplicate clusters, bitset writes landing in the inputs and frames); byte-for-byte memory, registers, flags and x87 state.
test_native_gather.c yes Leaf and cluster gathers against translated code with duplicate surfaces, masks, clipping, capacity exits, aliases, NaNs, signed counts, unusual rounding modes and pre-existing x87 registers; declines change nothing.
test_gather_pipeline.c yes The whole light/shadow gather 00553D80 with all callees on randomized BSP trees, leaves, clusters, portals and per-view bitsets (read-only BSP pages), plus every BSP of a supplied retail map (HALO_TEST_GATHER_MAP, default b30) when available; tree fallback cases.

Related pages

Clone this wiki locally