-
Notifications
You must be signed in to change notification settings - Fork 5
Textures and Texture Packs
Halo creates its textures through IDirect3DDevice9::CreateTexture and fills them with LockRect, exactly as on Windows; the pixel bytes stay in guest memory in their original D3D formats (mostly DXT). When a draw samples a texture, the Direct3D 9 Bridge decodes level 0 to BGRA8 on the CPU, builds a mip chain, and uploads it into a content-keyed cache in the Metal Renderer. Before that, an optional native texture pack (.hvt) can replace the texture wholesale with higher-resolution artwork, matched by the CRC32 of its original level-0 bytes.
| File | Role |
|---|---|
d3d9.c |
Texture object creation, LockRect/UnlockRect, level aliases, content generations. |
d3d9_render.inc |
uploaded_texture, uploaded_cube_texture, uploaded_volume_texture: flush, key, decode, upload. |
texture_decode.c / .h
|
D3D surface bytes to tightly packed BGRA8 (uncompressed and DXT). |
texture_mips.h |
Exact rounded box-filter mip generation. |
texture_content_hash.h |
XXH3-64 content hashing for cache keys. |
third_party/xxhash |
Vendored xxHash v0.8.3 (xxhash.h, BSD-2-Clause; UPSTREAM.json records the source URL and SHA-256). |
texture_mod_pack.h |
.hvt format parsing, binary search, IEEE CRC32 (hardware-accelerated on ARM64). |
texture_mod_runtime.inc |
Pack mapping and per-texture replacement lookup. |
metalrenderer.m |
Texture table, key slots, LRU eviction, cached 2D/cube/volume creation, mip upload. |
flowchart TD
A["Draw samples stage N (gpu_program_draw / ff3d_draw / RHW path)"] --> B{"Texture kind"}
B -->|"2D"| C{"Render-target texture whose level 0 is GPU-dirty?"}
C -->|"yes"| D["mr_target_texture: sample the GPU target directly"]
C -->|"no"| E["flush_data: read back any GPU-dirty surface sharing level-0 storage"]
E --> F{"Texture pack loaded and texture not a render target?"}
F -->|"yes"| G{"mod_checked for this content_generation?"}
G -->|"no"| H["CRC32 of level-0 guest bytes, binary search the pack"]
H --> I["Remember mod_entry for this generation"]
G -->|"yes"| I
I --> J{"Match?"}
J -->|"yes"| K["Key = 0x48565458_00000000 OR crc, find or upload replacement BGRA (with mips)"]
K -->|"upload ok"| Z["Texture id for the sampler"]
K -->|"upload failed"| L
J -->|"no"| L
F -->|"no"| L["Key: reuse content_key if hashed_generation == content_generation, else XXH3 over guest, format, w, h, level-0 bytes"]
L --> M{"mr_texture_find_cached(key)?"}
M -->|"hit"| Z
M -->|"miss"| N["halo_texture_decode level 0 to BGRA8"]
N --> O["mr_texture_create_cached (with CPU mips) or _nomip for render targets"]
O --> Z
B -->|"cube"| P["All six faces level 0: flush, XXH3 key, decode, mr_texture_create_cube_cached"]
B -->|"volume"| Q["Level 0 slices: flush, XXH3 key, decode, mr_texture_create_volume_cached (no mips)"]
P --> Z
Q --> Z
-
create_texturerecords width, height, depth, levels (0 means the full chain), usage, format and pool; it allocates nothing. Limits: 8192x8192, depth 512, 16 levels,pool <= 3. - Level storage is allocated from guest pages on first
LockRectorGetSurfaceLevel(method_texture). Sizes come fromlevel_size: DXT1 isceil(w/4)·ceil(h/4)·8bytes, DXT2-5 the same with 16-byte blocks, everything else width × bits-per-pixel. -
content_generationadvances onLockRect,UnlockRect,AddDirtyRect, any lock/unlock of a level-alias surface,surface_copy/ColorFillinto it, and every GPU readback into its storage (content_changedpropagates a level's change to its owner texture). This generation is what lets upload skip rehashing an unchanged texture. - Render-target textures (
D3DUSAGE_RENDERTARGET, usage bit 1) are drawn into through their level-0 surface alias. When that surface's GPU copy is newer, sampling uses the target directly (mr_target_texture), with no readback and no re-upload.
Only level 0 is consumed. Lower mip levels the game writes are never read by the upload path; the renderer regenerates the chain from level 0 instead.
halo_texture_decode is a heap-free, thread-safe C99 decoder that writes width·height·4 bytes in B, G, R, A order (a little-endian A8R8G8B8 pixel), with output pitch exactly width·4. Input pitch may be larger than a row (0 means tight); the last row only needs its own bytes; edge blocks of compressed textures are cropped.
| D3DFORMAT | Value | Decoding |
|---|---|---|
A8R8G8B8 |
21 | Copy. |
X8R8G8B8 |
22 | Copy, alpha forced to 255. |
R5G6B5 |
23 | Expand 5/6/5 by bit replication; alpha 255. |
A1R5G5B5 |
25 | Expand 5/5/5; alpha 0 or 255. |
A4R4G4B4 |
26 | Expand each nibble (`v<<4 |
A8 |
28 | Black, alpha from the byte. |
L8 |
50 | Grey, alpha 255. |
A8L8 |
51 | Grey from L, alpha from A. |
DXT1 |
'DXT1' |
4-colour mode when c0 > c1, else 3 colours + transparent black. |
DXT2 / DXT3
|
'DXT2'/'DXT3'
|
Explicit 4-bit alpha; colour block always 4-colour. DXT2 is decoded as DXT3 and keeps its stored premultiplied RGB. |
DXT4 / DXT5
|
'DXT4'/'DXT5'
|
Interpolated 8- or 6-step alpha; DXT4 decoded as DXT5, premultiplied RGB kept. |
Palette interpolation uses integer (2a+b)/3, (a+b)/2 and the /7, /5 alpha ramps. Dimensions must be 1-8192. Errors are E_FORMAT, E_DIMENSION, E_NULL, E_PITCH, E_SOURCE_SIZE and E_OUTPUT_SIZE; any failure makes the draw fail with texture upload rather than drawing with a wrong texture. Formats the bridge can size but the decoder does not know (for example X1R5G5B5, 24, or any other FOURCC) therefore cannot be sampled.
Cached 2D and cube textures are created with a full mip chain, filled once at upload (upload_mip_chain) with replaceRegion per level. The comment explains the choice: cached textures are immutable, so building levels once avoids a Metal command buffer per texture and keeps transient driver allocations bounded during a large level load.
halo_texture_mip halves each dimension (minimum 1) with an exact rounded box filter: for each destination pixel it averages the source rectangle [x·w/nw, (x+1)·w/nw) × [y·h/nh, (y+1)·h/nh) per channel with (sum + count/2) / count. This handles odd and non-power-of-two sizes and padded rows. When both dimensions are even and greater than 1, a fast path sums four pixels in two 32-bit registers holding independent 16-bit lanes (R/B and G/A), so there are no carries between channels; it produces the same bytes.
Exceptions: render-target textures are uploaded without mips (mr_texture_create_cached_nomip), because they are re-sampled at 1:1 every frame and CPU mip generation dominated frame time; volume textures have level 0 only.
Byte accounting includes the whole chain (mip_chain_bytes), roughly 4/3 of the base level (a 256x256 texture accounts 349524 bytes).
Cache keys are computed by uploaded_texture with halo_texture_hash_more, which is XXH3_64bits_withSeed(bytes, size, seed) from the vendored xxHash v0.8.3, compiled header-only (XXH_INLINE_ALL). The key is a seeded chain:
key = 1469598103934665603 (FNV offset basis used as the initial seed)
key = XXH3(&guest_address, 4, key)
key = XXH3(&format, 4, key)
key = XXH3(&width, 4, key)
key = XXH3(&height, 4, key) (cube: no height; volume: also depth)
key = XXH3(level-0 bytes, size, key) (cube: each face in turn; volume: all slices)
if key == 0: key = 1
- The guest address is part of the key, so two textures with identical bytes at different addresses do not share a cache entry (texture-pack replacements, by contrast, are keyed by CRC alone and are shared).
- The key is recomputed only when
content_generationhas changed since the last hash; otherwise the storedcontent_keyis reused. Any write through a host API therefore produces a new key, and an unchanged texture costs no hashing. - Before hashing,
flush_datareads back any render surface whose storage is the texture's level 0 and whose GPU copy is newer, so a texture that aliases a render target hashes its current pixels. A failed readback fails the upload.
The header comment explains why XXH3 replaced FNV: keys still hash every source byte (including readbacks) rather than relying on dirty tracking, and XXH3 removes the serial per-byte FNV multiply from that hot path. Note that the comment's "keep hashing every source byte" refers to what goes into a key when one is computed; reuse across draws is governed by content_generation as described above.
The cache lives in the renderer's shared state, so all contexts share it.
| Structure | Purpose |
|---|---|
textures[4096] |
Slot table; id 0 is "none". Holds cached textures, uncached textures, and render-target aliases (uncached, never evicted). |
texture_keys[], texture_use[], texture_bytes[]
|
Key (0 for uncached), LRU tick, accounted bytes per slot. |
key_slots[16384] |
Open-addressed key-to-id index (key_slot_insert/remove); removal re-inserts the rest of the probe run. |
cached_texture_count, cached_texture_bytes
|
Totals checked against the budgets. |
mr_texture_find_cached updates the LRU tick on a hit. texture_create_cached_impl (and the cube/volume variants) checks the cache again, rejects a single texture larger than the byte budget, makes room, takes a free slot (evicting once more if the table is full), creates a shared-storage BGRA8Unorm texture, uploads level 0 and the mip chain, and inserts the key.
- At most 1024 cached textures and 512 MiB of cached bytes (
make_cached_texture_room). -
evict_oldest_cached_textureevicts the least recently used cached entry. Dropping the cache's reference is safe without waiting for the GPU because the shared command buffer retains every texture already encoded. -
Binding pin. The bridge brackets each draw with
mr_texture_bindings_begin/end. While a scope is open, entries used after the scope began (texture_use > texture_binding_floor) cannot be evicted, so a texture id resolved for stage 0 is still valid when stage 3's upload needs room. If nothing evictable remains, the new upload fails withtexture cache pinned by current drawand the draw is rejected rather than drawing with a recycled id. - Cube textures are 6-slice mipmapped; volume textures are limited to 2048 per side, level 0 only.
mr_texture_cache_counters reports hits, uploads and evictions; mr_cached_texture_count/bytes the current totals.
texture_mod_initialize runs once (pthread_once) on the first texture upload. If HALO_TEXTURE_MODS is not 0 and HALO_TEXTURE_PACK names a file, it mmaps the file read-only (MAP_PRIVATE), validates it, and allocates a "seen" byte per entry. Packs are demand-paged rather than copied into resident memory: full-game artwork exceeds 2 GiB, so the file limit is 4 GiB. Any failure logs a reason (pack unavailable, unsupported pack size, invalid pack) and the original textures are used. The visionOS app sets HALO_TEXTURE_PACK to the bundled TextureMods.hvt when it exists.
All integers little-endian (htm_open):
| Offset | Size | Field | Rule |
|---|---|---|---|
| 0 | 8 | magic | HVTEX001 |
| 8 | 4 | entry count | 1-65536 |
| 12 | 4 | entry size | must be 32 |
| 16 | 8 | directory offset | must be 32 |
| 24 | 8 | file size | must equal the actual size |
| 32 | 32·count | directory entries | see below |
| ... | pixel data |
Directory entry (32 bytes):
| Offset | Size | Field | Rule |
|---|---|---|---|
| 0 | 4 | CRC32 key | strictly increasing across entries |
| 4 | 4 | width | 1-8192 |
| 8 | 4 | height | 1-8192 |
| 12 | 4 | format | must be 1 (BGRA8) |
| 16 | 8 | data offset | at or after the end of the previous entry's data (or the directory), within the file |
| 24 | 8 | data bytes | exactly width·height·4
|
Entries must be sorted by key with no duplicates, and pixel blocks may not overlap. The replacement is one BGRA8 level; mips are generated at upload like any other cached texture.
The key is "TexMod's uncomplemented IEEE CRC32 of mip zero": the reflected IEEE polynomial 0xEDB88320, initial value 0xFFFFFFFF, and no final XOR, computed over the raw guest bytes of level 0 in the texture's original D3D format (for DXT textures, the compressed blocks). htm_crc uses the ARMv8 __crc32d/__crc32b instructions eight bytes at a time when the target advertises the CRC extension (the IEEE polynomial, not CRC32C), and a table-driven scalar loop otherwise. host_texture_crc_set_fast(0) lets a diagnostic probe force the scalar path for comparison.
uploaded_texture_mod runs inside uploaded_texture after the GPU-alias check and before content hashing:
- Skip if no pack is loaded or the texture is a render target.
- If not yet checked for the current
content_generation, compute the CRC, binary search the directory (htm_find) and store the one-based entry index (0 for no match). A texture rewritten by the game is re-checked, so modified contents stop matching. - On a match, use cache key
0x4856545800000000 | crc("HVTX"in the high word), find or create the cached BGRA texture from the mapped pixels, and log the first match of each entry with original and replacement sizes. - If the upload fails (for example the budget is exhausted), return 0 so the original texture is decoded and used instead. A replacement can therefore never turn a texture black.
The original guest pixels are never modified. How packs are produced is described in Visual Mods Pipeline.
| Variable | Default | Effect | Read at |
|---|---|---|---|
HALO_TEXTURE_PACK |
unset (visionOS: bundled TextureMods.hvt) |
Path of the .hvt pack. |
texture_mod_runtime.inc:17 |
HALO_TEXTURE_MODS |
enabled |
0 ignores the pack. |
texture_mod_runtime.inc:17 |
HALO_FF3D_DUMPTEX |
unset | Directory to dump the first 40 decoded textures (tex_<guest>_<w>x<h>.bgra). |
d3d9_render.inc:343 |
HALO_NO_ANISO, HALO_TEX_FALLBACK
|
unset | Sampler quality override and missing-texture fallback (see Metal Renderer). | metalrenderer.m |
| Test | Run by run_source_checks.py
|
What it asserts |
|---|---|---|
test_texture_mod_pack.c |
yes | The native CRC equals the scalar reference for every length 0-8192 at 16 alignments, exact-size buffers 1-256 (for ASan over-read detection) and 4 MiB; lookups; bounds, truncation, overlap and duplicate rejection. |
test_d3d9_texture_mod.c |
yes | Through the real D3D upload boundary: the replacement pixels are uploaded while guest bytes stay unchanged; a cached match is not re-uploaded; changing the guest texture (and its generation) stops matching; restoring it matches again; eviction re-uploads; an upload failure returns 0 so the original path is used; render targets are never replaced. |
test_texture_mips.c |
yes |
halo_texture_mip equals a reference box filter for all 65x65 size combinations with padded rows and for a 2048x2048 image; reports timing. |
test_texture_content_hash.c |
no | Fixed XXH3 vectors; unaligned inputs and size boundaries hash identically; changing guest address, format, width, height or any of the first/middle/last byte changes the key. HALO_HASH_BENCH adds an FNV comparison. |
test_d3d9_texture_cache.c |
no | 2D and cube uploads through the bridge with a fake cache; GPU-alias readback and readback failure. It also asserts that flipping a byte without a generation change yields a new key; the current uploaded_texture reuses content_key while the generation is unchanged, so this standalone fixture appears to predate generation-based key reuse. |
test_d3d9_texture_lifetime.c |
yes (macOS) | Parent/level alias ownership for 2D, cube and volume textures. |
test_metalrenderer_texture_bindings.m |
yes (macOS) | The binding pin, budget bound and post-scope eviction on real Metal. |
test_metalrenderer_mips.c, test_d3d9_texture_upload.m
|
no | Mip accounting and filtering; cached uploads copy CPU bytes and survive cache removal while queued. |
Documents master-chef at commit 9f915af (v1.0.3). Unofficial project, not affiliated with Microsoft, Bungie, Gearbox or Apple. Original code is MIT licensed; game content is not included.
Overview
- Architecture Overview
- Repository Layout
- Glossary
- Environment Variables
- Contributing Guide
- Open Questions
Translation
- Static Translation Pipeline
- XWA Decoder and Lifter
- Function Address Lists
- EngineReuse Runtime
- x87 Floating Point
Host runtime
- EngineHost Overview
- Win32 Compatibility Layer
- Threading and Synchronization
- Guest Memory and Heap
- Engine Overrides and Hooks
- Runtime Settings
Graphics
- Direct3D9 Bridge
- Metal Renderer
- Shader Translation
- Textures and Texture Packs
- Geometry Fast Paths
- Radial Fog
Panorama and presentation
- Panorama System
- Panorama Budget and LOD
- Frame Pacing
- visionOS App
- Immersive Presenter
- Layer Alignment
Audio and input
Tooling and process