Skip to content

Metal Renderer

M T edited this page Oct 4, 2026 · 1 revision

Metal Renderer

metalrenderer.m is the plain-C-API Metal rasteriser underneath the Direct3D 9 Bridge. It owns one Metal device, command queue and set of caches shared by every D3D render-target surface; each D3D surface that is drawn to gets an mr_context, which is just a colour target plus a Depth32Float_Stencil8 target viewed onto that shared state. Draws from all contexts are encoded into one shared command buffer in submission order, so a render-target texture can be sampled by a later draw without a CPU readback. It compiles unchanged for macOS and visionOS (Metal + Foundation only); the panorama code publishes its targets into the visionOS presenter through the blit/FXAA entry points described below.

Source files

File Role
metalrenderer.h Public C API: vertex formats, blend enum, draw-state and program-state structs, fixed-stage helpers, every entry point.
metalrenderer.m Implementation: shared state, command buffer/encoder management, texture cache, draw entry points, generated fixed-function shaders, pipeline/depth/sampler caches, the pipeline compiler service, readback/blit/FXAA/commit.
metalshader.c / .h D3D shader bytecode to MSL via MojoShader (see Shader Translation).
texture_mips.h CPU box-filter mip generation used at upload (see Textures and Texture Packs).
radial_fog.h Vertex-shader fog rewrite applied during translation (see Radial Fog).

Limits and constants

Constant Value Meaning
MR_MAX_VERTS 65536 Vertices per draw.
MR_MAX_INDICES 196608 16-bit indices per draw.
MR_MAX_TEX 4096 Texture table slots (cached textures, uncached textures and render-target aliases).
MR_MAX_CACHED_TEXTURES 1024 Content-keyed cache entries.
MR_MAX_CACHED_TEXTURE_BYTES 512 MiB Cached texture bytes including mip chains (overridable at compile time).
MR_MAX_PROGRAMS / MR_PROGRAM_SLOTS 4096 / 8192 Programmable pipeline entries / open-addressed index.
MR_MAX_FIXED_RHW_PROGRAMS / MR_FIXED_RHW_SLOTS 2048 / 4096 Fixed-function and pretransformed pipeline entries / index.
MR_MAX_PROGRAM_SAMPLERS 256 Sampler-state cache entries.
depth/stencil cache 64 Depth-stencil-state entries.
custom RHW blend cache 32 Pipelines for raw-factor blending on the simple RHW shader.
MR_MAX_VERTEX_STRIDE 256 Bytes per vertex per stream.
MR_ARENA_BYTES / MR_ARENA_POOL 32 MiB / 3 Transient vertex/index arena per command buffer and the number kept for reuse.

Defined at metalrenderer.m:15-28 and :332-333.

Shared state and contexts

mr_shared is created once by shared_state: the system default device, a command queue, the built-in library compiled from kSrc (a simple RHW vertex/fragment pair), four legacy samplers (nearest/linear x wrap/clamp), a 1x1 white texture, 1x1 opaque-black 2D/cube/3D textures (D3D samples an unbound stage as opaque black), and the pipeline compiler service, which immediately opens the on-disk pipeline store and queues the built-in fixed vertex functions for prewarming.

An mr_context holds:

  • target (BGRA8, shared storage, render target + shader read) and depth_target (Depth32Float_Stencil8, private storage), created by make_target. The colour target also occupies an uncached slot target_id in the shared texture table, so other draws can sample it by id.
  • The current viewport (vx, vy, vw, vh), clamped to the target by mr_set_viewport. It is applied as the scissor rectangle; programmable and pretransformed draws use a full-target Metal viewport, while clip-space draws use it as the Metal viewport.
  • overlay_target, set by mr_set_overlay_target for the panorama HUD layer. Overlay pipelines accumulate coverage in alpha (source-alpha blending writes srcA·1 + dstA·(1-srcA) to alpha) instead of D3D's world semantics.
  • Pending clears (see below) and a triangle counter.

mr_create accepts 1-8192 per side and initialises the target to opaque black and depth 1.0. mr_resize waits for pending GPU work and replaces only the targets; caches survive (this is what D3D Reset relies on). mr_destroy flushes and waits, then frees the table slot.

Threading: the API is single-threaded by contract ("call all functions from the same thread", metalrenderer.h:15), which is the engine thread. The exceptions are the pipeline compiler (its own lock and worker threads), the arena pool (os_unfair_lock, returned from Metal completion handlers), the fast-path switch and traffic counters (atomics readable from the diagnostics thread), and mr_commit_async completion callbacks.

Command buffers, encoders and the arena

sequenceDiagram
    participant D as D3D bridge on the engine thread
    participant R as metalrenderer
    participant CB as Shared pending command buffer
    participant GPU as GPU
    D->>R: mr_clear and mr_clear_depth (deferred when fast paths are on)
    D->>R: mr_draw into context A
    R->>CB: begin_draw_encoder for A ends any other encoder, encodes other contexts' clears, opens a pass with A's clears as load actions
    R->>CB: bind only changed state, copy vertices to the arena, draw
    D->>R: mr_draw into context B, sampling A by target id
    R->>CB: end A's encoder and open B's pass
    D->>R: mr_fxaa_target_to or mr_blit_target_to for panorama publication
    R->>CB: render or blit encoder
    D->>R: mr_commit_async with a done callback
    R->>CB: retire the arena to the pool on completion, then commit
    CB->>GPU: execute in submission order
    GPU-->>R: completion handler calls done(arg, ok)
Loading
  • One command buffer (pending_cb) collects all work until a commit, readback or mr_resize. It is created with retainedReferences = YES (ensure_command_buffer), so every per-draw buffer and texture stays alive until the GPU finishes; this is why cached textures and resident vertex copies can be released by the CPU while draws that use them are still queued.
  • One open encoder (pending_encoder) for whichever context is being drawn into. begin_draw_encoder reuses it when the target is unchanged; otherwise it ends it and opens a pass that attaches colour, depth and stencil. Before any draw it encodes the waiting clears of every other context, because a target id acquired earlier may be sampled by this draw.
  • Arena: vertices and indices are copied into a 32 MiB shared MTLBuffer at 256-byte aligned offsets (arena_alloc). When it fills, or at commit, it is retired to a three-buffer pool once the command buffer completes. The comment explains why: a fresh 32 MB buffer every frame faulted and zero-filled every page inside the draw copies. Draws larger than the arena get a dedicated buffer that is not pooled.
  • flush_pending commits and optionally waits (waitUntilCompleted), turning a command-buffer error into MR_ERR_DEVICE with the Metal error text.

Fast paths (HALO_DRAW_FASTPATH)

mr_fast_paths_enabled reads HALO_DRAW_FASTPATH once (1 enables; default off on desktop, set to 1 by the visionOS app, which also calls mr_set_fast_paths on the engine thread before any draw). Three shortcuts depend on it, all designed to feed the GPU exactly what the plain path does:

  1. Folded clears. mr_clear, mr_clear_depth and mr_clear_stencil store the value and add the context to clear_list (defer_clear) instead of encoding a pass that clears, stores, and is immediately loaded again. The next pass that draws into the context uses the clear as its load action (describe_pass). Any other consumer (sampling via mr_target_texture, blits, FXAA, uploads, readback, commit) first encodes the waiting clears as a pass of their own, so ordering is preserved. With fast paths off every clear is its own pass.
  2. Encoder state reuse. bind_draw_state and bind_fragment_texture skip setRenderPipelineState, setDepthStencilState, stencil reference, viewport, scissor, cull mode, winding, and per-slot texture/sampler calls whose value the open encoder already holds. The record (mr_bound_state) is kept either way and forgotten when an encoder ends, so the switch can change mid-encoder.
  3. Resident vertex copies. resident_range lets mr_draw_program bind a persistent MTLBuffer (created by mr_buffer_create) at a 4-byte-aligned offset instead of copying the stream into the arena. Ownership and correctness rules are on Geometry Fast Paths.

mr_draw_traffic_stats reports cumulative bytes bound from resident copies, bytes copied to the arena, and clears folded into a following pass.

Draw entry points

Function Input vertices Pipeline Used by
mr_draw_rhw mr_vertex_rhw (x, y, z, rhw, colour, u, v), stride at least 28 Built-in v_main/f_main, one cached PSO per (overlay, blend, write mask), or a custom-blend PSO Bridge RHW path without pixel shader
mr_draw_program Raw stream bytes described by a D3D declaration (stream 0 plus secondary streams) Translated VS + translated PS, or translated VS + generated fixed-function fragment Bridge programmable path
mr_draw_fixed_rhw mr_vertex_fixed_rhw (screen-space, colour, specular, 8 UV sets, fog) Generated RHW vertex + generated fixed fragment Tests; same core as below
mr_draw_program_rhw mr_vertex_fixed_rhw Generated RHW vertex + translated PS Bridge RHW path with pixel shader (HUD, screen effects)
mr_draw_fixed_clip mr_vertex_fixed_clip (homogeneous clip position) Generated clip-space vertex + generated fixed fragment Bridge FF3D path

The last three share mr_draw_fixed_common.

Common rules: nv 1-65536, ni at most 196608, every index below nv (checked with a vectorisable maximum, indices_below), primitive one of triangle list/strip, line list/strip or point list (mr_primitive_supported); fewer than 3 vertices/indices or an empty viewport is a successful no-op. Nothing is drawn on error. Error codes: MR_ERR_ARGS, MR_ERR_BOUNDS, MR_ERR_PRIMITIVE, MR_ERR_TEXTURE, MR_ERR_DEVICE, MR_ERR_SHADER, MR_ERR_DECL, MR_ERR_UNSUPPORTED, MR_ERR_STATE, with text in mr_last_error().

Note that the D3D bridge currently passes only triangle topologies through (see Direct3D 9 Bridge); the line/point support in the renderer is reachable through the C API.

Programmable draw, step by step

flowchart TD
    A["mr_draw_program(state, prim, vertices, stride, nv, indices, ni)"] --> B["Validate args, blend, mask, cull, bounds"]
    B --> C["depth_stencil_state_for (64-entry LRU cache)"]
    C --> D["program_for: key = program_key(state, stride, overlay)"]
    D --> E{"Cache entry state"}
    E -->|"READY"| H
    E -->|"BUILDING by worker"| F["Wait for that pipeline only"]
    F --> E
    E -->|"missing / QUEUED / retryable FAILED"| G["program_build on engine thread"]
    G --> H["Bind stream 0: resident copy or arena copy"]
    H --> I["Secondary streams used by the declaration: resident or arena"]
    I --> J["pack_float_uniforms for VS and PS"]
    J --> K["sampler_for per sampled stage (256-entry LRU cache)"]
    K --> L["begin_draw_encoder, bind_draw_state"]
    L --> M["Vertex buffers at 1+stream, VS constants at vertex buffer 0, PS constants at fragment buffer 0, extras at fragment buffer 1"]
    M --> N["Bind textures (black fallback for unbound or mismatched kinds)"]
    N --> O["drawIndexedPrimitives / drawPrimitives"]
Loading
  • Vertex layout. make_vertex_descriptor matches every attribute the translated VS consumes (usage, usage index) against the D3D declaration and maps D3D declaration types to Metal vertex formats (vertex_format): FLOAT1-4, D3DCOLOR as UChar4Normalized_BGRA, UBYTE4, SHORT2/4, UBYTE4N, SHORT2N/4N, USHORT2N/4N, FLOAT16_2/4. UDEC3/DEC3N are not supported. Only D3DDECLMETHOD_DEFAULT and streams 0-15 are accepted; an element must fit inside its stream's stride; a missing semantic or missing end marker fails the draw. Stream s is bound at vertex-buffer index 1 + s.
  • Constants. pack_float_uniforms copies only the register ranges MojoShader reports as used, in its packing order, into a contiguous float4 array. Integer and boolean constants are rejected (integer/bool shader constants are unsupported).
  • Pixel-shader samplers. ps_sampler_texture binds the stage's texture only if its kind matches the shader's declared sampler type; an unbound or mismatched stage gets opaque black of the declared kind (as the D3D9 runtime does) instead of failing the draw.
  • Fixed-function stages without a pixel shader must have their textures available; otherwise the draw fails with fixed stage N texture is unavailable, unless the diagnostic HALO_TEX_FALLBACK is set, in which case the 1x1 white texture is bound (white on a lightmap stage reads as fully lit).

Pipeline caches

Pipelines are immutable and keyed by everything that affects the pipeline descriptor; per-draw values (constants, textures, stencil reference, viewport) are not part of the key.

Cache Key Capacity / policy
pso[2][MR_BLEND_COUNT][16] (and the retained pso_d split) overlay, blend mode, write mask Fixed array; built on demand by pso_for2. All passes attach depth now, so the depth/no-depth split only matters for cache identity.
custom_rhw[32] src, dst, op, mask, overlay Linear search; full cache fails the draw ("custom blend pipeline cache is full"). pso_custom.
programs[4096] program_key: VS key (token hash), PS key, declaration bytes, stream-0 stride, all 16 secondary strides, overlay, blend (+ raw factors/op when custom), write mask, alpha-test enabled (+ function), fog enable, radial-fog mode when non-zero, sampler kinds of bound stages, and the eight fixed stages when there is no PS Open addressing in 8192 slots. Entries are never evicted; a full cache fails new pipelines ("programmable pipeline cache is full"). Background prewarm stops at 3/4 capacity so the engine always has room.
fixed_rhw_programs[2048] fixed_rhw_key: PS key or fixed stages, has-PS, blend (+ raw), mask, sampler kinds, alpha test, overlay, clip-space flag, and fog/radial mode only when radial fog is active Same policy, 4096 slots.

A key of 0 is mapped to 1 so 0 can mean "empty". Pipelines are built by the compiler service described in Shader Translation: functions are compiled once per distinct MSL text and shared by every variant, background workers prewarm from shader creation and from the manifest of earlier sessions, and a draw whose pipeline is not ready builds it (or waits for the single worker already building it). Draws are never skipped because a pipeline is late.

test_metalrenderer_overlay_cache.m verifies that world and HUD (overlay) pipelines never share a cache row in either creation order.

Blend modes

configure_program_color sets BGRA8Unorm, the write mask, and the factors:

Mode RGB src / dst Alpha src / dst
MR_BLEND_NONE blending off blending off
MR_BLEND_SRC_ALPHA SrcA / 1-SrcA SrcA (overlay: One) / 1-SrcA
MR_BLEND_ADD One / One One / One
MR_BLEND_MODULATE DstColor / Zero DstA / Zero
MR_BLEND_MODULATE2 DstColor / SrcColor DstA / SrcA
MR_BLEND_SRC_ALPHA_ZERO SrcA / Zero SrcA / Zero
MR_BLEND_SRC_ALPHA_ADD SrcA / One SrcA / One
MR_BLEND_PREMULTIPLIED One / 1-SrcA One / 1-SrcA
MR_BLEND_DEST_ALPHA_ADD DstA / One DstA / One
MR_BLEND_CUSTOM raw D3D factors via d3d_blend_factor same factors with colour factors mapped to their alpha forms; op via d3d_blend_operation (ADD, SUBTRACT, REVSUBTRACT, MIN, MAX)

D3DBLEND_BLENDFACTOR/INVBLENDFACTOR map to Metal's blend colour factors; the bridge never sets a blend colour, so they use Metal's default. The RGB-only write mask (7) used by Halo's destination-alpha passes keeps destination alpha unchanged.

Fixed-function fragment stages

When no pixel shader is bound, make_fixed_fragment generates an MSL fragment function that evaluates the D3D9 texture-stage cascade for up to eight stages, stopping at the first D3DTOP_DISABLE. The stage parameters are part of the pipeline key, so each combination is compiled once.

  • Inputs. COLOR0/COLOR1 and TEXCOORD0..7 are declared only if the vertex function outputs them (for clip-space and RHW draws a synthetic description provides all of them). diffuse defaults to white when the vertex stage has no COLOR0; specular defaults to zero.
  • Textures. A stage declares texture2d or texturecube (volume textures are rejected in fixed stages) only when the stage actually consumes D3DTA_TEXTURE. mr_fixed_op_arg_mask defines which operands an operation reads (SELECTARG1 only arg1, SELECTARG2 only arg2, MULTIPLYADD/LERP all three, everything else arg1+arg2); BLENDTEXTUREALPHA and BLENDTEXTUREALPHAPM always read the texture. Stale values left in unused argument slots therefore never create a texture or TEXCOORD dependency. A stage that does sample requires the vertex stage to provide its coordinate set and rejects generated coordinates or texture transforms at this level.
  • Arguments (fixed_arg): DIFFUSE, CURRENT, TEXTURE, TFACTOR (uniform), SPECULAR, TEMP, CONSTANT (per-stage literal), with the COMPLEMENT (0x10) and ALPHAREPLICATE (0x20) modifiers.
  • Result. D3DTSS_RESULTARG may be CURRENT or TEMP. Each stage writes saturate(float4(rgb, alpha)). If the alpha op is DISABLE alpha passes current.a through.
  • After the cascade: vertex fog (mix(fog_color, current.rgb, saturate(fog))) when fog is enabled and the vertex stage outputs fog, then the D3D alpha test with any of the eight D3DCMP functions against ALPHAREF.

Supported D3DTEXTUREOPs (fixed_op):

Op Value Expression
DISABLE 1 current
SELECTARG1 / SELECTARG2 2 / 3 a1 / a2
MODULATE / 2X / 4X 4 / 5 / 6 a1·a2 (·2, ·4)
ADD 7 a1+a2
ADDSIGNED / 2X 8 / 9 a1+a2-0.5 (·2)
SUBTRACT 10 a1-a2
ADDSMOOTH 11 a1+a2·(1-a1)
BLENDDIFFUSEALPHA / TEXTUREALPHA / FACTORALPHA / CURRENTALPHA 12 / 13 / 14 / 16 lerp by that alpha
BLENDTEXTUREALPHAPM 15 a1+a2·(1-texture.a)
MODULATEALPHA_ADDCOLOR 18 a1.rgb + a1.a·a2 (colour only)
MODULATECOLOR_ADDALPHA 19 a1·a2 + a1.a (colour only)
MODULATEINVALPHA_ADDCOLOR 20 a1 + (1-a1.a)·a2 (colour only)
MODULATEINVCOLOR_ADDALPHA 21 (1-a1)·a2 + a1.a (colour only)
DOTPRODUCT3 24 dot(2a1-1, 2a2-1), replicated
MULTIPLYADD 25 a1·a2+a0
LERP 26 a0·a1+(1-a0)·a2

PREMODULATE (17), the bump-mapping ops (22, 23) and ops 18-21 used as alpha operations fail pipeline creation with fixed stage N uses unsupported color/alpha operation. test_fixed_modulate_ops.m checks ops 18/20/21 against independently computed D3D equations on GPU pixels.

Uniforms for the fixed fragment (MRFixedUniforms): TEXTUREFACTOR as float4, the alpha reference, padding, and the fog colour (fixed_uniforms). With radial fog active the padding is spelled as three floats so the fog colour sits at byte 32 as the CPU writes it; the original float3 padding layout (which Metal aligns to 16, placing the colour at byte 48 in the generated struct) is kept for existing pipelines when radial fog is off.

Pretransformed and clip-space vertex functions

fixed_rhw_vertex_source generates the vertex stage for the three mr_vertex_fixed_* paths. In RHW mode it converts pixel coordinates to NDC and multiplies by w = 1/rhw (guarding |rhw| <= 1e-20) so colour and UVs interpolate perspective-correctly; in clip-space mode it passes the position through. It outputs both colours and all eight TEXCOORDs, plus fog when radial fog is active. The GPU vertex record is 112 bytes, or 128 bytes when it carries fog.

Depth/stencil state cache

depth_stencil_state_for maps D3DCMP 1-8 and D3DSTENCILOP 1-8 directly (incr/decr saturate and wrap, invert, replace, zero, keep); the same stencil descriptor is used for front and back faces. Every draw binds a depth state, even a disabled one (compare ALWAYS, no write), because encoders are shared across D3D draws and the previous draw's state must not leak. Inactive stencil fields are normalised so materials that leave different stencil values behind do not fill the cache with equivalent states.

Pressure rule: 64 entries; when full, the least-recently-used entry (by state_tick) is replaced. Replacing a cache reference is safe because encoders and the bound-state record retain the Metal objects already in use. test_metalrenderer_state_pressure.m cycles 192 distinct states and checks the cache stays at 64 while queued pixels remain correct.

Sampler state cache

sampler_for translates mr_program_sampler (filled from D3DSAMP_* by the bridge's program_sampler_settings):

D3D D3DTADDRESS Value Metal address mode
WRAP 1 Repeat
MIRROR 2 MirrorRepeat
CLAMP 3 ClampToEdge
BORDER 4 ClampToZero (only with border colour 0; any other border colour fails with MR_ERR_UNSUPPORTED)
MIRRORONCE 5 MirrorClampToEdge

U and V are independent (address_u sets S and R, address_v sets T). An address of 0 means WRAP.

Filters: D3DTEXF_POINT (1) → nearest, LINEAR (2) and ANISOTROPIC (3) → linear; mip filter NONE/POINT/LINEAR → not mipmapped/nearest/linear. D3DSAMP_MAXMIPLEVEL becomes lodMinClamp (D3D's "most detailed level allowed"). MAXANISOTROPY is clamped to 1-16 and forced to 1 unless min or mag filter is anisotropic.

Quality override. Unless HALO_NO_ANISO is set (any value other than empty or starting with 0), every sampler with a mip filter and a non-point min filter becomes 8x anisotropic with linear mips. The comment explains that Halo's hardware database only enabled anisotropy for a few 2003-era cards, so the engine asked for bilinear with point mips almost everywhere, and that 16x at 2560x1920 contributed to an earlier GPU-bound build on the headset.

Pressure rule: 256 entries keyed by (address U/V, min, mag, mip, anisotropy, max mip level), LRU replacement when full. test_metalrenderer_state_pressure.m issues 768 distinct requests and checks every draw succeeds with the cache bounded, then verifies all 25 U/V addressing combinations with 450 exact GPU pixel checks on both draw paths. test_rhw_diffuse_alpha.m checks that BORDER with a zero border colour adds nothing outside [0,1] while WRAP/CLAMP sample the edge texel, and that a non-zero border colour fails explicitly.

The legacy mr_draw_state.linear_filter/address_clamp booleans still select one of four prebuilt samplers when no full sampler description is supplied.

Textures

The content-keyed texture cache (mr_texture_find_cached, mr_texture_create_cached[_nomip], cube and volume variants), its LRU eviction, the per-draw binding pin, and CPU mip generation are described in Textures and Texture Packs. Uncached textures (mr_texture_create/update/destroy) are used only by tests and the RHW path's legacy texture ids; mr_texture_update waits for the GPU before overwriting.

Render targets, readback and publication

Function Behaviour
mr_target_texture Encodes the context's pending clears, ends its encoder if open, and returns the table id aliasing its colour target, so another context can sample it in submission order.
mr_read_framebuffer Commits, waits for completion, and copies the BGRA target to CPU memory. This is a synchronous stall; the bridge counts every call.
mr_write_framebuffer Stages CPU pixels in the arena (rows padded to 256 bytes) and blits them into the target on the shared command buffer, ordered after earlier readers and before later draws; never writes a texture from the CPU or drains the GPU.
mr_blit_target_to Blits the colour target into an external MTLTexture of the same size/format (zero-copy hand-off to the presenter).
mr_fxaa_target_to Same, through an FXAA 3.11 quality pass (full-screen triangle, linear clamp sampler, subpix parameter). Falls back to the plain blit if the destination is not renderable or the pipeline failed to build. The comment explains the need: each engine pixel spreads over about 1.3 display pixels across and 3.5-4 down on the headset, so near-horizontal edges become tall crawling staircases. The pass runs on gamma-encoded bytes before the presenter's sRGB view decodes them; the HUD layer is never filtered.
mr_blit_texture_copy Texture-to-texture copy on the shared command buffer.
mr_commit_async Encodes pending clears, ends the encoder, retires the arena, commits, and calls done(arg, ok) exactly once on completion (immediately if nothing was pending). On failure mr_last_commit_error() holds code N: text, which distinguishes a GPU fault from the compositor discarding work when the app is backgrounded.
mr_shared_device The shared MTLDevice, so the presenter allocates textures on the same device.

How the panorama passes use these is covered in Panorama System and Immersive Presenter.

Environment variables

Variable Default Effect Read at
HALO_DRAW_FASTPATH off (1 enables; visionOS app sets 1) Folded clears, encoder state reuse, resident vertex copies. metalrenderer.m:197
HALO_NO_ANISO unset Disables the forced 8x anisotropic trilinear override. metalrenderer.m:2602
HALO_TEX_FALLBACK unset Diagnostic: bind white for missing fixed-stage textures instead of failing the draw. metalrenderer.m:2603
HALO_PIPELINE_WORKERS, HALO_SHADER_PREWARM, HALO_PIPELINE_CACHE, HALO_PIPELINE_CACHE_DIR, HALO_PIPELINE_ARCHIVE, HALO_PIPELINE_ARCHIVE_SAVE_MS, HALO_PIPELINE_TRACE see Shader Translation Pipeline compiler service.
HALO_RADIAL_FOG off Radial-fog pipeline variants (read through halo_settings_radial_fog). halo_settings.c

Tests

Test Run by tools/run_source_checks.py What it asserts
test_metalrenderer_fastpaths.m yes (macOS) One seeded sequence of programmable, pretransformed and clip-space draws, clears at every point, depth/stencil tests, blending, cross-target sampling, uploads, readbacks, blits, FXAA and commits runs with and without fast paths; every readback and outside texture must match byte for byte. Targeted checks cover cached target ids after clears, destination aliasing, encoder restarts and releasing resident buffers before commit. Uses hand-written MSL in place of MojoShader.
test_metalrenderer_state_pressure.m yes (macOS) Sampler and depth/stencil cache bounds under pressure; independent U/V addressing for all five modes on both paths.
test_metalrenderer_texture_bindings.m yes (macOS) With a 1 KiB budget, a texture resolved inside a binding scope cannot be evicted (creation fails instead), and eviction resumes after the scope ends.
test_metalrenderer_overlay_cache.m yes (macOS) Basic, depth and programmable caches isolate world and HUD alpha semantics in either creation order (fake Metal objects).
test_metalrenderer_pipeline_cache.m yes (macOS) See Shader Translation.
test_metalrenderer_mips.c no Mip-chain byte accounting (NPOT 3x5 = 72 bytes; 256x256 = 349524 bytes), mip filtering reduces checkerboard minification variance, an anisotropic tuple renders.
test_d3d9_resource_upload.m no Overwriting/freeing source vertex and index bytes after mr_draw_rhw but before commit does not change the result (they were copied to the arena).
test_d3d9_texture_upload.m no Cached uploads copy CPU bytes; a queued draw keeps a texture removed from the cache, and a destroyed render-target context's aliased texture alive.
test_framebuffer_upload.m no mr_write_framebuffer ordering against draws/clears, odd row pitch, staging lifetime.
test_fixed_modulate_ops.m no Ops 18/20/21 GPU pixels within 1 of the D3D equations.
test_fixed_stage_dependencies.m no Unused texture operands add no texture or TEXCOORD (even with generated-coordinate flags on a stage that does not sample); a consumed missing coordinate still fails.
test_rhw_diffuse_alpha.m no texture_color_only keeps diffuse alpha; legacy alpha and additive blending; border/wrap/clamp pixels; non-zero border fails.
test_destalpha_blend.m no Exact DESTALPHA/ONE equations and RGB-only alpha preservation. It asserts MR_BLEND_COUNT == 9 and calls configure_program_color with four arguments, which no longer matches the current header (MR_BLEND_COUNT is 10 after MR_BLEND_CUSTOM) and seven-argument signature, so as written it appears to predate the custom-blend change.

Related pages

Clone this wiki locally