Skip to content

Audio System

M T edited this page Oct 4, 2026 · 1 revision

Audio System

Halo PC 1.10 produces all of its sound through DirectSound 8 (DSOUND.dll) and decodes its Ogg Vorbis streams through VORBISFILE.dll. Neither exists on macOS or visionOS, so the host replaces both: directsound.c emulates the DirectSound COM objects the translated engine calls, directsound_mixer.c mixes every secondary buffer into one stereo float bus, and an AudioToolbox AudioQueue plays that bus at 48 kHz. On visionOS an AVAudioSession coordinator (audio_session.inc) owns session activation, interruptions, backgrounding and recovery. The mixer is also the input to the haptics synthesiser (see Haptics).

The engine-side sound system (sources, channels, voices, the sound cache) is the original translated code; the host only implements the API underneath it. Several host behaviours exist specifically to match what that original code expects from Windows (see GetStatus and cursor semantics).

Source files

File Role
native/EngineHost/directsound.h Public host API (create, shutdown, watchdog, pause/resume/rebuild, statistics) and the HOST_DSOUND_SHIM_ENTRIES table that names every COM method as an import.
native/EngineHost/directsound.c DirectSound 8 COM emulation (device, buffers, 3D listener, 3D buffers), the AudioQueue output path, watchdog, WAV capture and voice trace.
native/EngineHost/directsound_mixer.h DsMixer, DsMixerVoice, DsMixer3D, DsMixerListener and the mixer API.
native/EngineHost/directsound_mixer.c Format conversion, Catmull-Rom resampling, millibel gain, pan, 3D distance rolloff, cones, head-yaw panning, own-weapon lift, soft limiter, per-voice onset detection.
native/EngineHost/directsound_engine_state.inc Read-only inspection of the original engine's own sound tables in guest memory, for the device report and the "voices held" log line.
native/EngineHost/audio_ownership_trace.inc Default-off trace of four original sound-ownership functions (HALO_AUDIO_OWNERSHIP_TRACE=1).
native/EngineHost/audio_source_identity.h XXH3-64 hash of each compressed stream handed to the Vorbis decoder (diagnostic identity, not a tag).
native/EngineHost/vorbis_shim.c / .h Native ov_open_callbacks, ov_read, ov_crosslap, ov_clear on top of stb_vorbis.
native/EngineHost/third_party/stb_vorbis.c Public-domain/MIT Ogg Vorbis decoder, v1.22, compiled into vorbis_shim.c with STB_VORBIS_NO_STDIO.
native/EngineHost/shims_misc.c Registers DirectSoundCreate8/DirectSoundCreate, enumeration stubs, all COM methods and the four VORBISFILE.dll exports in the import table.
native/EngineVision/Sources/audio_session.inc visionOS AVAudioSession coordinator: activation, interruptions, media-services reset, foreground/background, route logging, the external watchdog timer.
native/EngineHost/halo_settings.c Live settings read by the mixer: own-weapon gain (self_gain_db) and head yaw.

Architecture at a glance

flowchart LR
    ENG["Translated Halo sound engine"]
    VS["vorbis_shim.c (stb_vorbis)"]
    GD["Buffer guest_data (32-bit guest memory)"]
    DS["directsound.c COM shims"]
    HV["DsMixerVoice.data (host heap copy)"]
    MX["DsMixer: 768 voice slots"]
    CB["AudioQueue callback: 48 kHz float32 stereo, 3 x 1536 frames"]
    OUT["System audio output"]
    HAP["haptics.c detectors"]
    WD["host_dsound_watchdog"]
    SES["audio_session.inc (visionOS)"]

    ENG -->|"ov_open_callbacks / ov_read"| VS
    VS -->|"16-bit PCM"| ENG
    ENG -->|"Lock, write PCM"| GD
    ENG -->|"Unlock"| DS
    GD -->|"memcpy at Unlock"| HV
    ENG -->|"Play, Stop, SetVolume, 3D params"| DS
    DS --> MX
    HV --> MX
    MX -->|"ds_mixer_render"| CB
    CB --> OUT
    MX -->|"voice onsets + mixed signal"| HAP
    WD -->|"dispose + rebuild queue"| CB
    SES -->|"pause / resume / rebuild"| CB
    SES -->|"250 ms timer"| WD
Loading

Key points:

  • The engine writes PCM into guest memory (the buffer returned by Lock). At Unlock the host copies those bytes into a host-owned array that the mixer reads. The audio thread therefore never touches guest memory (directsound.c:811-829).
  • There is exactly one output: an AudioQueue (on macOS and on visionOS, per the file header at directsound.c:1-5). AVAudioSession is used only for session management on visionOS; it does not carry samples.
  • No spatial audio framework is used. 3D is reduced to a stereo gain pair per voice by the mixer.

How the COM objects are exposed to the guest

Halo calls COM methods through vtables in guest memory. build_vtables (directsound.c:211-225) allocates four guest vtables (IDirectSound8 with 12 methods, IDirectSoundBuffer8 with 24, IDirectSound3DListener8 with 18, IDirectSound3DBuffer8 with 21) and fills each slot with host_proc_address("DSOUND.dll", "Iface::Method"). Those names are registered as ordinary imports by HOST_DSOUND_SHIM_ENTRIES (directsound.h:89-118), which the header explains is so that "translated indirect calls use the same checked host boundary as ordinary DLL imports". See Win32 Compatibility Layer for the import/shim mechanism.

Each guest COM object is a 16-byte guest allocation: +0 vtable pointer, +4 index into the host's objects[] table. object_from_guest (directsound.c:481-487) validates the index, the alive flag and that the guest pointer matches either the buffer's main object or its 3D alias, so stale or forged pointers return NULL and the method returns DSERR_INVALIDPARAM.

Other DSOUND.dll exports (shims_misc.c:392-397):

Export Behaviour
DirectSoundCreate8, DirectSoundCreate (and ordinals #11, #1) host_dsound_create8: rejects a non-null device GUID with DSERR_NODRIVER, builds vtables, starts audio (audio_start), returns the single device object.
DirectSoundEnumerateA (#2) Calls the guest callback once with "Primary Sound Driver".
DirectSoundCaptureEnumerateA (#9) Reports no capture devices.
GetDeviceID (#14) Returns DSERR_NODRIVER.

Data structures

DsObject (host side of one COM object)

Defined at directsound.c:42-50; table objects[DS_OBJECT_MAX] with DS_OBJECT_MAX = 1024 live objects (slot 0 is never used).

Field Meaning
guest, guest_3d Guest COM object, and the lazily created 3D alias (listener for the primary buffer, IDirectSound3DBuffer8 for 3D secondaries). Both share the same slot index and reference count.
guest_data Guest buffer the engine locks and writes (same size as the voice).
refs, caps, primary, alive COM reference count, DSBCAPS_* flags, primary-buffer flag.
lock1, lock1_bytes, lock2_bytes, locked The outstanding Lock region (wrapping split into two parts).
unlock_*, last_*_frame Trace-only counters (refills, bytes written, non-zero samples, output frame of last write and last non-silent write).
voice The embedded DsMixerVoice the mixer reads.

DsMixerVoice and DsMixer

Defined in directsound_mixer.h:8-69.

Structure / field Meaning
DS_MIXER_MAX_VOICES = 768 Mixer slots; every secondary buffer occupies one slot for its lifetime.
DsMixerFormat tag (1 = PCM, 3 = IEEE float), channels, bits, block_align, sample_rate, avg_bytes_per_sec.
data, bytes Host copy of the buffer contents.
cursor_frames (double) Fractional play position in source frames.
end_count Natural completions of non-looping plays (kept across Play).
volume, pan, frequency DirectSound millibels (-10000..0), pan (-10000..10000), playback rate in Hz.
playing, looping, has_3d, pending_3d State flags.
onset_* Per-voice onset detector state for haptics (see Haptics).
self_voice Latched when the voice has sounded within 0.5 units of the listener since its last Play (the player's own weapon).
current_3d, deferred_3d (DsMixer3D) Position, velocity, cone orientation, min/max distance, cone angles, outside volume, mode.
DsMixer.lock One mutex protecting all voices and the listener.
DsMixer.current_listener, deferred_listener, pending_listener Listener position, velocity, front/up, distance/rolloff/doppler factors.
rendered_frames, ended_voices, limited_samples, rendered_samples Counters for diagnostics.

Object lifecycle

Creating buffers

IDirectSound8::CreateSoundBuffer (host_dsound_device_3, directsound.c:706-723):

  1. Validates the DSBUFFERDESC (size ≥ 20). A primary buffer must pass zero bytes and no format.
  2. For a secondary buffer, rejects 0 bytes, more than 128 MiB, an unreadable format, or a size that is not a multiple of block_align (DSERR_BADFORMAT).
  3. DSBCAPS_LOCHARDWARE is not refused. The comment explains that Vista-and-later Windows mixes every buffer in software and never fails this; refusing it "made Halo believe the device had no hardware voices at startup".
  4. object_new (directsound.c:518-549) finds the first free slot (the limit is concurrent buffers, not buffers ever created), callocs host data, adds the voice to the mixer, allocates guest_data and the 16-byte guest object. Any failure unwinds every allocation and the mixer slot.

read_format (directsound.c:492-506) accepts mono or stereo; 8- or 16-bit PCM or 32-bit float; 100..200000 Hz; WAVE_FORMAT_EXTENSIBLE (0xFFFE) whose subformat is PCM or float. block_align and avg_bytes_per_sec must be consistent.

DuplicateSoundBuffer creates a new secondary with the same caps/format and copies the host data.

QueryInterface and the 3D aliases

buffer_qi (directsound.c:586-599) matches only the first 32 bits (Data1) of the IID:

IID Data1 Interface Result
0x00000000 IUnknown The buffer itself
0x279AFA85 IDirectSoundBuffer The buffer itself
0x6825A449 IDirectSoundBuffer8 The buffer itself
0x279AFA84 IDirectSound3DListener Primary buffer only: a guest alias with the listener vtable
0x279AFA86 IDirectSound3DBuffer 3D secondary only: a guest alias with the 3D-buffer vtable

Every successful QueryInterface adds a reference to the shared DsObject.

Release

object_release (directsound.c:565-583) frees everything when the shared count reaches zero: it removes the voice from the mixer under mixer.lock before freeing voice.data, and keeps holding objects_lock until the slot is zeroed. The comment states why: "Keep the slot owned until the callback can no longer see its voice; otherwise a creator could overwrite it during removal."

The device object (device_guest) is created once and never freed; its AddRef/Release only move a counter.

Buffer methods

Method (index) Implementation notes
GetCaps (3) Size, caps, buffer bytes.
GetCurrentPosition (4) Play cursor and write cursor; see below.
GetFormat (5) Writes an 18-byte WAVEFORMATEX.
GetVolume/GetPan/GetFrequency (6-8) Read under mixer.lock.
GetStatus (9) PLAYING, plus LOOPING only while playing. Counts probes and busy answers.
Initialize (10) DSERR_ALREADYINITIALIZED.
Lock (11) Region from an offset, from the write cursor (DSLOCK_FROMWRITECURSOR) or the whole buffer (DSLOCK_ENTIREBUFFER); splits at the end of the ring into two parts; a second Lock before Unlock is DSERR_INVALIDCALL.
Play (12) Only flag allowed is DSBPLAY_LOOPING. Sets playing, looping and calls ds_mixer_voice_restart (resets onset state and self_voice). Does not move the cursor.
SetCurrentPosition (13) Sets cursor_frames = pos / block_align; out-of-range positions rejected.
SetFormat (14) Primary buffer only.
SetVolume (15) Requires DSBCAPS_CTRLVOLUME; -10000..0 mB.
SetPan (16) Requires DSBCAPS_CTRLPAN; -10000..10000.
SetFrequency (17) Requires DSBCAPS_CTRLFREQUENCY; 0 restores the format rate; otherwise 100..200000.
Stop (18) Clears playing; keeps the cursor (Stop then Play resumes).
Unlock (19) Must exactly match the outstanding Lock; copies both regions into host data under mixer.lock.
Restore (20) DS_OK for a valid object.
SetFX, AcquireResources, GetObjectInPath (21-23) DSERR_UNSUPPORTED.

IDirectSound8::GetCaps reports all primary formats (dwFlags = 0x1F), a continuous rate range 100..200000 Hz and one primary buffer, with all hardware-voice counts zero (directsound.c:724-730). GetSpeakerConfig reports 4 (stereo).

GetStatus and cursor semantics (why they matter)

Halo streams every sound through a fixed pool of looping DirectSound buffers ("voices"). The original allocator 005482E0 only reuses a stopped voice when 00547FF0 sees neither PLAYING nor LOOPING from GetStatus. voice_status (directsound.c:168-180) therefore reports LOOPING only together with PLAYING, as Windows and Wine do. The comment records the failure this prevents: reporting LOOPING for a stopped buffer "silenced Build75": each of the 24 positional voices played one sound and was then held until the next sound_stop_all, about 35 seconds of effects after each level load, pause or checkpoint revert.

Similarly, voice_write_cursor (directsound.c:181-195) returns a write cursor that leads the play cursor only while playing; a stopped buffer reports both cursors at the same position. Halo's 00547C80 compares the pair to decide whether a draining voice is still running; unequal cursors on a stopped voice made it skip the Play.

The playing write lead is sample_rate / 20 frames (50 ms), computed in whole frames (directsound_mixer.c:190-199); computing it in bytes had produced half-frame cursors for 22,050 Hz streams.

Typical sound start, as the engine drives it

sequenceDiagram
    participant E as Halo (00547C80 / 005482E0)
    participant D as directsound.c
    participant M as DsMixer
    participant Q as AudioQueue callback
    E->>D: GetStatus(voice) (allocator probe)
    D-->>E: 0 (stopped, free)
    E->>D: Lock(0, ring bytes)
    D-->>E: pointer(s) into guest_data
    E->>E: write PCM into guest memory
    E->>D: Unlock(regions)
    D->>M: memcpy guest_data to voice.data (mixer.lock)
    E->>D: SetCurrentPosition(0)
    E->>D: Play(DSBPLAY_LOOPING)
    D->>M: playing=1, looping=1, restart onset state
    loop every buffer period
        Q->>M: ds_mixer_render(frames, 48000)
    end
    E->>D: Lock(FROMWRITECURSOR) / Unlock (refill ring)
    E->>D: Stop()
    D->>M: playing=0 (cursor kept)
Loading

3D listener and 3D buffers

All 3D setters exist for both immediate and deferred application (directsound.c:835-882):

  • Every setter writes the deferred copy (deferred_listener or voice.deferred_3d). With DS3D_DEFERRED (non-zero last argument) it only marks pending_*; with DS3D_IMMEDIATE it copies the whole deferred state to current.
  • IDirectSound3DListener8::CommitDeferredSettings calls ds_mixer_commit (directsound_mixer.c:53-67), which applies the pending listener and every voice's pending 3D parameters under one lock.
  • Float arguments are read from the guest stack at ESP+8 (after the return address and this), and vectors as three consecutive floats. All values are validated: vectors must be finite; distance factor > 0; rolloff and doppler ≥ 0; cone angles ≤ 360 with inside ≤ outside; outside volume -10000..0; min distance > 0 and ≤ max; mode 0..2.
  • Voice defaults (directsound.c:511-517): min distance 1, max distance 1e9, cone angles 360/360.
  • Listener defaults (directsound_mixer.c:25-34): front +Z, up +Y, all factors 1.

Velocities and the Doppler factor are stored and returned by the getters, but the mixer does not apply Doppler shift. With HALO_AUDIO_VOICE_TRACE on, listener_changed logs the first eight listener writes and every 500th ([audio-listener]), because, as the comment says, whether DirectSound should also attenuate depends on the rolloff factor the game sets.

The mixer

ds_mixer_render (directsound_mixer.c:230-341) runs on the AudioQueue callback thread with mixer.lock held for the whole pass.

flowchart TD
    S["For each of 768 slots"] --> P{"playing and has data?"}
    P -- no --> S
    P -- yes --> G["ds_mixer_voice_gains: volume, pan, 3D"]
    G --> Z{"both gains zero?"}
    Z -- yes --> ADV["Advance cursor by step x frames, wrap or end; no resampling"]
    Z -- no --> R["Per output frame: Catmull-Rom from 4 source frames"]
    R --> ACC["Accumulate left*gain_l, right*gain_r"]
    ACC --> ON["Positional voices: onset envelopes (haptics)"]
    ON --> R
    ADV --> S
    R --> S
    S --> LIM["After all voices: count samples past 1.0, apply soft knee"]
    LIM --> HO["Unlock, then halo_haptics_observe(mixed output)"]
Loading

Format conversion

read_channel (directsound_mixer.c:69-84) converts a sample to float: 8-bit unsigned (x-128)/128, 16-bit signed x/32768, 32-bit float clamped to ±1 (non-finite becomes 0). Mono voices feed the same value to both channels.

Resampling

Every voice is resampled to the 48 kHz output with step frequency / output_rate. The interpolator is four-tap Catmull-Rom (directsound_mixer.c:201-211). The comment explains: Halo's effects are 22,050 Hz, linear interpolation at that ratio acts as a low-pass "which is what makes the mix sound muffled". A looping voice wraps its neighbour taps across the loop seam; a one-shot holds its end frames. When a non-looping voice runs past its last frame it stops, end_count and ended_voices increment.

A voice whose computed gain is zero in both channels (for example one Halo has muted to -10000 mB but keeps streaming into) is advanced, wrapped or ended exactly as if mixed, without the resampling cost (directsound_mixer.c:253-263).

Gain, pan and 3D

ds_mixer_voice_gains (directsound_mixer.c:86-165):

  1. Volume: millibel_gain(v) = 10^(v/2000), with ≤ -10000 mapped to 0 and ≥ 0 to 1.
  2. Pan: positive pan attenuates the left channel by millibel_gain(-pan), negative pan attenuates the right.
  3. 3D (if has_3d and mode ≠ 2 DS3DMODE_DISABLE):
    • Relative position: mode 0 (normal) subtracts the listener position; mode 1 (head-relative) uses the position as is.
    • Distance rolloff: beyond min_distance, gain *= (min / min(distance, max)) ^ rolloff. The listener's distance factor is deliberately not applied. The comment: Halo sets it to 3.048 (ten feet in metres); multiplying the distance by it "quietly pushed every positional sound 9.7 dB down while 2D music was untouched".
    • Own-weapon lift: if the distance is < 0.5 units or self_voice is latched, gain *= halo_settings_self_gain(). Halo hands the player's own weapon, reload and melee to DirectSound 20 dB down (the assault rifle at -2000 mB); in a headset recording firing raised the mix by under a decibel.
    • Spatial pan: the direction is projected on the listener's right vector (up x front). Before that, the right/front pair is rotated by the viewer's head yaw (halo_settings_head_yaw()), so a sound keeps its place in the world when only the head turns in the panorama. p = dot(direction, right) in [-1, 1]; p > 0 multiplies the left gain by sqrt(1-p), p < 0 multiplies the right gain by sqrt(1+p).
    • Cones: if a cone orientation is set and outside_angle < 360, the full angle to the listener (twice the half-angle) is compared with the inside/outside angles and the outside volume is interpolated linearly in millibels between them.

self_voice is latched during rendering when a positional voice is within 0.5 units of the listener (directsound_mixer.c:250-251) and cleared by Play. The header comment explains why it is latched rather than judged per block: a moving player left the shot more than half a unit behind mid-sound, the 18 dB lift fell away, and the shot was heard "cutting out".

Head pitch is not used by the mixer; only yaw.

Limiter

After all voices are summed, every sample passes through soften (directsound_mixer.c:213-228): unity below a 0.90 knee, then knee + (1-knee) * over/(1+over) with over = (|x|-knee)/(1-knee), so the output never reaches ±1. A full-scale peak loses about 0.45 dB; a mix at twice full scale still lands under the ceiling. Samples that would have exceeded ±1 before the curve are counted in limited_samples; host_dsound_get_stats logs DirectSound limiter: N of M samples past full scale at most every 10 seconds of output (directsound.c:641-650).

Own-weapon gain setting

Setting Source Default Range Effect
HaloSettings.self_gain_db HALO_SELF_GAIN_DB seeds it (halo_settings.c:64); live slider "Your own weapon" in the Presentation settings window (EngineSettingsView.swift:108-111) 18 dB 0..24 dB (out-of-range environment values fall back to 18) Linear gain 10^(dB/20) applied to positional voices on the listener (halo_settings_self_gain, halo_settings.c:136). The haptics onset detector divides it back out so the player's weapon is judged at Halo's own level.

Output path (AudioQueue)

Constants

Constant Value Where Meaning
OUTPUT_RATE 48000 Hz directsound.c:35 Output sample rate; also the mixer's render rate.
OUTPUT_BUFFERS 3 same AudioQueue buffers in circulation.
output_frames_per_buffer 1536 directsound.c:36-40 Frames per buffer (32 ms). Total queued latency 96 ms.
OUTPUT_FRAMES 2048 same Upper bound for HALO_AUDIO_FRAMES and per-callback render clamp.
Format linear PCM, float32, packed, 2 channels, 8 bytes/frame directsound.c:317-324 Interleaved stereo float.

Why 1536: the comment at directsound.c:36-39 says Build80 recorded callback gaps of up to 50 ms; with 512-frame buffers the two remaining buffers covered only 21.3 ms, so a healthy queue could run dry without an enqueue error. 1536 frames keeps 64 ms of reserve after each completed buffer. test_audio_watchdog.c asserts (OUTPUT_BUFFERS-1) * frames * 1000 / OUTPUT_RATE > 50.

Queue lifecycle state

Variable Meaning
output_requested Set by the first DirectSoundCreate8 (audio_start, directsound.c:386-394). Until then no queue exists.
output_suspended Set by host_dsound_pause_output, cleared by host_dsound_resume_output. A suspended output never starts.
output_restart_required Forces an explicit AudioQueueStart on the next resume. Set at creation, on pause, on dispose and on every resume request.
output_callbacks_suspended, output_callbacks_discarding Callback gate (see below).
output_deferred_buffers[3] Buffers that completed while paused, held for re-enqueue on resume.
output_enqueue_fault Last failing AudioQueueEnqueueBuffer status from the callback, retained until a rebuild.
output_terminated Set by host_dsound_shutdown; nothing restarts afterwards.
output_generation Incremented on every successful AudioQueueNewOutput.

audio_prepare_locked (directsound.c:302-343) reads HALO_AUDIO_FRAMES and the capture variables once, creates the queue with AudioQueueNewOutput(..., NULL run loop, ...) (callbacks on AudioQueue's own thread), allocates the three buffers, primes each with silence and enqueues it. Any failure disposes the queue and returns 0.

audio_resume_locked (directsound.c:345-384) prepares if needed and, when a restart is required or kAudioQueueProperty_IsRunning reads 0, re-enqueues deferred buffers as silence, opens the callback gate and calls AudioQueueStart. It does not trust IsRunning: on macOS it stays true for a paused queue, and the Build60 headset report showed it true after backgrounding while callbacks had stopped for good (directsound.c:408-420).

The callback

audio_callback (directsound.c:227-275), under output_callback_lock:

  1. If terminated or discarding, return (the buffer is dropped).
  2. If suspended, remember the buffer in output_deferred_buffers (once) and return without rendering, so no voice advances while paused.
  3. Record the gap since the previous callback (max_callback_gap_ns; a gap over two buffer periods counts as a late callback).
  4. ds_mixer_render into the buffer, then measure peak and non-zero samples into atomics (output_frames, output_nonzero, output_peak_bits).
  5. If WAV capture is active, convert to 16-bit and append (bounded).
  6. Enqueue. A failure is recorded in telemetry_enqueue_failures and output_enqueue_fault.
  7. Record the callback's own work time (max_callback_work_ns).

The callback never writes host.log: "stderr shares its file lock with engine/diagnostic writers and may block behind disk I/O" (directsound.c:258-260). All logging happens on the diagnostics/control side.

Pause, resume, rebuild, shutdown

Function Lines Behaviour
host_dsound_pause_output 396-406 Marks suspended and restart-required, closes the callback gate (without discarding), AudioQueuePause. Voices, cursors, diagnostics and capture are preserved.
host_dsound_resume_output 408-420 Clears suspended, forces a restart, audio_resume_locked. Records last_recovery 0 or -1.
host_dsound_rebuild_output 469-479 Disposes the queue and prepares a new one (only if requested and not suspended); does not start it. Callers follow it with resume.
host_dsound_shutdown 601-608 Sets output_terminated, disposes, finalises the capture WAV header. Called by host.c after the engine has stopped (host.c:659).

audio_dispose_locked (directsound.c:288-300) closes the gate in discard mode first, because AudioQueueDispose can synchronously return buffers through the callback, and "those callbacks must not mix again" (directsound.c:277-279).

Lock order

  • output_lock (queue control) is taken before output_callback_lock (callback gate); output_callback_lock is released before AudioQueuePause/Dispose can call back.
  • objects_lock is taken before mixer.lock (host_dsound_get_stats, host_dsound_get_voice_stats).
  • The callback holds output_callback_lock and then mixer.lock (inside ds_mixer_render). COM methods take only mixer.lock (or objects_lock alone).

Watchdog

host_dsound_watchdog (directsound.c:426-466) rebuilds the queue when output is requested, not suspended, not terminated, and either:

  • a callback enqueue failed (output_enqueue_fault), or
  • output_frames has not advanced for 0.5 s and the queue has been started for at least 0.5 s (resume grace).

Rebuilds are rate-limited to one every 2 seconds. Progress is rechecked after taking output_lock, because a callback may have finished while the watchdog waited, and the fault is rechecked because another transaction may already have replaced the queue. A rebuild is audio_dispose_locked + audio_resume_locked inside one locked transaction; a failed preparation leaves no queue but stays eligible for the next retry. Each rebuild logs [audio-watchdog] ... rebuilt the queue (N so far): playing|FAILED and increments watchdog_rebuilds (reported as audio_rebuilds).

Who runs it:

  • By default, the engine thread once per IDirect3DDevice9::Present (d3d9.c:544-546).
  • On visionOS, audio_session_prepare calls host_dsound_watchdog_set_external() and runs it from a 250 ms timer on the audio control queue instead. The comment explains that the per-frame watchdog "slept through exactly the stalls it exists for: level loads and the engine's own stops (Build73's queue stopped at 290.4 s and nothing recovered it)" (audio_session.inc:62-75).

visionOS audio session coordinator

audio_session.inc is included into EngineVisionRuntime.m (EngineVisionRuntime.m:606) and is compiled only when TARGET_OS_VISION. enginevision_start calls audio_session_prepare() before creating the engine worker (EngineVisionRuntime.m:726-728).

All session mutations and queue recovery run on the serial queue HaloVision.audio-session, never on the render callback.

Element Lines Behaviour
audio_session_can_resume 13-17 True only if the runtime is STARTING or RUNNING, the app is foreground, not interrupted and resume is allowed.
audio_session_resume 29-52 setCategory:Playback mode:Default options:0, setActive:YES, optional host_dsound_rebuild_output, then host_dsound_resume_output. On any failure: pause the output and schedule a retry 2 s later (audio_retry_after_ns).
audio_session_tick 54-60 Every 250 ms: if a resume is pending and its retry time has come, retry the whole transaction; otherwise run host_dsound_watchdog.
audio_session_route 20-28 Logs [audio-route] with output port types only (never names or UIDs), volume, rate, channel count; at most 64 reports.

Notifications (audio_session.inc:76-137):

Notification Action
AVAudioSessionInterruptionNotification began audio_interrupted = YES, pause the output.
... ended Clear interrupted; allow resume only if ShouldResume was set; attempt resume.
AVAudioSessionMediaServicesWereResetNotification Mark rebuild needed (orphaned queue) and resume.
UIApplicationDidBecomeActiveNotification Clear interruption and resume_allowed=NO (coming back is the player's own action; keeping shouldResume=NO "left the mix silent for the rest of the session"), set foreground, resume.
UIApplicationDidEnterBackgroundNotification Clear foreground, pause. The app has no audio background mode; a queue left running stalls and every background rebuild failed in Build72 (OSStatus 'what').
AVAudioSessionRouteChangeNotification Log only.
stateDiagram-v2
    [*] --> Startup
    Startup --> Playing: activation + resume ok
    Startup --> RetryPending: activation/rebuild/resume failed
    RetryPending --> Playing: retry after 2 s succeeds
    RetryPending --> RetryPending: retry fails
    Playing --> Interrupted: interruption began (pause)
    Interrupted --> Playing: ended with ShouldResume
    Interrupted --> Waiting: ended without ShouldResume
    Waiting --> Playing: app became active
    Playing --> Background: entered background (pause)
    Background --> Playing: became active
    Playing --> Playing: watchdog rebuild (stall or enqueue fault)
    Playing --> Playing: media services reset (rebuild + resume)
Loading

On macOS (desktop host and probes) there is no session coordinator; the queue is started on device creation and the watchdog runs from Present.

Ogg Vorbis decoding (VORBISFILE.dll replacement)

Halo imports four libvorbisfile functions; vorbis_shim.c implements exactly those, registered in shims_misc.c:412-415 and shims_misc.c:444. All are cdecl.

Constant Value
VORBIS_MAX_STREAMS 16 simultaneous streams
VORBIS_READ_CHUNK 64 KiB guest read chunk
VORBIS_MAX_INPUT 256 MiB maximum compressed stream
Error codes OV_EREAD -128, OV_EFAULT -129, OV_EIMPL -130, OV_EINVAL -131, OV_ENOTVORBIS -132

ov_open_callbacks(source, vf, initial, ibytes, read, seek, close, tell) (vorbis_shim.c:73-92):

  1. read_source (vorbis_shim.c:38-71) pulls the entire compressed stream into host memory through the guest's own read callback (host_call_guest), using a 64 KiB guest scratch buffer. If seek and tell callbacks exist and there is no initial prefix, it seeks to the end to size the buffer, then back. It keeps reading until a zero read (a short read is not end of input; an older loop exited early), doubles capacity as needed, rejects a callback that returns more than asked, and rejects streams over 256 MiB rather than truncating them.
  2. stb_vorbis_open_memory decodes headers; only 1- or 2-channel streams are accepted.
  3. A free slot in streams[16] is claimed for this vf (the guest's OggVorbis_File address is used purely as a key).

ov_read(vf, out, length, bigendian, word, sgned, bitstream) (vorbis_shim.c:94-122): only little-endian signed 16-bit is supported (otherwise OV_EINVAL). Calls stb_vorbis_get_samples_short_interleaved for length/2 shorts and returns bytes produced (0 at end). A request smaller than one frame (length < channels*2) returns 0 without consuming input or marking end of stream. *bitstream is always 0. The last up to 4096 output samples are kept in tail for crosslapping.

ov_crosslap(old, new) (vorbis_shim.c:124-126): requires equal channel counts; copies the old stream's tail into the new stream's overlap. The next ov_read on the new stream linearly cross-fades its first n samples from the overlap (old*(1-t) + new*t, t = (i+1)/(n+1)). This is a simple linear blend, not libvorbisfile's windowed lapping.

ov_clear(vf) (vorbis_shim.c:128-137): frees the decoder and compressed copy, then calls the guest's close callback and returns its result.

The decoded PCM is returned into guest memory; the engine then writes it into a DirectSound buffer like any other sound. With voice tracing on, opens/reads/clears are logged as [audio-voice] vorbis-* with the stream generation, decoder offsets and an XXH3-64 identity-xxh3 of the compressed input (only computed when tracing).

Diagnostics

Statistics APIs

Function Returns
host_dsound_get_stats Total output frames, non-zero samples, peak; also emits the one-time "mix became active" line, the periodic limiter line, and (when tracing) the 5-second voice summary. Must not be called from the callback.
host_dsound_get_output_diagnostics HostDsOutputDiagnostics (directsound.h:26-35): generation, lifecycle flags, last OSStatus of prepare/start/enqueue/pause/dispose/recovery, callbacks, late callbacks, enqueue failures, last callback time, max gap, max work, sample rate, buffer frames.
host_dsound_get_output_state kAudioQueueProperty_IsRunning and its OSStatus.
host_dsound_get_voice_stats HostDsVoiceStats (directsound.h:36-59): host Play/Stop/GetStatus counters and buffer counts, plus the engine's own tables.
host_dsound_watchdog_rebuilds Count of watchdog rebuilds.

On visionOS these feed the device report through EngineDiagnosticsBridge.m:67-105, which also samples AVAudioSession.outputVolume. See Diagnostics and Telemetry.

Engine sound tables (read-only)

directsound_engine_state.inc documents and reads the original sound engine's state in guest memory (for halo.exe build c9acf0c46954, per its header). From its comment (directsound_engine_state.inc:1-34):

Guest address Table
0x007252C0 Pointer to the sound-source data array (512 x 0xB0), filled by sound_impulse_start (00549AF0).
0x00724A50 Looping sounds (128 x 0xE4).
0x00724A60, count at 0x007252B4 Channels (0x18 each; +0 source or -1).
0x007252E4 Per-channel voice map (+0 voice or -1, +2 voice type).
0x00725430, count at 0x00725428 Voice records (0x678 each): +0 state, +2 channel, +8 stopped, +9 draining, +0x38 type, +0x670 IDirectSoundBuffer8.
0x0069F528 Voice type table (the first type is the positional voices).
0x006AC528 Sound cache data array (512 x 16: +2 loaded, +5 lock count).
0x006AC4A0 Cache read requests (512 x 0x30: +0x1D queued, +0x20 file: level map, bitmaps.map, sounds.map).
0x00725200, 0x00725201, 0x00725202, 0x007252B6 Initialised, enabled, paused by the game, disabled flags.

Every data array is checked for the "d@t@" signature, capacity, element size and bounds before being walked; reads begin only after the guest has created its DirectSound device (engine_sound_observable). sound_engine_voices_locked classifies each unassigned voice as free or held by replicating 00547FF0's test against the host's GetStatus answer. voice_exhaustion_note logs [audio-engine] voices held once when some voice type has held voices but no free one (Build75's failure signature) and voices free when it clears (at most 64 notes).

Voice trace (HALO_AUDIO_VOICE_TRACE)

When enabled, host_dsound_get_stats emits, at most once per 5 s of output, an [audio-voice] summary line (active/looping voices, ends, Play/Stop/probe counters, engine flags, sources, channels, starved channels, voices free/held by type, cache, queued reads), then one voice line per playing or refilled buffer. Allocation, play, stop, position, refill (on silent/non-silent transitions), failures and Vorbis events are logged inline. Detail lines are capped at 256 per 5-second window, with a dropped counter (directsound.c:145-162). The switch is "lazy-on": host_dsound_trace_enabled re-reads the environment until it finds it set, so an early UI poll cannot cache it off before the engine worker exports its defaults.

Ownership trace (HALO_AUDIO_OWNERSHIP_TRACE=1)

host_audio_ownership_dispatch (audio_ownership_trace.inc:61-87) is called from engine_dispatch_override (overrides.c:249) for the hook addresses 0x00442550, 0x00544090, 0x00544120 and 0x0048A1A0 (registered in engine_hooks.h). It logs a table-observed line when the tag table pointer (0x0087BC14), header (0x006A8954) or valid flag (0x006A8150) changes, and before/after snapshots around the first three functions (for 0x00442550 only when its first stack argument is 0x0066A044), including the current scenario's name, the tag's name and definition, and a salted record from the pool at 0x007461A0. It executes the wrapped function exactly once, never writes guest memory, rejects stale salts, and stops after 512 reports. The code does not document the semantic name of each wrapped function beyond "original sound ownership observations"; see Engine Overrides and Hooks.

WAV capture

HALO_AUDIO_CAPTURE_WAV=<path> writes the final mix as 16-bit stereo 48 kHz WAV from the callback (no allocation; a stack buffer), for HALO_AUDIO_CAPTURE_SECONDS (default 30, 1..1800; the comment notes thirty seconds "only reaches the opening cutscene"). The header is rewritten with the real length at shutdown. The capture can be fed to tests/haptics_from_capture.c (see Haptics).

Environment variables

Variable Default Effect Read at
HALO_AUDIO_FRAMES 1536 Frames per AudioQueue buffer, accepted 64..2048; read once at first queue preparation. directsound.c:306-307
HALO_AUDIO_CAPTURE_WAV unset Path of a diagnostic WAV capture of the output mix. directsound.c:309-315
HALO_AUDIO_CAPTURE_SECONDS 30 Capture length, 1..1800 s. directsound.c:314
HALO_AUDIO_VOICE_TRACE off; 1 on visionOS (device defaults, overridable) Enables voice/listener/Vorbis tracing and the 5-second summaries. Any value other than empty or 0 enables it. directsound.c:99-106, EngineVisionRuntime.m:634
HALO_AUDIO_OWNERSHIP_TRACE off Exactly 1 enables the ownership trace (max 512 reports). audio_ownership_trace.inc:63-64
HALO_SELF_GAIN_DB 18 Seeds the own-weapon lift, 0..24 dB. halo_settings.c:64
HALO_SIM_HEAD unset Simulator only: yaw,pitch in degrees replaces the head pose; the yaw reaches the mixer's spatial pan. EngineImmersive.swift:355-360
HALO_PROBE_SUSPEND_AUDIO unset Menu-input probe only: 1 pauses output before the engine starts (rendering-only measurements). MenuInputProbe.m:844-848
HALO_ENGINE_TRANSLATION newest native/build/engine-reuse/*/sub_005482E0.c Source-check runner: directory of translated functions for the translated-allocator variant of the voice-reuse test. run_source_checks.py:155

Threading summary

Thread Audio work
Engine thread (and other translated guest threads) All COM calls, Vorbis calls (guest callbacks run on the caller's CPU context), per-frame watchdog on macOS.
AudioQueue callback thread audio_callback, ds_mixer_render, haptics onset and observe detectors. No logging, no guest memory access.
visionOS HaloVision.audio-session serial queue Session activation, interruptions, background/foreground, 250 ms watchdog/retry tick.
Diagnostics/UI host_dsound_get_*; all logging that could block.

Failure modes and recovery

Symptom Mechanism
Callbacks stop while IsRunning stays true (Build60) Resume always restarts; watchdog rebuilds after 0.5 s without progress.
A buffer lost to a failed callback enqueue output_enqueue_fault is retained across later successful enqueues; watchdog rebuilds even though frames advance.
Queue stalls during level loads (Build73) External 250 ms watchdog on visionOS, independent of Present.
Session activation fails Output paused, whole transaction retried every 2 s while the runtime runs and the app is in the foreground.
Media services reset Queue rebuilt on the next resume.
Backgrounding Output paused cleanly; restarted on foreground.
Effects go silent after ~35 s (Build75) GetStatus reports LOOPING only while playing; held voices logged by voice_exhaustion_note.
Positional sounds ~10 dB too quiet Distance factor excluded from rolloff.
Own weapon inaudible under music self_gain_db lift, latched per sound.
Muffled sound Catmull-Rom interpolation instead of linear.
Buzz in loud firefights Soft-knee limiter instead of hard clipping.

The release notes state the limits of this evidence: "Audio interruptions have been reported. Synthetic mixer checks do not establish uninterrupted audible playback on a headset" (docs/KNOWN_ISSUES.md).

Tests

C tests in native/EngineHost/tests/ are built and run by tools/run_source_checks.py unless noted (see Testing and Source Checks).

Test Runner What it asserts
test_directsound_mixer.c Mac Exact resampled values (Catmull-Rom 0.1875 between ramp points, a full-scale frame softened to 0.95), volume/pan attenuation, distance factor does not change rolloff, 4x min distance gives 0.25, own-weapon lift persists after the player moves away, muted voices still advance/end/wrap, removed voices produce silence.
test_mixer_quality.c portable 1 kHz tone 22,050 to 48,000 Hz has < 0.2 % non-fundamental power; quiet material passes the limiter untouched; eight full-scale voices stay ≤ 1.0 with no pinned samples; head yaw ±90° moves a frontal sound to the opposite ear; rolloff unaffected by distance factor.
test_audio_watchdog.c Mac Mocked AudioQueue and clock: buffer reserve > 50 ms, failed rebuild retries with a 2 s bound, single locked transaction, resume grace, concurrent progress while waiting for the lock, deferred paused buffers re-enqueued exactly once, failed Start keeps the gate closed, enqueue loss detected despite later progress, termination stops recovery, telemetry (late callbacks, gaps exclude pauses, 48000/1536).
test_directsound_lifetime.c Mac 2048 create/play/3D-alias/release cycles with a concurrent render thread; stale pointers rejected; allocation failure at each stage leaks nothing; 768-voice mixer capacity and 1023-object table recover after release.
test_directsound_voice_reuse.c Mac (plus translated-allocator variant when generated code exists) Drives Halo's allocator (C transcription of 005482E0/00547FF0, or the translated functions with -DHALO_TRANSLATED_ALLOCATOR) through the shim: 200 sounds on 24 positional voices all start; GetStatus/cursor semantics; engine-table statistics; held/free log lines once each. -DBUILD75_GET_STATUS reproduces the Build75 silence (24 sounds, then none).
test_guest_heap_sound_owners.c Mac Production guest heap with 4096 recycled buffer/3D-alias lifetimes and a concurrent renderer; shared references keep guest data alive; final release reclaims every block.
test_audio_shim_boundaries.c Manual (needs an owned sounds.map) Vorbis input reader (> 512 KiB, short reads, prefix, seek/no-seek, invalid callback); DirectSound ABI (lazy trace enable, 5 s summary cadence, 256-event cap, mono/stereo cursor alignment, split refills, rejected calls, Stop/Play resume, natural end); Vorbis ABI (undersized reads, full drain equals fresh stream, close/reuse); pause/resume/terminated diagnostics without queue mutation.
test_halo_vorbis.c Manual (needs sounds.map) stb_vorbis opens the first Ogg stream in sounds.map, 1-2 channels, 11025-48000 Hz, decodes non-zero PCM.
test_audio_source_identity.c Manual Equal inputs hash equally; a final-byte or length change differs; input untouched.
test_audio_ownership_trace.c Manual Off by default with no side effects; when on, the wrapped function runs exactly once, CPU state is preserved, guest memory is unchanged, stale salt yields flags=ffffffff, report cap respected.
AudioSessionRecoveryValidation.m Mac Includes the production audio_session.inc with a mocked AVAudioSession and host functions: failed activation retried after 2 s (not sooner), media reset rebuilds, failed rebuild/resume retried, background and interruption (without ShouldResume) suppress resume, stopped runtime suppresses resume.

Related pages

Clone this wiki locally