-
Notifications
You must be signed in to change notification settings - Fork 5
Audio System
Halo PC 1.10 produces all of its sound through DirectSound 8 (DSOUND.dll) and decodes its Ogg Vorbis streams through VORBISFILE.dll. Neither exists on macOS or visionOS, so the host replaces both: directsound.c emulates the DirectSound COM objects the translated engine calls, directsound_mixer.c mixes every secondary buffer into one stereo float bus, and an AudioToolbox AudioQueue plays that bus at 48 kHz. On visionOS an AVAudioSession coordinator (audio_session.inc) owns session activation, interruptions, backgrounding and recovery. The mixer is also the input to the haptics synthesiser (see Haptics).
The engine-side sound system (sources, channels, voices, the sound cache) is the original translated code; the host only implements the API underneath it. Several host behaviours exist specifically to match what that original code expects from Windows (see GetStatus and cursor semantics).
| File | Role |
|---|---|
native/EngineHost/directsound.h |
Public host API (create, shutdown, watchdog, pause/resume/rebuild, statistics) and the HOST_DSOUND_SHIM_ENTRIES table that names every COM method as an import. |
native/EngineHost/directsound.c |
DirectSound 8 COM emulation (device, buffers, 3D listener, 3D buffers), the AudioQueue output path, watchdog, WAV capture and voice trace. |
native/EngineHost/directsound_mixer.h |
DsMixer, DsMixerVoice, DsMixer3D, DsMixerListener and the mixer API. |
native/EngineHost/directsound_mixer.c |
Format conversion, Catmull-Rom resampling, millibel gain, pan, 3D distance rolloff, cones, head-yaw panning, own-weapon lift, soft limiter, per-voice onset detection. |
native/EngineHost/directsound_engine_state.inc |
Read-only inspection of the original engine's own sound tables in guest memory, for the device report and the "voices held" log line. |
native/EngineHost/audio_ownership_trace.inc |
Default-off trace of four original sound-ownership functions (HALO_AUDIO_OWNERSHIP_TRACE=1). |
native/EngineHost/audio_source_identity.h |
XXH3-64 hash of each compressed stream handed to the Vorbis decoder (diagnostic identity, not a tag). |
native/EngineHost/vorbis_shim.c / .h
|
Native ov_open_callbacks, ov_read, ov_crosslap, ov_clear on top of stb_vorbis. |
native/EngineHost/third_party/stb_vorbis.c |
Public-domain/MIT Ogg Vorbis decoder, v1.22, compiled into vorbis_shim.c with STB_VORBIS_NO_STDIO. |
native/EngineHost/shims_misc.c |
Registers DirectSoundCreate8/DirectSoundCreate, enumeration stubs, all COM methods and the four VORBISFILE.dll exports in the import table. |
native/EngineVision/Sources/audio_session.inc |
visionOS AVAudioSession coordinator: activation, interruptions, media-services reset, foreground/background, route logging, the external watchdog timer. |
native/EngineHost/halo_settings.c |
Live settings read by the mixer: own-weapon gain (self_gain_db) and head yaw. |
flowchart LR
ENG["Translated Halo sound engine"]
VS["vorbis_shim.c (stb_vorbis)"]
GD["Buffer guest_data (32-bit guest memory)"]
DS["directsound.c COM shims"]
HV["DsMixerVoice.data (host heap copy)"]
MX["DsMixer: 768 voice slots"]
CB["AudioQueue callback: 48 kHz float32 stereo, 3 x 1536 frames"]
OUT["System audio output"]
HAP["haptics.c detectors"]
WD["host_dsound_watchdog"]
SES["audio_session.inc (visionOS)"]
ENG -->|"ov_open_callbacks / ov_read"| VS
VS -->|"16-bit PCM"| ENG
ENG -->|"Lock, write PCM"| GD
ENG -->|"Unlock"| DS
GD -->|"memcpy at Unlock"| HV
ENG -->|"Play, Stop, SetVolume, 3D params"| DS
DS --> MX
HV --> MX
MX -->|"ds_mixer_render"| CB
CB --> OUT
MX -->|"voice onsets + mixed signal"| HAP
WD -->|"dispose + rebuild queue"| CB
SES -->|"pause / resume / rebuild"| CB
SES -->|"250 ms timer"| WD
Key points:
- The engine writes PCM into guest memory (the buffer returned by
Lock). AtUnlockthe host copies those bytes into a host-owned array that the mixer reads. The audio thread therefore never touches guest memory (directsound.c:811-829). - There is exactly one output: an AudioQueue (on macOS and on visionOS, per the file header at
directsound.c:1-5).AVAudioSessionis used only for session management on visionOS; it does not carry samples. - No spatial audio framework is used. 3D is reduced to a stereo gain pair per voice by the mixer.
Halo calls COM methods through vtables in guest memory. build_vtables (directsound.c:211-225) allocates four guest vtables (IDirectSound8 with 12 methods, IDirectSoundBuffer8 with 24, IDirectSound3DListener8 with 18, IDirectSound3DBuffer8 with 21) and fills each slot with host_proc_address("DSOUND.dll", "Iface::Method"). Those names are registered as ordinary imports by HOST_DSOUND_SHIM_ENTRIES (directsound.h:89-118), which the header explains is so that "translated indirect calls use the same checked host boundary as ordinary DLL imports". See Win32 Compatibility Layer for the import/shim mechanism.
Each guest COM object is a 16-byte guest allocation: +0 vtable pointer, +4 index into the host's objects[] table. object_from_guest (directsound.c:481-487) validates the index, the alive flag and that the guest pointer matches either the buffer's main object or its 3D alias, so stale or forged pointers return NULL and the method returns DSERR_INVALIDPARAM.
Other DSOUND.dll exports (shims_misc.c:392-397):
| Export | Behaviour |
|---|---|
DirectSoundCreate8, DirectSoundCreate (and ordinals #11, #1) |
host_dsound_create8: rejects a non-null device GUID with DSERR_NODRIVER, builds vtables, starts audio (audio_start), returns the single device object. |
DirectSoundEnumerateA (#2) |
Calls the guest callback once with "Primary Sound Driver". |
DirectSoundCaptureEnumerateA (#9) |
Reports no capture devices. |
GetDeviceID (#14) |
Returns DSERR_NODRIVER. |
Defined at directsound.c:42-50; table objects[DS_OBJECT_MAX] with DS_OBJECT_MAX = 1024 live objects (slot 0 is never used).
| Field | Meaning |
|---|---|
guest, guest_3d
|
Guest COM object, and the lazily created 3D alias (listener for the primary buffer, IDirectSound3DBuffer8 for 3D secondaries). Both share the same slot index and reference count. |
guest_data |
Guest buffer the engine locks and writes (same size as the voice). |
refs, caps, primary, alive
|
COM reference count, DSBCAPS_* flags, primary-buffer flag. |
lock1, lock1_bytes, lock2_bytes, locked
|
The outstanding Lock region (wrapping split into two parts). |
unlock_*, last_*_frame
|
Trace-only counters (refills, bytes written, non-zero samples, output frame of last write and last non-silent write). |
voice |
The embedded DsMixerVoice the mixer reads. |
Defined in directsound_mixer.h:8-69.
| Structure / field | Meaning |
|---|---|
DS_MIXER_MAX_VOICES = 768 |
Mixer slots; every secondary buffer occupies one slot for its lifetime. |
DsMixerFormat |
tag (1 = PCM, 3 = IEEE float), channels, bits, block_align, sample_rate, avg_bytes_per_sec. |
data, bytes
|
Host copy of the buffer contents. |
cursor_frames (double) |
Fractional play position in source frames. |
end_count |
Natural completions of non-looping plays (kept across Play). |
volume, pan, frequency
|
DirectSound millibels (-10000..0), pan (-10000..10000), playback rate in Hz. |
playing, looping, has_3d, pending_3d
|
State flags. |
onset_* |
Per-voice onset detector state for haptics (see Haptics). |
self_voice |
Latched when the voice has sounded within 0.5 units of the listener since its last Play (the player's own weapon). |
current_3d, deferred_3d (DsMixer3D) |
Position, velocity, cone orientation, min/max distance, cone angles, outside volume, mode. |
DsMixer.lock |
One mutex protecting all voices and the listener. |
DsMixer.current_listener, deferred_listener, pending_listener
|
Listener position, velocity, front/up, distance/rolloff/doppler factors. |
rendered_frames, ended_voices, limited_samples, rendered_samples
|
Counters for diagnostics. |
IDirectSound8::CreateSoundBuffer (host_dsound_device_3, directsound.c:706-723):
- Validates the
DSBUFFERDESC(size ≥ 20). A primary buffer must pass zero bytes and no format. - For a secondary buffer, rejects 0 bytes, more than 128 MiB, an unreadable format, or a size that is not a multiple of
block_align(DSERR_BADFORMAT). -
DSBCAPS_LOCHARDWAREis not refused. The comment explains that Vista-and-later Windows mixes every buffer in software and never fails this; refusing it "made Halo believe the device had no hardware voices at startup". -
object_new(directsound.c:518-549) finds the first free slot (the limit is concurrent buffers, not buffers ever created),callocs host data, adds the voice to the mixer, allocatesguest_dataand the 16-byte guest object. Any failure unwinds every allocation and the mixer slot.
read_format (directsound.c:492-506) accepts mono or stereo; 8- or 16-bit PCM or 32-bit float; 100..200000 Hz; WAVE_FORMAT_EXTENSIBLE (0xFFFE) whose subformat is PCM or float. block_align and avg_bytes_per_sec must be consistent.
DuplicateSoundBuffer creates a new secondary with the same caps/format and copies the host data.
buffer_qi (directsound.c:586-599) matches only the first 32 bits (Data1) of the IID:
| IID Data1 | Interface | Result |
|---|---|---|
0x00000000 |
IUnknown | The buffer itself |
0x279AFA85 |
IDirectSoundBuffer | The buffer itself |
0x6825A449 |
IDirectSoundBuffer8 | The buffer itself |
0x279AFA84 |
IDirectSound3DListener | Primary buffer only: a guest alias with the listener vtable |
0x279AFA86 |
IDirectSound3DBuffer | 3D secondary only: a guest alias with the 3D-buffer vtable |
Every successful QueryInterface adds a reference to the shared DsObject.
object_release (directsound.c:565-583) frees everything when the shared count reaches zero: it removes the voice from the mixer under mixer.lock before freeing voice.data, and keeps holding objects_lock until the slot is zeroed. The comment states why: "Keep the slot owned until the callback can no longer see its voice; otherwise a creator could overwrite it during removal."
The device object (device_guest) is created once and never freed; its AddRef/Release only move a counter.
| Method (index) | Implementation notes |
|---|---|
GetCaps (3) |
Size, caps, buffer bytes. |
GetCurrentPosition (4) |
Play cursor and write cursor; see below. |
GetFormat (5) |
Writes an 18-byte WAVEFORMATEX. |
GetVolume/GetPan/GetFrequency (6-8) |
Read under mixer.lock. |
GetStatus (9) |
PLAYING, plus LOOPING only while playing. Counts probes and busy answers. |
Initialize (10) |
DSERR_ALREADYINITIALIZED. |
Lock (11) |
Region from an offset, from the write cursor (DSLOCK_FROMWRITECURSOR) or the whole buffer (DSLOCK_ENTIREBUFFER); splits at the end of the ring into two parts; a second Lock before Unlock is DSERR_INVALIDCALL. |
Play (12) |
Only flag allowed is DSBPLAY_LOOPING. Sets playing, looping and calls ds_mixer_voice_restart (resets onset state and self_voice). Does not move the cursor. |
SetCurrentPosition (13) |
Sets cursor_frames = pos / block_align; out-of-range positions rejected. |
SetFormat (14) |
Primary buffer only. |
SetVolume (15) |
Requires DSBCAPS_CTRLVOLUME; -10000..0 mB. |
SetPan (16) |
Requires DSBCAPS_CTRLPAN; -10000..10000. |
SetFrequency (17) |
Requires DSBCAPS_CTRLFREQUENCY; 0 restores the format rate; otherwise 100..200000. |
Stop (18) |
Clears playing; keeps the cursor (Stop then Play resumes). |
Unlock (19) |
Must exactly match the outstanding Lock; copies both regions into host data under mixer.lock. |
Restore (20) |
DS_OK for a valid object. |
SetFX, AcquireResources, GetObjectInPath (21-23) |
DSERR_UNSUPPORTED. |
IDirectSound8::GetCaps reports all primary formats (dwFlags = 0x1F), a continuous rate range 100..200000 Hz and one primary buffer, with all hardware-voice counts zero (directsound.c:724-730). GetSpeakerConfig reports 4 (stereo).
Halo streams every sound through a fixed pool of looping DirectSound buffers ("voices"). The original allocator 005482E0 only reuses a stopped voice when 00547FF0 sees neither PLAYING nor LOOPING from GetStatus. voice_status (directsound.c:168-180) therefore reports LOOPING only together with PLAYING, as Windows and Wine do. The comment records the failure this prevents: reporting LOOPING for a stopped buffer "silenced Build75": each of the 24 positional voices played one sound and was then held until the next sound_stop_all, about 35 seconds of effects after each level load, pause or checkpoint revert.
Similarly, voice_write_cursor (directsound.c:181-195) returns a write cursor that leads the play cursor only while playing; a stopped buffer reports both cursors at the same position. Halo's 00547C80 compares the pair to decide whether a draining voice is still running; unequal cursors on a stopped voice made it skip the Play.
The playing write lead is sample_rate / 20 frames (50 ms), computed in whole frames (directsound_mixer.c:190-199); computing it in bytes had produced half-frame cursors for 22,050 Hz streams.
sequenceDiagram
participant E as Halo (00547C80 / 005482E0)
participant D as directsound.c
participant M as DsMixer
participant Q as AudioQueue callback
E->>D: GetStatus(voice) (allocator probe)
D-->>E: 0 (stopped, free)
E->>D: Lock(0, ring bytes)
D-->>E: pointer(s) into guest_data
E->>E: write PCM into guest memory
E->>D: Unlock(regions)
D->>M: memcpy guest_data to voice.data (mixer.lock)
E->>D: SetCurrentPosition(0)
E->>D: Play(DSBPLAY_LOOPING)
D->>M: playing=1, looping=1, restart onset state
loop every buffer period
Q->>M: ds_mixer_render(frames, 48000)
end
E->>D: Lock(FROMWRITECURSOR) / Unlock (refill ring)
E->>D: Stop()
D->>M: playing=0 (cursor kept)
All 3D setters exist for both immediate and deferred application (directsound.c:835-882):
- Every setter writes the deferred copy (
deferred_listenerorvoice.deferred_3d). WithDS3D_DEFERRED(non-zero last argument) it only markspending_*; withDS3D_IMMEDIATEit copies the whole deferred state tocurrent. -
IDirectSound3DListener8::CommitDeferredSettingscallsds_mixer_commit(directsound_mixer.c:53-67), which applies the pending listener and every voice's pending 3D parameters under one lock. - Float arguments are read from the guest stack at
ESP+8(after the return address andthis), and vectors as three consecutive floats. All values are validated: vectors must be finite; distance factor > 0; rolloff and doppler ≥ 0; cone angles ≤ 360 with inside ≤ outside; outside volume -10000..0; min distance > 0 and ≤ max; mode 0..2. - Voice defaults (
directsound.c:511-517): min distance 1, max distance 1e9, cone angles 360/360. - Listener defaults (
directsound_mixer.c:25-34): front +Z, up +Y, all factors 1.
Velocities and the Doppler factor are stored and returned by the getters, but the mixer does not apply Doppler shift. With HALO_AUDIO_VOICE_TRACE on, listener_changed logs the first eight listener writes and every 500th ([audio-listener]), because, as the comment says, whether DirectSound should also attenuate depends on the rolloff factor the game sets.
ds_mixer_render (directsound_mixer.c:230-341) runs on the AudioQueue callback thread with mixer.lock held for the whole pass.
flowchart TD
S["For each of 768 slots"] --> P{"playing and has data?"}
P -- no --> S
P -- yes --> G["ds_mixer_voice_gains: volume, pan, 3D"]
G --> Z{"both gains zero?"}
Z -- yes --> ADV["Advance cursor by step x frames, wrap or end; no resampling"]
Z -- no --> R["Per output frame: Catmull-Rom from 4 source frames"]
R --> ACC["Accumulate left*gain_l, right*gain_r"]
ACC --> ON["Positional voices: onset envelopes (haptics)"]
ON --> R
ADV --> S
R --> S
S --> LIM["After all voices: count samples past 1.0, apply soft knee"]
LIM --> HO["Unlock, then halo_haptics_observe(mixed output)"]
read_channel (directsound_mixer.c:69-84) converts a sample to float: 8-bit unsigned (x-128)/128, 16-bit signed x/32768, 32-bit float clamped to ±1 (non-finite becomes 0). Mono voices feed the same value to both channels.
Every voice is resampled to the 48 kHz output with step frequency / output_rate. The interpolator is four-tap Catmull-Rom (directsound_mixer.c:201-211). The comment explains: Halo's effects are 22,050 Hz, linear interpolation at that ratio acts as a low-pass "which is what makes the mix sound muffled". A looping voice wraps its neighbour taps across the loop seam; a one-shot holds its end frames. When a non-looping voice runs past its last frame it stops, end_count and ended_voices increment.
A voice whose computed gain is zero in both channels (for example one Halo has muted to -10000 mB but keeps streaming into) is advanced, wrapped or ended exactly as if mixed, without the resampling cost (directsound_mixer.c:253-263).
ds_mixer_voice_gains (directsound_mixer.c:86-165):
-
Volume:
millibel_gain(v) = 10^(v/2000), with ≤ -10000 mapped to 0 and ≥ 0 to 1. -
Pan: positive pan attenuates the left channel by
millibel_gain(-pan), negative pan attenuates the right. -
3D (if
has_3dand mode ≠ 2DS3DMODE_DISABLE):- Relative position: mode 0 (normal) subtracts the listener position; mode 1 (head-relative) uses the position as is.
-
Distance rolloff: beyond
min_distance,gain *= (min / min(distance, max)) ^ rolloff. The listener's distance factor is deliberately not applied. The comment: Halo sets it to 3.048 (ten feet in metres); multiplying the distance by it "quietly pushed every positional sound 9.7 dB down while 2D music was untouched". -
Own-weapon lift: if the distance is < 0.5 units or
self_voiceis latched,gain *= halo_settings_self_gain(). Halo hands the player's own weapon, reload and melee to DirectSound 20 dB down (the assault rifle at -2000 mB); in a headset recording firing raised the mix by under a decibel. -
Spatial pan: the direction is projected on the listener's right vector (
up x front). Before that, the right/front pair is rotated by the viewer's head yaw (halo_settings_head_yaw()), so a sound keeps its place in the world when only the head turns in the panorama.p = dot(direction, right)in [-1, 1];p > 0multiplies the left gain bysqrt(1-p),p < 0multiplies the right gain bysqrt(1+p). -
Cones: if a cone orientation is set and
outside_angle < 360, the full angle to the listener (twice the half-angle) is compared with the inside/outside angles and the outside volume is interpolated linearly in millibels between them.
self_voice is latched during rendering when a positional voice is within 0.5 units of the listener (directsound_mixer.c:250-251) and cleared by Play. The header comment explains why it is latched rather than judged per block: a moving player left the shot more than half a unit behind mid-sound, the 18 dB lift fell away, and the shot was heard "cutting out".
Head pitch is not used by the mixer; only yaw.
After all voices are summed, every sample passes through soften (directsound_mixer.c:213-228): unity below a 0.90 knee, then knee + (1-knee) * over/(1+over) with over = (|x|-knee)/(1-knee), so the output never reaches ±1. A full-scale peak loses about 0.45 dB; a mix at twice full scale still lands under the ceiling. Samples that would have exceeded ±1 before the curve are counted in limited_samples; host_dsound_get_stats logs DirectSound limiter: N of M samples past full scale at most every 10 seconds of output (directsound.c:641-650).
| Setting | Source | Default | Range | Effect |
|---|---|---|---|---|
HaloSettings.self_gain_db |
HALO_SELF_GAIN_DB seeds it (halo_settings.c:64); live slider "Your own weapon" in the Presentation settings window (EngineSettingsView.swift:108-111) |
18 dB | 0..24 dB (out-of-range environment values fall back to 18) | Linear gain 10^(dB/20) applied to positional voices on the listener (halo_settings_self_gain, halo_settings.c:136). The haptics onset detector divides it back out so the player's weapon is judged at Halo's own level. |
| Constant | Value | Where | Meaning |
|---|---|---|---|
OUTPUT_RATE |
48000 Hz | directsound.c:35 |
Output sample rate; also the mixer's render rate. |
OUTPUT_BUFFERS |
3 | same | AudioQueue buffers in circulation. |
output_frames_per_buffer |
1536 | directsound.c:36-40 |
Frames per buffer (32 ms). Total queued latency 96 ms. |
OUTPUT_FRAMES |
2048 | same | Upper bound for HALO_AUDIO_FRAMES and per-callback render clamp. |
| Format | linear PCM, float32, packed, 2 channels, 8 bytes/frame | directsound.c:317-324 |
Interleaved stereo float. |
Why 1536: the comment at directsound.c:36-39 says Build80 recorded callback gaps of up to 50 ms; with 512-frame buffers the two remaining buffers covered only 21.3 ms, so a healthy queue could run dry without an enqueue error. 1536 frames keeps 64 ms of reserve after each completed buffer. test_audio_watchdog.c asserts (OUTPUT_BUFFERS-1) * frames * 1000 / OUTPUT_RATE > 50.
| Variable | Meaning |
|---|---|
output_requested |
Set by the first DirectSoundCreate8 (audio_start, directsound.c:386-394). Until then no queue exists. |
output_suspended |
Set by host_dsound_pause_output, cleared by host_dsound_resume_output. A suspended output never starts. |
output_restart_required |
Forces an explicit AudioQueueStart on the next resume. Set at creation, on pause, on dispose and on every resume request. |
output_callbacks_suspended, output_callbacks_discarding
|
Callback gate (see below). |
output_deferred_buffers[3] |
Buffers that completed while paused, held for re-enqueue on resume. |
output_enqueue_fault |
Last failing AudioQueueEnqueueBuffer status from the callback, retained until a rebuild. |
output_terminated |
Set by host_dsound_shutdown; nothing restarts afterwards. |
output_generation |
Incremented on every successful AudioQueueNewOutput. |
audio_prepare_locked (directsound.c:302-343) reads HALO_AUDIO_FRAMES and the capture variables once, creates the queue with AudioQueueNewOutput(..., NULL run loop, ...) (callbacks on AudioQueue's own thread), allocates the three buffers, primes each with silence and enqueues it. Any failure disposes the queue and returns 0.
audio_resume_locked (directsound.c:345-384) prepares if needed and, when a restart is required or kAudioQueueProperty_IsRunning reads 0, re-enqueues deferred buffers as silence, opens the callback gate and calls AudioQueueStart. It does not trust IsRunning: on macOS it stays true for a paused queue, and the Build60 headset report showed it true after backgrounding while callbacks had stopped for good (directsound.c:408-420).
audio_callback (directsound.c:227-275), under output_callback_lock:
- If terminated or discarding, return (the buffer is dropped).
- If suspended, remember the buffer in
output_deferred_buffers(once) and return without rendering, so no voice advances while paused. - Record the gap since the previous callback (
max_callback_gap_ns; a gap over two buffer periods counts as a late callback). -
ds_mixer_renderinto the buffer, then measure peak and non-zero samples into atomics (output_frames,output_nonzero,output_peak_bits). - If WAV capture is active, convert to 16-bit and append (bounded).
- Enqueue. A failure is recorded in
telemetry_enqueue_failuresandoutput_enqueue_fault. - Record the callback's own work time (
max_callback_work_ns).
The callback never writes host.log: "stderr shares its file lock with engine/diagnostic writers and may block behind disk I/O" (directsound.c:258-260). All logging happens on the diagnostics/control side.
| Function | Lines | Behaviour |
|---|---|---|
host_dsound_pause_output |
396-406 |
Marks suspended and restart-required, closes the callback gate (without discarding), AudioQueuePause. Voices, cursors, diagnostics and capture are preserved. |
host_dsound_resume_output |
408-420 |
Clears suspended, forces a restart, audio_resume_locked. Records last_recovery 0 or -1. |
host_dsound_rebuild_output |
469-479 |
Disposes the queue and prepares a new one (only if requested and not suspended); does not start it. Callers follow it with resume. |
host_dsound_shutdown |
601-608 |
Sets output_terminated, disposes, finalises the capture WAV header. Called by host.c after the engine has stopped (host.c:659). |
audio_dispose_locked (directsound.c:288-300) closes the gate in discard mode first, because AudioQueueDispose can synchronously return buffers through the callback, and "those callbacks must not mix again" (directsound.c:277-279).
-
output_lock(queue control) is taken beforeoutput_callback_lock(callback gate);output_callback_lockis released beforeAudioQueuePause/Disposecan call back. -
objects_lockis taken beforemixer.lock(host_dsound_get_stats,host_dsound_get_voice_stats). - The callback holds
output_callback_lockand thenmixer.lock(insideds_mixer_render). COM methods take onlymixer.lock(orobjects_lockalone).
host_dsound_watchdog (directsound.c:426-466) rebuilds the queue when output is requested, not suspended, not terminated, and either:
- a callback enqueue failed (
output_enqueue_fault), or -
output_frameshas not advanced for 0.5 s and the queue has been started for at least 0.5 s (resume grace).
Rebuilds are rate-limited to one every 2 seconds. Progress is rechecked after taking output_lock, because a callback may have finished while the watchdog waited, and the fault is rechecked because another transaction may already have replaced the queue. A rebuild is audio_dispose_locked + audio_resume_locked inside one locked transaction; a failed preparation leaves no queue but stays eligible for the next retry. Each rebuild logs [audio-watchdog] ... rebuilt the queue (N so far): playing|FAILED and increments watchdog_rebuilds (reported as audio_rebuilds).
Who runs it:
- By default, the engine thread once per
IDirect3DDevice9::Present(d3d9.c:544-546). - On visionOS,
audio_session_preparecallshost_dsound_watchdog_set_external()and runs it from a 250 ms timer on the audio control queue instead. The comment explains that the per-frame watchdog "slept through exactly the stalls it exists for: level loads and the engine's own stops (Build73's queue stopped at 290.4 s and nothing recovered it)" (audio_session.inc:62-75).
audio_session.inc is included into EngineVisionRuntime.m (EngineVisionRuntime.m:606) and is compiled only when TARGET_OS_VISION. enginevision_start calls audio_session_prepare() before creating the engine worker (EngineVisionRuntime.m:726-728).
All session mutations and queue recovery run on the serial queue HaloVision.audio-session, never on the render callback.
| Element | Lines | Behaviour |
|---|---|---|
audio_session_can_resume |
13-17 |
True only if the runtime is STARTING or RUNNING, the app is foreground, not interrupted and resume is allowed. |
audio_session_resume |
29-52 |
setCategory:Playback mode:Default options:0, setActive:YES, optional host_dsound_rebuild_output, then host_dsound_resume_output. On any failure: pause the output and schedule a retry 2 s later (audio_retry_after_ns). |
audio_session_tick |
54-60 |
Every 250 ms: if a resume is pending and its retry time has come, retry the whole transaction; otherwise run host_dsound_watchdog. |
audio_session_route |
20-28 |
Logs [audio-route] with output port types only (never names or UIDs), volume, rate, channel count; at most 64 reports. |
Notifications (audio_session.inc:76-137):
| Notification | Action |
|---|---|
AVAudioSessionInterruptionNotification began |
audio_interrupted = YES, pause the output. |
| ... ended | Clear interrupted; allow resume only if ShouldResume was set; attempt resume. |
AVAudioSessionMediaServicesWereResetNotification |
Mark rebuild needed (orphaned queue) and resume. |
UIApplicationDidBecomeActiveNotification |
Clear interruption and resume_allowed=NO (coming back is the player's own action; keeping shouldResume=NO "left the mix silent for the rest of the session"), set foreground, resume. |
UIApplicationDidEnterBackgroundNotification |
Clear foreground, pause. The app has no audio background mode; a queue left running stalls and every background rebuild failed in Build72 (OSStatus 'what'). |
AVAudioSessionRouteChangeNotification |
Log only. |
stateDiagram-v2
[*] --> Startup
Startup --> Playing: activation + resume ok
Startup --> RetryPending: activation/rebuild/resume failed
RetryPending --> Playing: retry after 2 s succeeds
RetryPending --> RetryPending: retry fails
Playing --> Interrupted: interruption began (pause)
Interrupted --> Playing: ended with ShouldResume
Interrupted --> Waiting: ended without ShouldResume
Waiting --> Playing: app became active
Playing --> Background: entered background (pause)
Background --> Playing: became active
Playing --> Playing: watchdog rebuild (stall or enqueue fault)
Playing --> Playing: media services reset (rebuild + resume)
On macOS (desktop host and probes) there is no session coordinator; the queue is started on device creation and the watchdog runs from Present.
Halo imports four libvorbisfile functions; vorbis_shim.c implements exactly those, registered in shims_misc.c:412-415 and shims_misc.c:444. All are cdecl.
| Constant | Value |
|---|---|
VORBIS_MAX_STREAMS |
16 simultaneous streams |
VORBIS_READ_CHUNK |
64 KiB guest read chunk |
VORBIS_MAX_INPUT |
256 MiB maximum compressed stream |
| Error codes |
OV_EREAD -128, OV_EFAULT -129, OV_EIMPL -130, OV_EINVAL -131, OV_ENOTVORBIS -132
|
ov_open_callbacks(source, vf, initial, ibytes, read, seek, close, tell) (vorbis_shim.c:73-92):
-
read_source(vorbis_shim.c:38-71) pulls the entire compressed stream into host memory through the guest's own read callback (host_call_guest), using a 64 KiB guest scratch buffer. If seek and tell callbacks exist and there is no initial prefix, it seeks to the end to size the buffer, then back. It keeps reading until a zero read (a short read is not end of input; an older loop exited early), doubles capacity as needed, rejects a callback that returns more than asked, and rejects streams over 256 MiB rather than truncating them. -
stb_vorbis_open_memorydecodes headers; only 1- or 2-channel streams are accepted. - A free slot in
streams[16]is claimed for thisvf(the guest'sOggVorbis_Fileaddress is used purely as a key).
ov_read(vf, out, length, bigendian, word, sgned, bitstream) (vorbis_shim.c:94-122): only little-endian signed 16-bit is supported (otherwise OV_EINVAL). Calls stb_vorbis_get_samples_short_interleaved for length/2 shorts and returns bytes produced (0 at end). A request smaller than one frame (length < channels*2) returns 0 without consuming input or marking end of stream. *bitstream is always 0. The last up to 4096 output samples are kept in tail for crosslapping.
ov_crosslap(old, new) (vorbis_shim.c:124-126): requires equal channel counts; copies the old stream's tail into the new stream's overlap. The next ov_read on the new stream linearly cross-fades its first n samples from the overlap (old*(1-t) + new*t, t = (i+1)/(n+1)). This is a simple linear blend, not libvorbisfile's windowed lapping.
ov_clear(vf) (vorbis_shim.c:128-137): frees the decoder and compressed copy, then calls the guest's close callback and returns its result.
The decoded PCM is returned into guest memory; the engine then writes it into a DirectSound buffer like any other sound. With voice tracing on, opens/reads/clears are logged as [audio-voice] vorbis-* with the stream generation, decoder offsets and an XXH3-64 identity-xxh3 of the compressed input (only computed when tracing).
| Function | Returns |
|---|---|
host_dsound_get_stats |
Total output frames, non-zero samples, peak; also emits the one-time "mix became active" line, the periodic limiter line, and (when tracing) the 5-second voice summary. Must not be called from the callback. |
host_dsound_get_output_diagnostics |
HostDsOutputDiagnostics (directsound.h:26-35): generation, lifecycle flags, last OSStatus of prepare/start/enqueue/pause/dispose/recovery, callbacks, late callbacks, enqueue failures, last callback time, max gap, max work, sample rate, buffer frames. |
host_dsound_get_output_state |
kAudioQueueProperty_IsRunning and its OSStatus. |
host_dsound_get_voice_stats |
HostDsVoiceStats (directsound.h:36-59): host Play/Stop/GetStatus counters and buffer counts, plus the engine's own tables. |
host_dsound_watchdog_rebuilds |
Count of watchdog rebuilds. |
On visionOS these feed the device report through EngineDiagnosticsBridge.m:67-105, which also samples AVAudioSession.outputVolume. See Diagnostics and Telemetry.
directsound_engine_state.inc documents and reads the original sound engine's state in guest memory (for halo.exe build c9acf0c46954, per its header). From its comment (directsound_engine_state.inc:1-34):
| Guest address | Table |
|---|---|
0x007252C0 |
Pointer to the sound-source data array (512 x 0xB0), filled by sound_impulse_start (00549AF0). |
0x00724A50 |
Looping sounds (128 x 0xE4). |
0x00724A60, count at 0x007252B4
|
Channels (0x18 each; +0 source or -1). |
0x007252E4 |
Per-channel voice map (+0 voice or -1, +2 voice type). |
0x00725430, count at 0x00725428
|
Voice records (0x678 each): +0 state, +2 channel, +8 stopped, +9 draining, +0x38 type, +0x670 IDirectSoundBuffer8. |
0x0069F528 |
Voice type table (the first type is the positional voices). |
0x006AC528 |
Sound cache data array (512 x 16: +2 loaded, +5 lock count). |
0x006AC4A0 |
Cache read requests (512 x 0x30: +0x1D queued, +0x20 file: level map, bitmaps.map, sounds.map). |
0x00725200, 0x00725201, 0x00725202, 0x007252B6
|
Initialised, enabled, paused by the game, disabled flags. |
Every data array is checked for the "d@t@" signature, capacity, element size and bounds before being walked; reads begin only after the guest has created its DirectSound device (engine_sound_observable). sound_engine_voices_locked classifies each unassigned voice as free or held by replicating 00547FF0's test against the host's GetStatus answer. voice_exhaustion_note logs [audio-engine] voices held once when some voice type has held voices but no free one (Build75's failure signature) and voices free when it clears (at most 64 notes).
When enabled, host_dsound_get_stats emits, at most once per 5 s of output, an [audio-voice] summary line (active/looping voices, ends, Play/Stop/probe counters, engine flags, sources, channels, starved channels, voices free/held by type, cache, queued reads), then one voice line per playing or refilled buffer. Allocation, play, stop, position, refill (on silent/non-silent transitions), failures and Vorbis events are logged inline. Detail lines are capped at 256 per 5-second window, with a dropped counter (directsound.c:145-162). The switch is "lazy-on": host_dsound_trace_enabled re-reads the environment until it finds it set, so an early UI poll cannot cache it off before the engine worker exports its defaults.
host_audio_ownership_dispatch (audio_ownership_trace.inc:61-87) is called from engine_dispatch_override (overrides.c:249) for the hook addresses 0x00442550, 0x00544090, 0x00544120 and 0x0048A1A0 (registered in engine_hooks.h). It logs a table-observed line when the tag table pointer (0x0087BC14), header (0x006A8954) or valid flag (0x006A8150) changes, and before/after snapshots around the first three functions (for 0x00442550 only when its first stack argument is 0x0066A044), including the current scenario's name, the tag's name and definition, and a salted record from the pool at 0x007461A0. It executes the wrapped function exactly once, never writes guest memory, rejects stale salts, and stops after 512 reports. The code does not document the semantic name of each wrapped function beyond "original sound ownership observations"; see Engine Overrides and Hooks.
HALO_AUDIO_CAPTURE_WAV=<path> writes the final mix as 16-bit stereo 48 kHz WAV from the callback (no allocation; a stack buffer), for HALO_AUDIO_CAPTURE_SECONDS (default 30, 1..1800; the comment notes thirty seconds "only reaches the opening cutscene"). The header is rewritten with the real length at shutdown. The capture can be fed to tests/haptics_from_capture.c (see Haptics).
| Variable | Default | Effect | Read at |
|---|---|---|---|
HALO_AUDIO_FRAMES |
1536 | Frames per AudioQueue buffer, accepted 64..2048; read once at first queue preparation. | directsound.c:306-307 |
HALO_AUDIO_CAPTURE_WAV |
unset | Path of a diagnostic WAV capture of the output mix. | directsound.c:309-315 |
HALO_AUDIO_CAPTURE_SECONDS |
30 | Capture length, 1..1800 s. | directsound.c:314 |
HALO_AUDIO_VOICE_TRACE |
off; 1 on visionOS (device defaults, overridable) |
Enables voice/listener/Vorbis tracing and the 5-second summaries. Any value other than empty or 0 enables it. |
directsound.c:99-106, EngineVisionRuntime.m:634
|
HALO_AUDIO_OWNERSHIP_TRACE |
off | Exactly 1 enables the ownership trace (max 512 reports). |
audio_ownership_trace.inc:63-64 |
HALO_SELF_GAIN_DB |
18 | Seeds the own-weapon lift, 0..24 dB. | halo_settings.c:64 |
HALO_SIM_HEAD |
unset | Simulator only: yaw,pitch in degrees replaces the head pose; the yaw reaches the mixer's spatial pan. |
EngineImmersive.swift:355-360 |
HALO_PROBE_SUSPEND_AUDIO |
unset | Menu-input probe only: 1 pauses output before the engine starts (rendering-only measurements). |
MenuInputProbe.m:844-848 |
HALO_ENGINE_TRANSLATION |
newest native/build/engine-reuse/*/sub_005482E0.c
|
Source-check runner: directory of translated functions for the translated-allocator variant of the voice-reuse test. | run_source_checks.py:155 |
| Thread | Audio work |
|---|---|
| Engine thread (and other translated guest threads) | All COM calls, Vorbis calls (guest callbacks run on the caller's CPU context), per-frame watchdog on macOS. |
| AudioQueue callback thread |
audio_callback, ds_mixer_render, haptics onset and observe detectors. No logging, no guest memory access. |
visionOS HaloVision.audio-session serial queue |
Session activation, interruptions, background/foreground, 250 ms watchdog/retry tick. |
| Diagnostics/UI |
host_dsound_get_*; all logging that could block. |
| Symptom | Mechanism |
|---|---|
Callbacks stop while IsRunning stays true (Build60) |
Resume always restarts; watchdog rebuilds after 0.5 s without progress. |
| A buffer lost to a failed callback enqueue |
output_enqueue_fault is retained across later successful enqueues; watchdog rebuilds even though frames advance. |
| Queue stalls during level loads (Build73) | External 250 ms watchdog on visionOS, independent of Present. |
| Session activation fails | Output paused, whole transaction retried every 2 s while the runtime runs and the app is in the foreground. |
| Media services reset | Queue rebuilt on the next resume. |
| Backgrounding | Output paused cleanly; restarted on foreground. |
| Effects go silent after ~35 s (Build75) |
GetStatus reports LOOPING only while playing; held voices logged by voice_exhaustion_note. |
| Positional sounds ~10 dB too quiet | Distance factor excluded from rolloff. |
| Own weapon inaudible under music |
self_gain_db lift, latched per sound. |
| Muffled sound | Catmull-Rom interpolation instead of linear. |
| Buzz in loud firefights | Soft-knee limiter instead of hard clipping. |
The release notes state the limits of this evidence: "Audio interruptions have been reported. Synthetic mixer checks do not establish uninterrupted audible playback on a headset" (docs/KNOWN_ISSUES.md).
C tests in native/EngineHost/tests/ are built and run by tools/run_source_checks.py unless noted (see Testing and Source Checks).
| Test | Runner | What it asserts |
|---|---|---|
test_directsound_mixer.c |
Mac | Exact resampled values (Catmull-Rom 0.1875 between ramp points, a full-scale frame softened to 0.95), volume/pan attenuation, distance factor does not change rolloff, 4x min distance gives 0.25, own-weapon lift persists after the player moves away, muted voices still advance/end/wrap, removed voices produce silence. |
test_mixer_quality.c |
portable | 1 kHz tone 22,050 to 48,000 Hz has < 0.2 % non-fundamental power; quiet material passes the limiter untouched; eight full-scale voices stay ≤ 1.0 with no pinned samples; head yaw ±90° moves a frontal sound to the opposite ear; rolloff unaffected by distance factor. |
test_audio_watchdog.c |
Mac | Mocked AudioQueue and clock: buffer reserve > 50 ms, failed rebuild retries with a 2 s bound, single locked transaction, resume grace, concurrent progress while waiting for the lock, deferred paused buffers re-enqueued exactly once, failed Start keeps the gate closed, enqueue loss detected despite later progress, termination stops recovery, telemetry (late callbacks, gaps exclude pauses, 48000/1536). |
test_directsound_lifetime.c |
Mac | 2048 create/play/3D-alias/release cycles with a concurrent render thread; stale pointers rejected; allocation failure at each stage leaks nothing; 768-voice mixer capacity and 1023-object table recover after release. |
test_directsound_voice_reuse.c |
Mac (plus translated-allocator variant when generated code exists) | Drives Halo's allocator (C transcription of 005482E0/00547FF0, or the translated functions with -DHALO_TRANSLATED_ALLOCATOR) through the shim: 200 sounds on 24 positional voices all start; GetStatus/cursor semantics; engine-table statistics; held/free log lines once each. -DBUILD75_GET_STATUS reproduces the Build75 silence (24 sounds, then none). |
test_guest_heap_sound_owners.c |
Mac | Production guest heap with 4096 recycled buffer/3D-alias lifetimes and a concurrent renderer; shared references keep guest data alive; final release reclaims every block. |
test_audio_shim_boundaries.c |
Manual (needs an owned sounds.map) |
Vorbis input reader (> 512 KiB, short reads, prefix, seek/no-seek, invalid callback); DirectSound ABI (lazy trace enable, 5 s summary cadence, 256-event cap, mono/stereo cursor alignment, split refills, rejected calls, Stop/Play resume, natural end); Vorbis ABI (undersized reads, full drain equals fresh stream, close/reuse); pause/resume/terminated diagnostics without queue mutation. |
test_halo_vorbis.c |
Manual (needs sounds.map) |
stb_vorbis opens the first Ogg stream in sounds.map, 1-2 channels, 11025-48000 Hz, decodes non-zero PCM. |
test_audio_source_identity.c |
Manual | Equal inputs hash equally; a final-byte or length change differs; input untouched. |
test_audio_ownership_trace.c |
Manual | Off by default with no side effects; when on, the wrapped function runs exactly once, CPU state is preserved, guest memory is unchanged, stale salt yields flags=ffffffff, report cap respected. |
AudioSessionRecoveryValidation.m |
Mac | Includes the production audio_session.inc with a mocked AVAudioSession and host functions: failed activation retried after 2 s (not sooner), media reset rebuilds, failed rebuild/resume retried, background and interruption (without ShouldResume) suppress resume, stopped runtime suppresses resume. |
- Haptics: the mixer's onset detector and the mixed-signal detector.
- Input and Controllers: the trigger state the haptics fire window uses.
- Win32 Compatibility Layer: import table and shim calling conventions.
-
Guest Memory and Heap:
guest_alloc/guest_freeused for COM objects and buffers. - Threading and Synchronization
- Engine Overrides and Hooks: dispatch hooks used by the ownership trace.
- Immersive Presenter: source of the head yaw used for spatial panning.
-
Runtime Settings:
self_gain_dband the settings window. - Diagnostics and Telemetry: device report fields.
- Environment Variables
- Testing and Source Checks
Documents master-chef at commit 9f915af (v1.0.3). Unofficial project, not affiliated with Microsoft, Bungie, Gearbox or Apple. Original code is MIT licensed; game content is not included.
Overview
- Architecture Overview
- Repository Layout
- Glossary
- Environment Variables
- Contributing Guide
- Open Questions
Translation
- Static Translation Pipeline
- XWA Decoder and Lifter
- Function Address Lists
- EngineReuse Runtime
- x87 Floating Point
Host runtime
- EngineHost Overview
- Win32 Compatibility Layer
- Threading and Synchronization
- Guest Memory and Heap
- Engine Overrides and Hooks
- Runtime Settings
Graphics
- Direct3D9 Bridge
- Metal Renderer
- Shader Translation
- Textures and Texture Packs
- Geometry Fast Paths
- Radial Fog
Panorama and presentation
- Panorama System
- Panorama Budget and LOD
- Frame Pacing
- visionOS App
- Immersive Presenter
- Layer Alignment
Audio and input
Tooling and process