Repository navigation
Threading and Synchronization
Halo PC is a multi-threaded Windows program: besides the main game thread it creates a cache-file reader thread and a saved-game writer thread, and it synchronizes them with events, mutexes, critical sections and asynchronous file I/O. In the port, every guest thread is a real native pthread that runs translated code against its own EngineCPU, guest stack and guest TEB, and every Win32 synchronization object is a native mutex/condition-variable pair behind a handle. threading.c implements the objects and threads; shims_kernel32.c exposes them as Win32 APIs and adds TLS, APC delivery and Sleep; host.c provides the per-thread guest context. This page describes the thread model, each object type, waits, cancellation, quality of service, and the wait telemetry the device report consumes.
| File | Role |
|---|---|
threading.h |
Public API: wait result codes, event/mutex/thread/critical-section functions, host_thread_wait_ns/host_thread_waits, host_core_telemetry, QoS report accessors. |
threading.c |
HostEvent, HostMutex, HostThread, HostCriticalSection; cancellation-safe waits; QoS selection; wait telemetry. |
host.c |
host_initialize_guest_thread (guest stack + TEB + TLS), host_execute_guest_thread (run with a per-thread setjmp escape), thread-local host_active_cpu, host_exit. |
shims_kernel32.c |
Win32 wrappers, TLS slots, APC queue, Sleep/SleepEx, cache-read waits. |
d3d9.c |
Records the presenting thread (host_present_thread) at every Present. |
native/EngineVision/Sources/EngineVisionRuntime.m |
Creates the engine thread in the app and sets host_core_telemetry. |
flowchart TB
subgraph native["Native threads"]
main["Desktop main thread / visionOS app main thread (UI, run loop)"]
eng["Engine thread: host_run, tid 1, TEB 0x7FFDE000, stack 0x00100000..0x00200000, 1.5 GiB native stack, USER_INTERACTIVE on visionOS"]
rd["Guest thread: cache reader 00443940, tid 2+, own TEB and guest stack from page space"]
wr["Guest thread: saved-game writer 00538980, tid 2+, own TEB and guest stack"]
fw["Framework threads: AudioQueue callback, GameController notifications, CompositorServices render loop"]
end
eng -->|"CreateThread shim, host_thread_create"| rd
eng -->|"CreateThread shim"| wr
eng <-->|"events, mutexes, critical sections (pthread mutex+cond)"| rd
rd -->|"ReadFileEx completions (thread-local APC ring)"| rd
rd -->|"io_signal_completion (condvar)"| eng
main -->|"halo_settings atomics, pointer queue, host_quit_requested"| eng
fw <-->|"mixer lock, settings atomics"| eng
| Thread | Guest identity | Guest stack | Native stack | QoS |
|---|---|---|---|---|
Engine thread (runs host_run) |
host_current_thread_id() == 1 (no HostThread), TEB 0x7FFDE000
|
0x00100000–0x00200000 (1 MiB), initial esp 256 bytes below the top |
1.5 GiB (main.c and enginevision_start) |
QOS_CLASS_USER_INTERACTIVE in the app (EngineVisionRuntime.m:741); desktop leaves the default |
Guest threads (CreateThread) |
tid from next_tid, starting at 2 |
stack_size rounded up to 64 KiB, minimum 64 KiB, default 1 MiB, from the page allocator |
pthread default | Creator's class by default (HALO_GUEST_THREAD_QOS) |
| Any other native thread | Treated as tid 1 if it ever calls a guest-facing function | n/a | n/a | n/a |
Both the guest stack and the translated code's native C stack are used: translated functions keep x86 stack data in guest memory at esp, but each translated call is also a nested native C call, so the engine thread runs on a very large native stack (the app refuses to start it if the 1.5 GiB request is not honoured).
Why the engine thread is USER_INTERACTIVE: the comment at EngineVisionRuntime.m:734-740 records that created without a class it "ran as an ordinary default-priority thread beside a compositor that renders at 90 Hz at user-interactive priority, eligible for the efficiency cores", and that the headset's engine ran at four frames a second in play while the Mac ran at sixteen: "the thread was not short of work to skip, it was short of a core."
sequenceDiagram
participant C as Creator (engine thread)
participant S as CreateThread shim
participant T as threading.c
participant H as host.c
participant N as new pthread (thread_start)
C->>S: CreateThread(sa, stack, start, param, flags, &tid)
S->>T: host_thread_create(stack, start, param, flags, &tid)
T->>T: calloc HostThread, tid = next_tid++, exit_code = STILL_ACTIVE (259), suspend_count = flags & CREATE_SUSPENDED
T->>H: host_initialize_guest_thread(cpu, tid, stack): page-alloc stack + TEB, set esp/fs_base/flags, setup TEB+TLS
T->>T: handle = host_handle_new(HANDLE_THREAD)
T->>N: pthread_create (with QoS attr if a class was chosen), pthread_detach
T-->>S: handle (tid written back)
N->>N: current_thread = thread, push cancel cleanup
N->>N: wait while suspend_count > 0 and not terminate_requested
alt terminated before start
N->>T: thread_mark_done(termination code)
else runs
N->>H: host_execute_guest_thread(cpu, start, param, &exit)
H->>H: host_active_cpu = cpu, push param and return token, setjmp
H->>H: engine_dispatch(start) ... returns with pc == token
H-->>N: exit code = eax (or 1 / host_exit code after failure)
N->>T: thread_mark_done(exit code), broadcast
end
Details:
-
Context setup (host.c:561-575):
guest_page_allocprovides the stack and a separate 64 KiB granule for the 4 KiB TEB.espstarts atstack_base + stack_size - 0x100,eflags = 0x202, x87 state is initialised, andsetup_thread_tebwrites the SEH chain terminator, stack bounds, self pointer, pid 4660, tid, PEB pointer, and a fresh copy of the PE's static TLS template (host.c:535-557). Guest thread stacks and TEBs are never freed; the threads are created once per process. -
Execution (host.c:577-600): the thread installs its own thread-local
host_escapejump buffer.engine_failon that thread longjmps back here, logsguest thread stopped at ..., prints a backtrace, and reports exit code 1;host_exit(code)on a non-main context calls_exit(code)immediately (process exit), because only the main context has thehost_runescape (host.c:476-482). -
ExitThread→host_thread_exitmarks the thread done with the given code and callspthread_exit; on the engine thread (noHostThread) it becomeshost_exit(code)(threading.c:411-415). -
ResumeThreaddecrements the suspend count, wakes the thread at zero, and returns the previous count (0xFFFFFFFFfor a bad handle). There is noSuspendThreadshim. -
GetExitCodeThreadreturnsSTILL_ACTIVE(259) until the thread is done. -
WaitForSingleObject(thread)waits fordone. -
Handles:
CloseHandlereleases the table slot only;HostThreadmemory is never freed (the native thread is detached and may still be running).
The two threads Halo creates, documented in the QoS comment (threading.c:230-243): the cache file's reader (00443940), which the engine thread waits on whenever something it needs this frame has not been read, and the saved-game writer (00538980).
host_thread_create chooses a QoS class from HALO_GUEST_THREAD_QOS (read at every creation) (threading.c:244-304):
HALO_GUEST_THREAD_QOS |
Class |
|---|---|
unset, empty, or inherit (default) |
qos_class_self() of the creating thread |
interactive |
QOS_CLASS_USER_INTERACTIVE |
initiated |
QOS_CLASS_USER_INITIATED |
default |
QOS_CLASS_DEFAULT |
utility |
QOS_CLASS_UTILITY |
off or anything else |
no attributes (the pre-QoS behaviour; Darwin runs it at DEFAULT) |
If pthread_attr_set_qos_class_np or the attributed pthread_create fails, the thread is created without attributes. Rationale from the source: created without attributes, Halo's two threads ran at DEFAULT, "below the USER_INTERACTIVE engine thread that waits for them and below the 90 Hz compositor, free to be left on an efficiency core behind both while the game stood still for a read". Darwin does not pass the creator's class to a thread made without attributes, whereas on Windows both threads start at the process's normal priority, the same as their creator; inheriting the creator's class reproduces that. SetThreadPriority/SetPriorityClass are accepted and ignored.
For the report, host_guest_thread_qos() returns the class applied to the most recently created guest thread (0 if it had no attributes) and host_guest_threads() the number created; both are read by EngineDiagnosticsBridge.m.
All objects are heap-allocated native structures referenced from the handle table (threading.c:14-49). None are destroyed by CloseHandle.
| Object | Native state | Win32 APIs |
|---|---|---|
HostEvent |
pthread_mutex_t lock, pthread_cond_t changed, manual_reset, signaled
|
CreateEventA, SetEvent, waits |
HostMutex |
lock, cond, owner (guest tid), recursion
|
CreateMutexA, ReleaseMutex, waits |
HostThread |
lock, cond, pthread_t, EngineCPU, tid, start, parameter, exit_code, suspend_count, done, terminate_requested
|
CreateThread and friends, waits |
HostCriticalSection |
lock, cond, owner, recursion, keyed by the guest CRITICAL_SECTION address in a global linked list |
Initialize/Enter/Leave/DeleteCriticalSection |
-
SetEventsetssignaledand broadcasts for manual-reset events, signals one waiter for auto-reset events (threading.c:99-108). - A successful wait on an auto-reset event clears
signaled(threading.c:110-132). - There is no
ResetEventorPulseEventshim.
Owned by guest thread id, recursive. A wait succeeds when the mutex is unowned or already owned by the caller, then increments recursion. ReleaseMutex by a non-owner fails (ERROR_NOT_OWNER, 288); the last release clears the owner and signals one waiter (threading.c:134-187). Abandonment (WAIT_ABANDONED) is defined in threading.h but never produced: a thread that exits holding a mutex leaves it owned.
host_critical_section_* (threading.c:417-491) look the native section up by guest address under a global lock (creating it on first use, so a section entered without being initialised still works). The guest CRITICAL_SECTION fields are mirrored so translated code that inspects them sees plausible values:
| Guest offset | Field | Value maintained |
|---|---|---|
+4 |
LockCount |
-1 when free, else recursion - 1
|
+8 |
RecursionCount |
recursion |
+12 |
OwningThread |
owner tid or 0 |
InitializeCriticalSection zeroes the first 24 bytes and sets LockCount = -1. DeleteCriticalSection clears the mirrored fields if unowned but keeps the native node in the list. Leaving a section the caller does not own is ignored. Lookup is a linear walk of the list.
Every blocking wait goes through wait_changed (threading.c:66-87):
/* On Darwin, a signal/cancel race in a shared condition variable can run
* cancellation cleanup without reacquiring its mutex. Never cancel inside
* the native wait: use bounded slices, then deliver pending cancellation
* only after unlocking. */So each wait:
- computes a 50 ms slice deadline; if the caller's own deadline is earlier, that becomes the final slice;
- disables cancellation,
pthread_cond_timedwaits for the slice, unlocks the object mutex, restores the cancel state, callspthread_testcancel()if cancellation was enabled, and relocks; - reports
ETIMEDOUTto the caller only for the final slice; a non-final slice timeout returns 0 so the caller simply rechecks its predicate.
Deadlines are absolute CLOCK_REALTIME times computed once at the start of the wait (deadline_from_now), so internal 50 ms polling never turns into a guest-visible timeout early. tests/test_thread_cancellation.c checks that a 125 ms wait on an unsignalled event returns WAIT_TIMEOUT after at least 120 ms.
Wait results (threading.h): WAIT_OBJECT_0 (0), WAIT_TIMEOUT (0x102), WAIT_FAILED (0xFFFFFFFF, for a handle that is not an event, mutex or thread). A zero timeout is a pure poll and never sleeps.
TerminateThread → host_thread_terminate (threading.c:317-333): fails if the thread is already done; otherwise sets terminate_requested and the exit code, broadcasts (so a suspended thread wakes and exits without entering the guest), and calls pthread_cancel while still holding the thread's lock, so completion cannot publish done and retire the detached native thread between the done check and pthread_cancel's use of its ID. The cancelled thread runs thread_cancel_cleanup, which marks it done; because terminate_requested is set, the exit code stays the one passed to TerminateThread.
Cancellation is delivered only at explicit points: inside wait_changed after unlocking, in host_thread_yield (pthread_testcancel then sched_yield), and at system calls that are POSIX cancellation points. The cache-read wait in shims_kernel32.c disables cancellation around its own condition wait.
tests/test_thread_cancellation.c keeps native threads joinable (it redefines pthread_detach) so it can prove that, over 40 cycles, cancellation of a suspended thread, an infinite event wait, a timed event wait, a mutex wait, a timed thread wait and a critical-section wait all complete, pthread_join returns, the object locks are left unlocked and reusable, a terminated suspended thread never entered the guest, ExitThread reports its code, a thread spinning in host_thread_yield can be terminated, and terminating an already finished thread fails.
These live in shims_kernel32.c and are described in detail on Win32 Compatibility Layer:
-
TLS (details): 63 usable slots stored in each thread's TEB at
+0xE10, with per-slot generations so a newly allocated index reads 0 on every thread without cross-thread TEB writes. -
APCs (details): a 64-entry thread-local ring filled by
ReadFileExand drained, oldest first, by alertableSleepEx/WaitForSingleObjectExon the same thread, which then returnWAIT_IO_COMPLETION. -
Cache-read waits (details): the engine thread's
Sleep(0)at00443EEA/004446D8blocks up to 1 ms on a condition variable thatdeliver_apcssignals after each completion (HALO_IO_WAIT=0restores the spin). -
Sleep (details):
Sleep(0)returns immediately except for a 250 µs give-up every 256 calls; presenter idle is accounted for the bearing budget.
The thread-local host_last_error, host_active_cpu, host_current_import, the call ring, the TLS generation cache, the APC ring and the USER32 message queue are all _Thread_local, so each native thread has its own.
When the core telemetry is on, every wait with a nonzero timeout records its elapsed time and count in the thread-local host_thread_wait_ns and host_thread_waits, and critical-section entry records only contended acquisitions (threading.c:367-400, threading.c:454-471). The engine thread samples its own totals at each Present (engine_thread_sample in EngineVisionRuntime.m) into the frame-split report; see Diagnostics and Telemetry.
| Rule | Reason / where enforced |
|---|---|
host_core_telemetry is 0 by default; the app sets it once before the engine runs: on unless HALO_CORE_TELEMETRY=0 (EngineVisionRuntime.m:657-659). The desktop host leaves it 0. |
"While it is 0 no wait reads a clock or counts, exactly as before the telemetry." |
| Zero-timeout polls are never timed. | They do not wait. |
| Non-polling waits are timed even if they return immediately. | Counts elapsed API time, not exact blocked CPU time. |
| Uncontended or recursive critical-section entries are not timed. | Keeps the clock off the fast path. |
Timestamps preserve errno. |
wait_clock_ns saves and restores errno so the instrumented call returns exactly the errno the underlying wait left (the test notes Darwin can leave 0x13C after a timed-out slice). |
| Read with relaxed atomics. | Written before any guest thread exists. |
tests/test_thread_wait_telemetry.c wraps the clock to deliberately clobber errno and verifies: polls read no clock; a failed non-polling wait and an immediately satisfied one each count once with two clock reads; a 20 ms timeout records at least 10 ms; blocked event and critical-section waiters on another thread record one wait and two clock reads in their own thread-local counters without affecting the main thread's; errno survives in both directions; and with telemetry off the same waits read no clock and count nothing.
| Lock | Protects | Notes |
|---|---|---|
heap_lock |
guest heap metadata | Never held while escaping (heap_fail unlocks before host_exit). |
page_lock |
page allocator bitmap | |
handles_lock |
handle table | Held only for the table lookup. |
procs_lock |
import table | Released before calling the shim. |
per-object lock
|
event/mutex/thread/critical-section state | Waited on with wait_changed. |
critical_sections_lock |
critical-section list | Held only for lookup/creation. |
tls_lock |
TLS allocation bitmap and slot access | |
io_lock |
cache-wait completion counter | "Only ever held around the counter, never while taking another lock." |
Documents master-chef at commit 9f915af (v1.0.3). Unofficial project, not affiliated with Microsoft, Bungie, Gearbox or Apple. Original code is MIT licensed; game content is not included.
Overview
- Architecture Overview
- Repository Layout
- Glossary
- Environment Variables
- Contributing Guide
- Open Questions
Translation
- Static Translation Pipeline
- XWA Decoder and Lifter
- Function Address Lists
- EngineReuse Runtime
- x87 Floating Point
Host runtime
- EngineHost Overview
- Win32 Compatibility Layer
- Threading and Synchronization
- Guest Memory and Heap
- Engine Overrides and Hooks
- Runtime Settings
Graphics
- Direct3D9 Bridge
- Metal Renderer
- Shader Translation
- Textures and Texture Packs
- Geometry Fast Paths
- Radial Fog
Panorama and presentation
- Panorama System
- Panorama Budget and LOD
- Frame Pacing
- visionOS App
- Immersive Presenter
- Layer Alignment
Audio and input
Tooling and process