Skip to content

Threading and Synchronization

M T edited this page Oct 4, 2026 · 1 revision

Threading and Synchronization

Halo PC is a multi-threaded Windows program: besides the main game thread it creates a cache-file reader thread and a saved-game writer thread, and it synchronizes them with events, mutexes, critical sections and asynchronous file I/O. In the port, every guest thread is a real native pthread that runs translated code against its own EngineCPU, guest stack and guest TEB, and every Win32 synchronization object is a native mutex/condition-variable pair behind a handle. threading.c implements the objects and threads; shims_kernel32.c exposes them as Win32 APIs and adds TLS, APC delivery and Sleep; host.c provides the per-thread guest context. This page describes the thread model, each object type, waits, cancellation, quality of service, and the wait telemetry the device report consumes.

Source files

File Role
threading.h Public API: wait result codes, event/mutex/thread/critical-section functions, host_thread_wait_ns/host_thread_waits, host_core_telemetry, QoS report accessors.
threading.c HostEvent, HostMutex, HostThread, HostCriticalSection; cancellation-safe waits; QoS selection; wait telemetry.
host.c host_initialize_guest_thread (guest stack + TEB + TLS), host_execute_guest_thread (run with a per-thread setjmp escape), thread-local host_active_cpu, host_exit.
shims_kernel32.c Win32 wrappers, TLS slots, APC queue, Sleep/SleepEx, cache-read waits.
d3d9.c Records the presenting thread (host_present_thread) at every Present.
native/EngineVision/Sources/EngineVisionRuntime.m Creates the engine thread in the app and sets host_core_telemetry.

Thread model

flowchart TB
    subgraph native["Native threads"]
        main["Desktop main thread / visionOS app main thread (UI, run loop)"]
        eng["Engine thread: host_run, tid 1, TEB 0x7FFDE000, stack 0x00100000..0x00200000, 1.5 GiB native stack, USER_INTERACTIVE on visionOS"]
        rd["Guest thread: cache reader 00443940, tid 2+, own TEB and guest stack from page space"]
        wr["Guest thread: saved-game writer 00538980, tid 2+, own TEB and guest stack"]
        fw["Framework threads: AudioQueue callback, GameController notifications, CompositorServices render loop"]
    end
    eng -->|"CreateThread shim, host_thread_create"| rd
    eng -->|"CreateThread shim"| wr
    eng <-->|"events, mutexes, critical sections (pthread mutex+cond)"| rd
    rd -->|"ReadFileEx completions (thread-local APC ring)"| rd
    rd -->|"io_signal_completion (condvar)"| eng
    main -->|"halo_settings atomics, pointer queue, host_quit_requested"| eng
    fw <-->|"mixer lock, settings atomics"| eng
Loading
Thread Guest identity Guest stack Native stack QoS
Engine thread (runs host_run) host_current_thread_id() == 1 (no HostThread), TEB 0x7FFDE000 0x00100000–0x00200000 (1 MiB), initial esp 256 bytes below the top 1.5 GiB (main.c and enginevision_start) QOS_CLASS_USER_INTERACTIVE in the app (EngineVisionRuntime.m:741); desktop leaves the default
Guest threads (CreateThread) tid from next_tid, starting at 2 stack_size rounded up to 64 KiB, minimum 64 KiB, default 1 MiB, from the page allocator pthread default Creator's class by default (HALO_GUEST_THREAD_QOS)
Any other native thread Treated as tid 1 if it ever calls a guest-facing function n/a n/a n/a

Both the guest stack and the translated code's native C stack are used: translated functions keep x86 stack data in guest memory at esp, but each translated call is also a nested native C call, so the engine thread runs on a very large native stack (the app refuses to start it if the 1.5 GiB request is not honoured).

Why the engine thread is USER_INTERACTIVE: the comment at EngineVisionRuntime.m:734-740 records that created without a class it "ran as an ordinary default-priority thread beside a compositor that renders at 90 Hz at user-interactive priority, eligible for the efficiency cores", and that the headset's engine ran at four frames a second in play while the Mac ran at sixteen: "the thread was not short of work to skip, it was short of a core."

Guest thread lifecycle

sequenceDiagram
    participant C as Creator (engine thread)
    participant S as CreateThread shim
    participant T as threading.c
    participant H as host.c
    participant N as new pthread (thread_start)
    C->>S: CreateThread(sa, stack, start, param, flags, &tid)
    S->>T: host_thread_create(stack, start, param, flags, &tid)
    T->>T: calloc HostThread, tid = next_tid++, exit_code = STILL_ACTIVE (259), suspend_count = flags & CREATE_SUSPENDED
    T->>H: host_initialize_guest_thread(cpu, tid, stack): page-alloc stack + TEB, set esp/fs_base/flags, setup TEB+TLS
    T->>T: handle = host_handle_new(HANDLE_THREAD)
    T->>N: pthread_create (with QoS attr if a class was chosen), pthread_detach
    T-->>S: handle (tid written back)
    N->>N: current_thread = thread, push cancel cleanup
    N->>N: wait while suspend_count > 0 and not terminate_requested
    alt terminated before start
        N->>T: thread_mark_done(termination code)
    else runs
        N->>H: host_execute_guest_thread(cpu, start, param, &exit)
        H->>H: host_active_cpu = cpu, push param and return token, setjmp
        H->>H: engine_dispatch(start) ... returns with pc == token
        H-->>N: exit code = eax (or 1 / host_exit code after failure)
        N->>T: thread_mark_done(exit code), broadcast
    end
Loading

Details:

  • Context setup (host.c:561-575): guest_page_alloc provides the stack and a separate 64 KiB granule for the 4 KiB TEB. esp starts at stack_base + stack_size - 0x100, eflags = 0x202, x87 state is initialised, and setup_thread_teb writes the SEH chain terminator, stack bounds, self pointer, pid 4660, tid, PEB pointer, and a fresh copy of the PE's static TLS template (host.c:535-557). Guest thread stacks and TEBs are never freed; the threads are created once per process.
  • Execution (host.c:577-600): the thread installs its own thread-local host_escape jump buffer. engine_fail on that thread longjmps back here, logs guest thread stopped at ..., prints a backtrace, and reports exit code 1; host_exit(code) on a non-main context calls _exit(code) immediately (process exit), because only the main context has the host_run escape (host.c:476-482).
  • ExitThread → host_thread_exit marks the thread done with the given code and calls pthread_exit; on the engine thread (no HostThread) it becomes host_exit(code) (threading.c:411-415).
  • ResumeThread decrements the suspend count, wakes the thread at zero, and returns the previous count (0xFFFFFFFF for a bad handle). There is no SuspendThread shim.
  • GetExitCodeThread returns STILL_ACTIVE (259) until the thread is done.
  • WaitForSingleObject(thread) waits for done.
  • Handles: CloseHandle releases the table slot only; HostThread memory is never freed (the native thread is detached and may still be running).

The two threads Halo creates, documented in the QoS comment (threading.c:230-243): the cache file's reader (00443940), which the engine thread waits on whenever something it needs this frame has not been read, and the saved-game writer (00538980).

Quality of service for guest threads

host_thread_create chooses a QoS class from HALO_GUEST_THREAD_QOS (read at every creation) (threading.c:244-304):

HALO_GUEST_THREAD_QOS Class
unset, empty, or inherit (default) qos_class_self() of the creating thread
interactive QOS_CLASS_USER_INTERACTIVE
initiated QOS_CLASS_USER_INITIATED
default QOS_CLASS_DEFAULT
utility QOS_CLASS_UTILITY
off or anything else no attributes (the pre-QoS behaviour; Darwin runs it at DEFAULT)

If pthread_attr_set_qos_class_np or the attributed pthread_create fails, the thread is created without attributes. Rationale from the source: created without attributes, Halo's two threads ran at DEFAULT, "below the USER_INTERACTIVE engine thread that waits for them and below the 90 Hz compositor, free to be left on an efficiency core behind both while the game stood still for a read". Darwin does not pass the creator's class to a thread made without attributes, whereas on Windows both threads start at the process's normal priority, the same as their creator; inheriting the creator's class reproduces that. SetThreadPriority/SetPriorityClass are accepted and ignored.

For the report, host_guest_thread_qos() returns the class applied to the most recently created guest thread (0 if it had no attributes) and host_guest_threads() the number created; both are read by EngineDiagnosticsBridge.m.

Synchronization objects

All objects are heap-allocated native structures referenced from the handle table (threading.c:14-49). None are destroyed by CloseHandle.

Object Native state Win32 APIs
HostEvent pthread_mutex_t lock, pthread_cond_t changed, manual_reset, signaled CreateEventA, SetEvent, waits
HostMutex lock, cond, owner (guest tid), recursion CreateMutexA, ReleaseMutex, waits
HostThread lock, cond, pthread_t, EngineCPU, tid, start, parameter, exit_code, suspend_count, done, terminate_requested CreateThread and friends, waits
HostCriticalSection lock, cond, owner, recursion, keyed by the guest CRITICAL_SECTION address in a global linked list Initialize/Enter/Leave/DeleteCriticalSection

Events

  • SetEvent sets signaled and broadcasts for manual-reset events, signals one waiter for auto-reset events (threading.c:99-108).
  • A successful wait on an auto-reset event clears signaled (threading.c:110-132).
  • There is no ResetEvent or PulseEvent shim.

Mutexes

Owned by guest thread id, recursive. A wait succeeds when the mutex is unowned or already owned by the caller, then increments recursion. ReleaseMutex by a non-owner fails (ERROR_NOT_OWNER, 288); the last release clears the owner and signals one waiter (threading.c:134-187). Abandonment (WAIT_ABANDONED) is defined in threading.h but never produced: a thread that exits holding a mutex leaves it owned.

Critical sections

host_critical_section_* (threading.c:417-491) look the native section up by guest address under a global lock (creating it on first use, so a section entered without being initialised still works). The guest CRITICAL_SECTION fields are mirrored so translated code that inspects them sees plausible values:

Guest offset Field Value maintained
+4 LockCount -1 when free, else recursion - 1
+8 RecursionCount recursion
+12 OwningThread owner tid or 0

InitializeCriticalSection zeroes the first 24 bytes and sets LockCount = -1. DeleteCriticalSection clears the mirrored fields if unowned but keeps the native node in the list. Leaving a section the caller does not own is ignored. Lookup is a linear walk of the list.

Waits and cancellation

Every blocking wait goes through wait_changed (threading.c:66-87):

/* On Darwin, a signal/cancel race in a shared condition variable can run
 * cancellation cleanup without reacquiring its mutex. Never cancel inside
 * the native wait: use bounded slices, then deliver pending cancellation
 * only after unlocking. */

So each wait:

  1. computes a 50 ms slice deadline; if the caller's own deadline is earlier, that becomes the final slice;
  2. disables cancellation, pthread_cond_timedwaits for the slice, unlocks the object mutex, restores the cancel state, calls pthread_testcancel() if cancellation was enabled, and relocks;
  3. reports ETIMEDOUT to the caller only for the final slice; a non-final slice timeout returns 0 so the caller simply rechecks its predicate.

Deadlines are absolute CLOCK_REALTIME times computed once at the start of the wait (deadline_from_now), so internal 50 ms polling never turns into a guest-visible timeout early. tests/test_thread_cancellation.c checks that a 125 ms wait on an unsignalled event returns WAIT_TIMEOUT after at least 120 ms.

Wait results (threading.h): WAIT_OBJECT_0 (0), WAIT_TIMEOUT (0x102), WAIT_FAILED (0xFFFFFFFF, for a handle that is not an event, mutex or thread). A zero timeout is a pure poll and never sleeps.

Termination

TerminateThread → host_thread_terminate (threading.c:317-333): fails if the thread is already done; otherwise sets terminate_requested and the exit code, broadcasts (so a suspended thread wakes and exits without entering the guest), and calls pthread_cancel while still holding the thread's lock, so completion cannot publish done and retire the detached native thread between the done check and pthread_cancel's use of its ID. The cancelled thread runs thread_cancel_cleanup, which marks it done; because terminate_requested is set, the exit code stays the one passed to TerminateThread.

Cancellation is delivered only at explicit points: inside wait_changed after unlocking, in host_thread_yield (pthread_testcancel then sched_yield), and at system calls that are POSIX cancellation points. The cache-read wait in shims_kernel32.c disables cancellation around its own condition wait.

tests/test_thread_cancellation.c keeps native threads joinable (it redefines pthread_detach) so it can prove that, over 40 cycles, cancellation of a suspended thread, an infinite event wait, a timed event wait, a mutex wait, a timed thread wait and a critical-section wait all complete, pthread_join returns, the object locks are left unlocked and reusable, a terminated suspended thread never entered the guest, ExitThread reports its code, a thread spinning in host_thread_yield can be terminated, and terminating an already finished thread fails.

Thread-local storage, APCs and Sleep

These live in shims_kernel32.c and are described in detail on Win32 Compatibility Layer:

  • TLS (details): 63 usable slots stored in each thread's TEB at +0xE10, with per-slot generations so a newly allocated index reads 0 on every thread without cross-thread TEB writes.
  • APCs (details): a 64-entry thread-local ring filled by ReadFileEx and drained, oldest first, by alertable SleepEx/WaitForSingleObjectEx on the same thread, which then return WAIT_IO_COMPLETION.
  • Cache-read waits (details): the engine thread's Sleep(0) at 00443EEA/004446D8 blocks up to 1 ms on a condition variable that deliver_apcs signals after each completion (HALO_IO_WAIT=0 restores the spin).
  • Sleep (details): Sleep(0) returns immediately except for a 250 µs give-up every 256 calls; presenter idle is accounted for the bearing budget.

The thread-local host_last_error, host_active_cpu, host_current_import, the call ring, the TLS generation cache, the APC ring and the USER32 message queue are all _Thread_local, so each native thread has its own.

Wait telemetry

When the core telemetry is on, every wait with a nonzero timeout records its elapsed time and count in the thread-local host_thread_wait_ns and host_thread_waits, and critical-section entry records only contended acquisitions (threading.c:367-400, threading.c:454-471). The engine thread samples its own totals at each Present (engine_thread_sample in EngineVisionRuntime.m) into the frame-split report; see Diagnostics and Telemetry.

Rule Reason / where enforced
host_core_telemetry is 0 by default; the app sets it once before the engine runs: on unless HALO_CORE_TELEMETRY=0 (EngineVisionRuntime.m:657-659). The desktop host leaves it 0. "While it is 0 no wait reads a clock or counts, exactly as before the telemetry."
Zero-timeout polls are never timed. They do not wait.
Non-polling waits are timed even if they return immediately. Counts elapsed API time, not exact blocked CPU time.
Uncontended or recursive critical-section entries are not timed. Keeps the clock off the fast path.
Timestamps preserve errno. wait_clock_ns saves and restores errno so the instrumented call returns exactly the errno the underlying wait left (the test notes Darwin can leave 0x13C after a timed-out slice).
Read with relaxed atomics. Written before any guest thread exists.

tests/test_thread_wait_telemetry.c wraps the clock to deliberately clobber errno and verifies: polls read no clock; a failed non-polling wait and an immediately satisfied one each count once with two clock reads; a 20 ms timeout records at least 10 ms; blocked event and critical-section waiters on another thread record one wait and two clock reads in their own thread-local counters without affecting the main thread's; errno survives in both directions; and with telemetry off the same waits read no clock and count nothing.

Locking summary

Lock Protects Notes
heap_lock guest heap metadata Never held while escaping (heap_fail unlocks before host_exit).
page_lock page allocator bitmap
handles_lock handle table Held only for the table lookup.
procs_lock import table Released before calling the shim.
per-object lock event/mutex/thread/critical-section state Waited on with wait_changed.
critical_sections_lock critical-section list Held only for lookup/creation.
tls_lock TLS allocation bitmap and slot access
io_lock cache-wait completion counter "Only ever held around the counter, never while taking another lock."

Related pages

Clone this wiki locally