-
Notifications
You must be signed in to change notification settings - Fork 3
kithara test utils
Documentation reviewed from source revision 19ca073f2. This records the documented contract at that revision; it is not a new runtime validation. API and usage · All crates.
#[kithara::test(serial)] expansions acquire the hidden test-binary-wide async
mutex before entering the generated body. The guard is deliberately outside the
per-test timeout: serialization controls overlap, while each timeout measures
only its own test body.
While at least one Recorder is alive, the capture layer records every tracing::event! whose target ends with _probe (all #[kithara::probe] expansions emit to <crate_name>_probe, e.g. kithara_stream_probe) into a process-wide Vec, so a test can snapshot the full sequence and assert on it.
-
Why a tracing layer, not
EventBus:kithara_events::EventBusis atokio::sync::broadcast— under load, lagged subscribers drop events. Probes fire at the decision site and the tracing layer records every emission without a bounded channel. -
Why a process-wide subscriber:
tracing::subscriber::set_defaultis thread-local, but probes fire on tokio worker threads (e.g. those spawned byDownloader::run) that do not inherit a per-test default. Because#[kithara::test]initialises a global subscriber viasetup_tracing_with_filter, the probe layer must be composed inside that init path —test::init_tracingattaches it alongside the fmt layer. A separateset_global_defaultwould fail withSetGlobalDefault. -
Activation: probe sites compile to no-ops unless the crate's
probefeature is enabled in the test build. Thecapturemodule itself is gated oncfg(any(test, feature = "probe"))and is absent onwasm32. - Under
--cfg rtsaninit_tracinginstalls no subscriber at all: a capturing/formatting subscriber allocates on the forbid-blocking audio worker, so that lane deliberately has no probe capture.
Isolation is by install id, not by serializing tests:
-
#[kithara::test]callsbump_install_id()once and enters theOWNED_INSTALL_IDtask-local scope before the test body runs. - Every probe firing stamps
current_install_id()into its event;Recorder::snapshotkeeps only events whoseinstall_idmatches the recorder's and whose timestamp is>= start_at. - The task-local is what makes this correct:
tokio::spawninherits it, so orphan tasks from a just-finished test (downloader on-complete, audio worker draining its last buffer) freeze the previous id and drop out of the next test's snapshot.spawn_blockingand non-tokio threads do not inherit it and fall back to the global atomic. - Probe capture is lease-bound. Each
install()acquires a lease, cloned recorders share that lease, and independently installed recorders share the global log. Dropping the last lease clears the log and releases its backing allocation; dropping one overlapping recorder leaves the log intact for its live siblings. The lease is what bounds the log's lifetime: the layer is composed into every test binary's subscriber, HLS playback fires ~10k probes/second, and a retainedProbeEventcosts ~700 bytes — unleased capture added ~70 MB of RSS per playback session and never gave it back.
Recorder::wait_for_probe / wait_for_probe_async are the sanctioned way to advance a test: they block until a recorded event matches the predicate or the budget elapses, including events that arrived before the call. Tests should use probe arrival as their clock instead of polling Audio::read / Stream::len() on wall time, and should fail when the budget elapses rather than relaxing it. Use the async variant on a current_thread runtime — the blocking one starves the tasks the test is waiting on.
Decode packed probe arguments with T::from_probe_arg(event.u64("field")?) rather than hand-written decoders next to the IntoProbeArg impls.
The native hang implementation owns the versioned kithara.hang.v1 artifact.
Each JSON envelope carries its exact nextest attempt identity, diagnostic and
last-progress context, plus an optional Flash wait snapshot. Publication is
no-clobber: a unique temporary file is flushed before it is linked to its final
name, so an interrupted write cannot replace earlier evidence.
After a successful publish, stderr contains only the artifact path. If publish
fails, stderr retains a bounded 64 KiB envelope excerpt for degraded diagnosis;
the reporter still treats the missing durable envelope as incomplete evidence.
The encoded envelope is always below 4 MiB. Each nextest value and the label is
bounded to 8 KiB, the diagnostic to 32 KiB, the context payload to 192 KiB, and
the Flash snapshot to 256 KiB before JSON escaping. Oversized context and Flash
values retain UTF-8-safe head and tail excerpts with an explicit omitted-byte
marker. Canonical nextest identity remains byte-for-byte exact within its 8 KiB
field budget; an oversized value is visibly marked rather than silently joined
to another attempt. A final size guard can discard context and Flash, but never
the captured nextest identity.
If a context producer returns malformed JSON, the envelope preserves those
bytes as a JSON string. This is an intentional degraded diagnostic: losing the
typed context must not prevent publication of the attempt identity and Flash
snapshot that can still explain the hang.
PreKillGuard is installed only for native #[kithara::test] expansions that
do not declare their own timeout. When KITHARA_HANG_PREKILL_SECS is set, its
real-clock worker records evidence shortly before the outer nextest deadline;
dropping the guard cancels and joins that worker. Explicit wall-clock, hard and
sync timeout paths call the same artifact owner directly before panic or abort.
The hook records evidence but never terminates the process.
The hook answers only panics whose payload is a &str or a String. Every
panic this workspace raises carries a message; a payload that is neither is a
dependency unwinding for control flow rather than failing, which is how loom
cancels each suspended generator at the end of every execution. Building a dump
for such an unwind reaches state that execution owned, so it is skipped.
The flight-recorder rings are std::sync::Mutex, not the platform one. They
are process-global, outlive every loom::model execution and are read from the
panic hook; a loom-modelled lock panics the moment anything touches it after
the execution that created it ended, which under a panic hook aborts the
process. Ring poisoning is ignored: a dump that drops its tail because another
thread died mid-record loses exactly the evidence it was taken for.