Repository navigation
Tuning
All knobs live in the Web UI's Settings tab and take effect live — no restarts. Since 0.8.2 the tab is split into four groups (Recognition, Transcription, Home Assistant, System); everything on this page lives under Recognition, with the rarely-touched knobs folded into Advanced blocks inside that card.
A transcript you disagree with has two possible causes — the audio was already bad, or the engine misread good audio — and the text alone cannot tell you which. The recognition log therefore keeps two recordings of the newest utterances:
- Microphone — the untouched capture, exactly as it arrived.
- Sent to the service — the conditioned copy, present only when Condition audio before upload changed anything.
Play them against the transcript. If the recording is clear and the transcript is not, the engine is the problem: compare a shadow service's reading of the same utterance, which the log shows underneath. If the recording itself is quiet, clipped or full of room noise, no engine will save it — move the satellite, or switch on the upload conditioning so a quiet utterance reaches the service at a usable level.
How many utterances keep their audio is set in the log (20 by default, roughly 200 kB each; 0 stores nothing).
Every session's AudioStop line in the add-on log says how the audio arrived:
AudioStop from HA — 2.25s audio in 8.29s stream (27%), 225 chunks,
longest gap 1840ms, slowest chunk 0.4ms, loop lag 12ms
Far below 100 % means the stream ran behind real time, and Home Assistant's end-of-speech detection — which runs on the same stream — ran behind with it: long listening phases, turns cut at the 15-second limit. The other figures say whose fault it was:
-
loop lag / slowest chunk high (hundreds of ms): Murdock's event
loop was busy and stopped reading. Look for a
Event loop held … by …warning just before; it names what held it. - both low, longest gap high: Murdock read every chunk the moment it came — the audio itself arrived late. Home Assistant was busy, or the satellite's network link is weak; compare the satellites, since the recognition log records the figures per utterance and marks late audio with a badge.
Below 80 % the add-on log also prints a warning naming the side.
The core decision: maximum cosine distance for a match (0 = identical, default 0.30). Lower = stricter (more false rejects), higher = looser (more false accepts).
Don't tune by feel — use "Suggest from log" next to the threshold field. It analyses the recognition log: distances of accepted matches (your voice) vs. best distances of blocked/unknown utterances (strangers, TV), and proposes the midpoint of the gap between the genuine 95th percentile and the impostor 5th percentile. If the distributions overlap it says so — that usually means a speaker needs more/better samples, not a different threshold.
A noisy kitchen needs a looser threshold than a quiet study. Once satellite identification is set up, every satellite appears under Per-satellite thresholds with its own override. Two satellites in the same room are tuned separately (the entity id, not the room, is the key).
The threshold answers "is this voice close enough to Jonas?" — it does not ask "is it closer to Jonas than to anyone else?" With two similar-sounding household members, a distance under the threshold can be a coin flip.
The margin gate (default 0, i.e. off) adds that second question:
the gap between the best and the second-best speaker must reach a
minimum, otherwise the result is unsicher instead of a match.
best: Jonas d=0.31
second: Nils d=0.35 margin = 0.04
gate = 0.10 → uncertain (not a match)
Effects of an uncertain result:
- the transcript is still forwarded (or blocked, following
require_speaker_match— same as an unknown voice) - the recognition log records outcome
uncertain-forwarded/blocked-uncertaintogether with the measured margin - the event fires with
speaker: null,uncertain: trueandreason: "uncertain"; the integration turns that intoSprecher: unsicher - the sample is not filed in the unknown postbox — it's a near-match of an enrolled voice, not a stranger
- the early-match shortcut is skipped, so the full verify decides
Start around 0.05–0.10 if two people in your household get confused;
leave it at 0 if everyone sounds clearly different. Look at
all_distances in the recognition log to see how much room you actually
have. Per-satellite overrides exist via the API
(PATCH /api/settings/satellite-margin-gates).
When an utterance contains more than one voice (you + TV, or two
people), a single whole-clip embedding blends them and drifts toward
"unknown". Extraction (default on) splits the utterance into speech
regions, scores each against the enrolled speakers at the stricter
extraction threshold (default 0.25), keeps only the dominant
speaker's regions, and verifies that clean concatenation at the normal
threshold. Single-region utterances skip all of it — no added latency.
The confidence reported to HA is a calibrated probability (Platt scaling), fitted automatically from your enrollments: genuine pairs are each sample vs. its speaker's leave-one-out centroid, impostor pairs are cross-speaker distances. It refits in the background on enrollment changes. Gating still uses the distance threshold — calibration only makes the reported confidence meaningful (and is the groundwork for context fusion).
While a TV/radio plays in the active satellite's room, Murdock tightens the threshold so TV voices don't sneak through:
- Set up the media context automation.
- By default, any playing source in the satellite's room tightens by the global boost (0.05).
-
Media restrictions (per satellite × source) refines that: per
(satellite, media player) pair, set how much that source tightens —
e.g. living-room TV → living-room satellite
0.15, bedroom radio →0(no effect). The strongest playing source wins; an explicit0disables a source for that satellite.
There's also a spectral liveness gate (min liveness score, default 0.35) that rejects obviously-played-back audio before embedding.
The composite sample-quality score (speech ratio, SNR, liveness, embedding consistency, centroid fit) drives auto-enroll's smart replacement. The component weights are tunable under Sample quality (advanced) — defaults are sensible; touch only if your environment skews one component (e.g. a permanently noisy room tanking SNR).
Three settings, in the order they are worth reaching for.
Cancel phrases (Recognition → advanced). Saying one of them drops the turn — for a wake word that fired on a phone call or the television. The phrase has to lead the utterance, so "kein Abbruch nötig" is still a normal request, and a mangled "Abruch" is matched phonetically. Where the upstream sends interim results it is caught mid-sentence, before the command it was never meant to carry gets acted on. Empty disables it.
Give up after silence (Recognition → advanced, default 3 s). Ends a turn in which nobody said anything. It only triggers when essentially no speech was detected, so drawing breath before speaking never costs the turn.
Liveness bar while media plays (Recognition → advanced, default +0.15). A playing TV is exactly when a marginal liveness score is most likely to be the TV, so the bar rises with the room.
Neither of the first two can shorten the recording — Home Assistant's pipeline decides when to stop sending audio, and its own silence detection is the knob for that. What they skip is the transcript, the conversation agent and the spoken reply, which is the part that makes a phantom trigger feel like a malfunction.
A wake word that fires while you are talking to somebody else is real, live, human speech. No acoustic measure separates "a person is speaking" from "a person is speaking, but not to you" — the liveness heuristic in particular scores a good loudspeaker highly, because reproducing those properties is exactly what a good loudspeaker does. Cancel phrases are the honest answer there today.