Skip to content
Bob McGlobus edited this page Oct 11, 2026 · 6 revisions

Tuning recognition

All knobs live in the Web UI's Settings tab and take effect live — no restarts. Since 0.8.2 the tab is split into four groups (Recognition, Transcription, Home Assistant, System); everything on this page lives under Recognition, with the rarely-touched knobs folded into Advanced blocks inside that card.

Listening to what was actually said

A transcript you disagree with has two possible causes — the audio was already bad, or the engine misread good audio — and the text alone cannot tell you which. The recognition log therefore keeps two recordings of the newest utterances:

  • Microphone — the untouched capture, exactly as it arrived.
  • Sent to the service — the conditioned copy, present only when Condition audio before upload changed anything.

Play them against the transcript. If the recording is clear and the transcript is not, the engine is the problem: compare a shadow service's reading of the same utterance, which the log shows underneath. If the recording itself is quiet, clipped or full of room noise, no engine will save it — move the satellite, or switch on the upload conditioning so a quiet utterance reaches the service at a usable level.

How many utterances keep their audio is set in the log (20 by default, roughly 200 kB each; 0 stores nothing).

Audio arriving slower than real time

Every session's AudioStop line in the add-on log says how the audio arrived:

AudioStop from HA — 2.25s audio in 8.29s stream (27%), 225 chunks,
longest gap 1840ms, slowest chunk 0.4ms, loop lag 12ms

Far below 100 % means the stream ran behind real time, and Home Assistant's end-of-speech detection — which runs on the same stream — ran behind with it: long listening phases, turns cut at the 15-second limit. The other figures say whose fault it was:

  • loop lag / slowest chunk high (hundreds of ms): Murdock's event loop was busy and stopped reading. Look for a Event loop held … by … warning just before; it names what held it.
  • both low, longest gap high: Murdock read every chunk the moment it came — the audio itself arrived late. Home Assistant was busy, or the satellite's network link is weak; compare the satellites, since the recognition log records the figures per utterance and marks late audio with a badge.

Below 80 % the add-on log also prints a warning naming the side.

Verify threshold

The core decision: maximum cosine distance for a match (0 = identical, default 0.30). Lower = stricter (more false rejects), higher = looser (more false accepts).

Don't tune by feel — use "Suggest from log" next to the threshold field. It analyses the recognition log: distances of accepted matches (your voice) vs. best distances of blocked/unknown utterances (strangers, TV), and proposes the midpoint of the gap between the genuine 95th percentile and the impostor 5th percentile. If the distributions overlap it says so — that usually means a speaker needs more/better samples, not a different threshold.

Per-satellite overrides

A noisy kitchen needs a looser threshold than a quiet study. Once satellite identification is set up, every satellite appears under Per-satellite thresholds with its own override. Two satellites in the same room are tuned separately (the entity id, not the room, is the key).

Margin gate

The threshold answers "is this voice close enough to Jonas?" — it does not ask "is it closer to Jonas than to anyone else?" With two similar-sounding household members, a distance under the threshold can be a coin flip.

The margin gate (default 0, i.e. off) adds that second question: the gap between the best and the second-best speaker must reach a minimum, otherwise the result is unsicher instead of a match.

best:   Jonas   d=0.31
second: Nils    d=0.35     margin = 0.04
gate = 0.10                → uncertain (not a match)

Effects of an uncertain result:

  • the transcript is still forwarded (or blocked, following require_speaker_match — same as an unknown voice)
  • the recognition log records outcome uncertain-forwarded / blocked-uncertain together with the measured margin
  • the event fires with speaker: null, uncertain: true and reason: "uncertain"; the integration turns that into Sprecher: unsicher
  • the sample is not filed in the unknown postbox — it's a near-match of an enrolled voice, not a stranger
  • the early-match shortcut is skipped, so the full verify decides

Start around 0.05–0.10 if two people in your household get confused; leave it at 0 if everyone sounds clearly different. Look at all_distances in the recognition log to see how much room you actually have. Per-satellite overrides exist via the API (PATCH /api/settings/satellite-margin-gates).

Adaptive speaker extraction

When an utterance contains more than one voice (you + TV, or two people), a single whole-clip embedding blends them and drifts toward "unknown". Extraction (default on) splits the utterance into speech regions, scores each against the enrolled speakers at the stricter extraction threshold (default 0.25), keeps only the dominant speaker's regions, and verifies that clean concatenation at the normal threshold. Single-region utterances skip all of it — no added latency.

Confidence calibration

The confidence reported to HA is a calibrated probability (Platt scaling), fitted automatically from your enrollments: genuine pairs are each sample vs. its speaker's leave-one-out centroid, impostor pairs are cross-speaker distances. It refits in the background on enrollment changes. Gating still uses the distance threshold — calibration only makes the reported confidence meaningful (and is the groundwork for context fusion).

Media-aware gating

While a TV/radio plays in the active satellite's room, Murdock tightens the threshold so TV voices don't sneak through:

  1. Set up the media context automation.
  2. By default, any playing source in the satellite's room tightens by the global boost (0.05).
  3. Media restrictions (per satellite × source) refines that: per (satellite, media player) pair, set how much that source tightens — e.g. living-room TV → living-room satellite 0.15, bedroom radio → 0 (no effect). The strongest playing source wins; an explicit 0 disables a source for that satellite.

There's also a spectral liveness gate (min liveness score, default 0.35) that rejects obviously-played-back audio before embedding.

Quality-score weights

The composite sample-quality score (speech ratio, SNR, liveness, embedding consistency, centroid fit) drives auto-enroll's smart replacement. The component weights are tunable under Sample quality (advanced) — defaults are sensible; touch only if your environment skews one component (e.g. a permanently noisy room tanking SNR).

Turning down false wakes

Three settings, in the order they are worth reaching for.

Cancel phrases (Recognition → advanced). Saying one of them drops the turn — for a wake word that fired on a phone call or the television. The phrase has to lead the utterance, so "kein Abbruch nötig" is still a normal request, and a mangled "Abruch" is matched phonetically. Where the upstream sends interim results it is caught mid-sentence, before the command it was never meant to carry gets acted on. Empty disables it.

Give up after silence (Recognition → advanced, default 3 s). Ends a turn in which nobody said anything. It only triggers when essentially no speech was detected, so drawing breath before speaking never costs the turn.

Liveness bar while media plays (Recognition → advanced, default +0.15). A playing TV is exactly when a marginal liveness score is most likely to be the TV, so the bar rises with the room.

Neither of the first two can shorten the recording — Home Assistant's pipeline decides when to stop sending audio, and its own silence detection is the knob for that. What they skip is the transcript, the conversation agent and the spoken reply, which is the part that makes a phantom trigger feel like a malfunction.

What none of this fixes

A wake word that fires while you are talking to somebody else is real, live, human speech. No acoustic measure separates "a person is speaking" from "a person is speaking, but not to you" — the liveness heuristic in particular scores a good loudspeaker highly, because reproducing those properties is exactly what a good loudspeaker does. Cancel phrases are the honest answer there today.

Clone this wiki locally