Skip to content

Experimental

Bob McGlobus edited this page Aug 12, 2026 · 5 revisions

Experimental features

Unproven or unfinished features live in their own tab, away from the knobs you tune daily. They are safe to leave off and safe to switch on — but expect rough edges, and don't build automations you rely on around them yet.

Whisper detection

Flags utterances that were whispered, so your conversation agent can answer quietly.

Whispering is acoustically unambiguous in a way most voice properties are not: the vocal folds don't vibrate, so there is no fundamental frequency and no harmonic structure — just filtered noise. The detector scores

  • harmonicity (is there periodicity at all — by far the strongest signal),
  • zero-crossing rate (noise crosses zero far more often),
  • spectral tilt (whispering is comparatively flat and bright).

Loudness only refines the decision, never drives it: a distant speaker is quiet too, and a whisper close to the microphone is not.

Two things worth knowing

It disables two gates for that utterance. Whispered speech is quiet and spectrally flat, which is exactly what the TV/playback liveness heuristic and the early reject throw away. Both are skipped when a whisper is detected — without that, Murdock would discard precisely the utterances this feature exists to notice.

A whisper usually comes back as unbekannt. Speaker embeddings are built from voiced speech; removing the pitch and shifting the formants moves the embedding a long way, so verification fails. That is left as-is on purpose: relaxing the threshold because "it's only a whisper" would turn the feature into a way past the gate. Whisper is reported, not trusted.

Recognising who whispers is possible — see Recognising who whispers below — but it needs a whisper profile per speaker. Without one, a whispered command is reported as whispered and attributed to nobody.

Reading the score

The recognition log shows the measured score, not just a verdict:

Shown Means
geflüstert 0.87 (highlighted) over the threshold — treated as a whisper
Flüstern 0.41 (muted) measured, but under the bar

The muted case is the useful one for tuning: it tells you how far a real whisper actually landed from the threshold. Unknown samples carry the score too, so a quiet rejected clip is explainable.

Default threshold 0.62; higher is stricter.

A noisy room used to fake it

Until 0.8.9 a running fan or air conditioner could push ordinary speech over the bar — one install saw 0.78 on a normally spoken sentence. The cause was frame selection, not acoustics: the silence floor was a fixed level that room noise sits above, so the noise frames counted as speech and, having no periodicity, dragged the voicing measure down.

The activity bar now rises with the room's own noise floor, and the voicing measure is weighted by energy rather than counting frames, so a long tail of barely-audible frames cannot outvote a few loud voiced syllables. The same voice reads the same in any room.

If a score still looks wrong, the clip is in the recognition log with its number — that is the evidence to tune against.

Recognising who whispers

Whispered speech moves the embedding far from a voiceprint built from normal speech, so by default a whisper comes back as unbekannt. Enroll a whisper profile to fix that:

Speakers → Enroll → speaking style Geflüstert → record two or three whispered samples.

  • Whispered samples never enter the normal voiceprint or the per-satellite profiles. Averaging in a voice with no pitch would make ordinary recognition worse for everybody.
  • The whisper centroid is consulted only when the detector says the utterance was whispered, and the threshold is unchanged. Someone who never enrolled a whisper stays unknown while whispering — which is what keeps this a recognition feature rather than a way past the gate.
  • Two samples are enough (whispering varies less than normal speech); below that no profile is built.

The samples list badges every whispered sample with geflüstert, so a profile stays maintainable: delete the whispered ones and the whisper profile disappears with them.

Assigning from the postbox

Clips assigned from Unknown or the recognition log are filed under the style they were spoken in, not the style you happen to be looking at: the clip's own whisper score decides. A whispered clip assigned to Jonas lands in Jonas's whisper profile and leaves his normal voiceprint alone.

That is the behaviour you want by default — without it, assigning a few whispered clips would quietly degrade ordinary recognition. The assignment dialog still lets you override it to Normale Sprache or Geflüstert when the detector got it wrong.

Telling Home Assistant

The whisper state leaves Murdock on every path, so an automation can duck the TTS volume without parsing JSON:

Path What you get
MQTT binary_sensor.murdock_whispering and sensor.murdock_whisper_score
Integration binary_sensor.murdock_<satellite>_whisper (attribute whisper_score)
Event / REST whisper and whisper_score in the recognition payload
Conversation agent an extra prompt line via the Murdock LLM API

The prompt line tells the agent to answer quietly and briefly — and explicitly not to read whispering as an instruction to keep something secret, which is the failure mode worth pre-empting.

Whispering is reported even when the speaker is unbekannt. Answering more quietly does not depend on knowing who asked.

Emotion detection

Classifies the emotional tone of verified speech and pushes it to Home Assistant.

The model is not bundled — it is ~356 MB and most installs will never want it. Fetch it explicitly:

DOWNLOAD_EMOTION_MODEL=1 bash scripts/download_models.sh ./models

Two files are downloaded, and both are required:

File What it is
emotion.onnx emotion2vec+ base — a feature extractor, emits frame-level 768-dim features
emotion_head.bin the 9-class linear head (~27 KB)

The ONNX alone cannot name an emotion; classification is mean-pooling the frames and applying that head. If the head is missing the download script deletes the ONNX too, and the classifier refuses to run rather than guessing — an early version invented label names like class_44737 and reported them with full confidence, which is worse than reporting nothing.

Labels: angry, disgusted, fearful, happy, neutral, sad, surprised, plus other and unknown, which are filtered out before anything reaches Home Assistant.

Accuracy on German is unverified. The model is multilingual by training, but nobody has measured it on a German household. The label lands in the recognition log — judge it there before building anything on it.

Multiple speakers per utterance

Not a toggle: when speaker extraction is enabled and more than one enrolled voice is present in a clip, the recognition event carries a speakers list with each name and how long they spoke, longest first.

"speakers": [
  {"speaker": "Jonas", "seconds": 2.4},
  {"speaker": "Anna",  "seconds": 0.9}
]

The gate still follows the dominant speaker — this is "who else was in the room", not a change to who the command is attributed to.

Extraction already scored every speech region against the enrolled speakers in order to pick that dominant voice, so this costs nothing extra; the information used to be discarded.

Clone this wiki locally