Repository navigation
Experimental
Unproven or unfinished features live in their own tab, away from the knobs you tune daily. They are safe to leave off and safe to switch on — but expect rough edges, and don't build automations you rely on around them yet.
Flags utterances that were whispered, so your conversation agent can answer quietly.
Whispering is acoustically unambiguous in a way most voice properties are not: the vocal folds don't vibrate, so there is no fundamental frequency and no harmonic structure — just filtered noise. The detector scores
- harmonicity (is there periodicity at all — by far the strongest signal),
- zero-crossing rate (noise crosses zero far more often),
- spectral tilt (whispering is comparatively flat and bright).
Loudness only refines the decision, never drives it: a distant speaker is quiet too, and a whisper close to the microphone is not.
It disables two gates for that utterance. Whispered speech is quiet and spectrally flat, which is exactly what the TV/playback liveness heuristic and the early reject throw away. Both are skipped when a whisper is detected — without that, Murdock would discard precisely the utterances this feature exists to notice.
A whisper usually comes back as unbekannt. Speaker embeddings are
built from voiced speech; removing the pitch and shifting the formants
moves the embedding a long way, so verification fails. That is left
as-is on purpose: relaxing the threshold because "it's only a whisper"
would turn the feature into a way past the gate. Whisper is reported,
not trusted.
Recognising who whispers is possible — see Recognising who whispers below — but it needs a whisper profile per speaker. Without one, a whispered command is reported as whispered and attributed to nobody.
The recognition log shows the measured score, not just a verdict:
| Shown | Means |
|---|---|
geflüstert 0.87 (highlighted) |
over the threshold — treated as a whisper |
Flüstern 0.41 (muted) |
measured, but under the bar |
The muted case is the useful one for tuning: it tells you how far a real whisper actually landed from the threshold. Unknown samples carry the score too, so a quiet rejected clip is explainable.
Default threshold 0.62; higher is stricter.
Until 0.8.9 a running fan or air conditioner could push ordinary speech over the bar — one install saw 0.78 on a normally spoken sentence. The cause was frame selection, not acoustics: the silence floor was a fixed level that room noise sits above, so the noise frames counted as speech and, having no periodicity, dragged the voicing measure down.
The activity bar now rises with the room's own noise floor, and the voicing measure is weighted by energy rather than counting frames, so a long tail of barely-audible frames cannot outvote a few loud voiced syllables. The same voice reads the same in any room.
If a score still looks wrong, the clip is in the recognition log with its number — that is the evidence to tune against.
Whispered speech moves the embedding far from a voiceprint built from
normal speech, so by default a whisper comes back as unbekannt. Enroll
a whisper profile to fix that:
Speakers → Enroll → speaking style Geflüstert → record two or three whispered samples.
- Whispered samples never enter the normal voiceprint or the per-satellite profiles. Averaging in a voice with no pitch would make ordinary recognition worse for everybody.
- The whisper centroid is consulted only when the detector says the utterance was whispered, and the threshold is unchanged. Someone who never enrolled a whisper stays unknown while whispering — which is what keeps this a recognition feature rather than a way past the gate.
- Two samples are enough (whispering varies less than normal speech); below that no profile is built.
The samples list badges every whispered sample with geflüstert, so a profile stays maintainable: delete the whispered ones and the whisper profile disappears with them.
Clips assigned from Unknown or the recognition log are filed under the style they were spoken in, not the style you happen to be looking at: the clip's own whisper score decides. A whispered clip assigned to Jonas lands in Jonas's whisper profile and leaves his normal voiceprint alone.
That is the behaviour you want by default — without it, assigning a few whispered clips would quietly degrade ordinary recognition. The assignment dialog still lets you override it to Normale Sprache or Geflüstert when the detector got it wrong.
The whisper state leaves Murdock on every path, so an automation can duck the TTS volume without parsing JSON:
| Path | What you get |
|---|---|
| MQTT |
binary_sensor.murdock_whispering and sensor.murdock_whisper_score
|
| Integration |
binary_sensor.murdock_<satellite>_whisper (attribute whisper_score) |
| Event / REST |
whisper and whisper_score in the recognition payload |
| Conversation agent | an extra prompt line via the Murdock LLM API |
The prompt line tells the agent to answer quietly and briefly — and explicitly not to read whispering as an instruction to keep something secret, which is the failure mode worth pre-empting.
Whispering is reported even when the speaker is unbekannt. Answering
more quietly does not depend on knowing who asked.
Classifies the emotional tone of verified speech and pushes it to Home Assistant.
The model is not bundled — it is ~356 MB and most installs will never want it. Fetch it explicitly:
DOWNLOAD_EMOTION_MODEL=1 bash scripts/download_models.sh ./modelsTwo files are downloaded, and both are required:
| File | What it is |
|---|---|
emotion.onnx |
emotion2vec+ base — a feature extractor, emits frame-level 768-dim features |
emotion_head.bin |
the 9-class linear head (~27 KB) |
The ONNX alone cannot name an emotion; classification is mean-pooling the
frames and applying that head. If the head is missing the download script
deletes the ONNX too, and the classifier refuses to run rather than
guessing — an early version invented label names like class_44737 and
reported them with full confidence, which is worse than reporting nothing.
Labels: angry, disgusted, fearful, happy, neutral, sad,
surprised, plus other and unknown, which are filtered out before
anything reaches Home Assistant.
Accuracy on German is unverified. The model is multilingual by training, but nobody has measured it on a German household. The label lands in the recognition log — judge it there before building anything on it.
Not a toggle: when speaker extraction is enabled and more than one
enrolled voice is present in a clip, the recognition event carries a
speakers list with each name and how long they spoke, longest first.
"speakers": [
{"speaker": "Jonas", "seconds": 2.4},
{"speaker": "Anna", "seconds": 0.9}
]The gate still follows the dominant speaker — this is "who else was in the room", not a change to who the command is attributed to.
Extraction already scored every speech region against the enrolled speakers in order to pick that dominant voice, so this costs nothing extra; the information used to be discarded.