Repository navigation
Architecture
Murdock registers as a Wyoming ASR provider. HA streams satellite audio to it; Murdock forwards the stream to the real STT engine in parallel with its own speaker verification, so recognition adds almost no latency on top of transcription.
AudioStart ─► open upstream lazily, forward chunks live
AudioChunk ─► buffer + forward; after ~1.5 s: early-probe embedding
AudioStop ─► gates (below) ─► Transcript (or "") back to the satellite
- Too short (< 1 s) → passthrough (no verification possible).
- No speakers enrolled → capture to the unknown postbox for training; passthrough or block (configurable).
- Liveness — spectral heuristic rejects obvious playback/TV noise.
- Adaptive extraction — multi-voice utterances are reduced to the dominant enrolled speaker's regions.
- Verify — CAM++ embedding vs. sqlite-vec centroids (cosine distance, threshold per satellite, tightened while media plays).
- Match → transcript forwarded, result published (MQTT + REST), optional auto-enroll. No match → empty transcript (if require-match), audio logged to the unknown postbox.
An early probe at ~1.5 s of audio can confirm a high-confidence match before the utterance even ends, shaving seconds off the perceived latency of speaker-aware automations.
| Module | Role |
|---|---|
wyoming_murdock/handler.py |
the Wyoming session: gates, upstream streaming, publishing |
murdock/core/embeddings.py |
CAM++ ONNX (192-dim, CPU, ~20 ms) |
murdock/core/fbank.py |
Kaldi-compatible log-mel features |
murdock/core/vad.py |
Silero VAD (segments, enrollment QC) |
murdock/core/extraction.py |
adaptive target-speaker region extraction |
murdock/core/speaker_store.py |
enrollment, verification, centroids, health |
murdock/core/calibration.py |
Platt scaling distance → P(same speaker) |
murdock/core/embedding_map.py |
PCA projection for the voice map |
murdock/core/unknown_store.py |
unknown postbox + TTL cleanup |
murdock/core/unknown_cluster.py |
greedy cosine clustering of unknowns |
murdock/core/liveness.py |
spectral playback/TV heuristic |
murdock/core/mqtt_integration.py |
discovery out, context in (media, satellite, presence) |
murdock/core/ha_integration.py |
legacy REST push |
murdock/core/recognition_log.py |
audit log (drives stats + threshold recommendation) |
murdock/api/ |
FastAPI backend for the Web UI |
murdock/ui/static/ |
vanilla HTML/CSS/JS frontend, DE/EN |
A single SQLite file (murdock.db) with the sqlite-vec extension
for SIMD-accelerated KNN over the speaker centroids. Sample audio is
stored as WAV blobs; per-sample embeddings are recomputed on demand
(centroid rebuilds, health, voice map) so they always reflect the
current embedder. Settings live in a key/value table — which is why the
backup can capture the whole deployment.
- Proxy stays proxy — no native HA integration; gating happens inline in the audio path. Context flows in/out over MQTT.
- Never worse — extraction, calibration and context only refine decisions; every failure path falls back to the plain verify gate.
- CPU-first — CAM++ (7 M params) + Silero run comfortably on a NAS; no GPU anywhere.