Skip to content

Architecture

Bob McGlobus edited this page Jun 11, 2026 · 2 revisions

Architecture

How Murdock works

The pipeline

Murdock registers as a Wyoming ASR provider. HA streams satellite audio to it; Murdock forwards the stream to the real STT engine in parallel with its own speaker verification, so recognition adds almost no latency on top of transcription.

AudioStart ─► open upstream lazily, forward chunks live
AudioChunk ─► buffer + forward; after ~1.5 s: early-probe embedding
AudioStop  ─► gates (below) ─► Transcript (or "") back to the satellite

The gates, in order

  1. Too short (< 1 s) → passthrough (no verification possible).
  2. No speakers enrolled → capture to the unknown postbox for training; passthrough or block (configurable).
  3. Liveness — spectral heuristic rejects obvious playback/TV noise.
  4. Adaptive extraction — multi-voice utterances are reduced to the dominant enrolled speaker's regions.
  5. Verify — CAM++ embedding vs. sqlite-vec centroids (cosine distance, threshold per satellite, tightened while media plays).
  6. Match → transcript forwarded, result published (MQTT + REST), optional auto-enroll. No match → empty transcript (if require-match), audio logged to the unknown postbox.

An early probe at ~1.5 s of audio can confirm a high-confidence match before the utterance even ends, shaving seconds off the perceived latency of speaker-aware automations.

Modules

Module Role
wyoming_murdock/handler.py the Wyoming session: gates, upstream streaming, publishing
murdock/core/embeddings.py CAM++ ONNX (192-dim, CPU, ~20 ms)
murdock/core/fbank.py Kaldi-compatible log-mel features
murdock/core/vad.py Silero VAD (segments, enrollment QC)
murdock/core/extraction.py adaptive target-speaker region extraction
murdock/core/speaker_store.py enrollment, verification, centroids, health
murdock/core/calibration.py Platt scaling distance → P(same speaker)
murdock/core/embedding_map.py PCA projection for the voice map
murdock/core/unknown_store.py unknown postbox + TTL cleanup
murdock/core/unknown_cluster.py greedy cosine clustering of unknowns
murdock/core/liveness.py spectral playback/TV heuristic
murdock/core/mqtt_integration.py discovery out, context in (media, satellite, presence)
murdock/core/ha_integration.py legacy REST push
murdock/core/recognition_log.py audit log (drives stats + threshold recommendation)
murdock/api/ FastAPI backend for the Web UI
murdock/ui/static/ vanilla HTML/CSS/JS frontend, DE/EN

Storage

A single SQLite file (murdock.db) with the sqlite-vec extension for SIMD-accelerated KNN over the speaker centroids. Sample audio is stored as WAV blobs; per-sample embeddings are recomputed on demand (centroid rebuilds, health, voice map) so they always reflect the current embedder. Settings live in a key/value table — which is why the backup can capture the whole deployment.

Design principles

  • Proxy stays proxy — no native HA integration; gating happens inline in the audio path. Context flows in/out over MQTT.
  • Never worse — extraction, calibration and context only refine decisions; every failure path falls back to the plain verify gate.
  • CPU-first — CAM++ (7 M params) + Silero run comfortably on a NAS; no GPU anywhere.

Clone this wiki locally