Skip to content

Training

Bob McGlobus edited this page Sep 23, 2026 · 3 revisions

Training speakers

Murdock needs 3–5 good samples per person to recognise reliably. There are three ways to get them — all end up in the same profile.

1. Over the voice satellite (recommended, works in the add-on)

No microphone, no files — just talk:

  1. Optionally pre-create the speaker: Speakers → Create speaker (no samples).
  2. Speak 5–10 normal commands to your satellite. With no (or no matching) speaker, each utterance is captured for training.
  3. Assign them:
    • Recognition log → blocked entries carry their recording and an Add to speaker button, or
    • Voice map → click a gray cross → Assign.
  4. Done — the centroid builds automatically. Lower the verify threshold back to ~0.30 if you raised it for bootstrapping (you don't need to: capture works regardless of the threshold).

2. File upload

Speakers → Enroll speaker accepts WAV, MP3, MP4/M4A, OGG, FLAC and WebM. Each sample gets a quality score (speech ratio, SNR, liveness, consistency, centroid fit) so you can tell training-worthy recordings from junk.

3. Browser recording

Works when the UI is served over HTTPS or http://localhost — i.e. in docker-compose setups. Inside the HA add-on, ingress is plain HTTP and browsers block getUserMedia there; the record button is disabled with a hint (use methods 1 or 2).

Auto-enroll (profile aging)

With Auto-enroll on (default), every high-confidence match adds a fresh embedding to the profile, replacing the lowest-quality auto sample once the cap (20) is reached. Your profile follows your voice over time — morning voice, colds, aging microphones.

Keeping profiles healthy

  • Speakers → Health (per speaker) shows each sample's drift from the centroid, its age and quality, plus a quality trend. Samples flagged drifted sit far from the profile center — listen and delete bad takes; the centroid rebuilds automatically.
  • Voice map (Speakers tab) renders the whole embedding space in 2-D: tight, separated clusters = reliable recognition. Click any point to play / delete / assign it.
  • After larger changes, run Recalibrate now (Settings → Confidence calibration) — or just let the automatic background refit handle it.

Profile health

The Speakers tab has a Profile health panel that turns the data Murdock already has into concrete advice: too few samples, poor recording quality, a satellite the speaker uses often but has no samples from (so no per-satellite profile can be built), matches passing only just, or a profile drifting worse over time.

An empty panel is the normal state — it only speaks up when there is something specific to fix.

Clone this wiki locally