Repository navigation
Transcript Correction
Your STT engine has never seen "Bett-Lightstrip". Murdock has: the Home Assistant registry tells it exactly which names exist. Three mechanisms use that knowledge, in this order.
| Stage | When | Works with |
|---|---|---|
| 1. Bias prompt | before transcription | OpenAI-compatible backends only |
| 2. Correction dictionary | after transcription, exact rules | every backend |
| 3. Automatic name correction | after transcription, fuzzy | every backend |
All of them are off by default and configured under Settings → Transcription → Transcript quality.
Terms are handed to the engine before it transcribes, so it can recognise custom names at the source. That only works where the API has a prompt field:
| Backend | Prompt sent? |
|---|---|
| Wyoming | ✗ no prompt field exists |
| voxtral | ✗ Mistral doesn't document one |
| openrouter | ✗ its request shape has no place for it |
| openai / groq / speaches | ✓ |
Before 0.8.3 this was the main use of the mirrored vocabulary, which means on a default install it did nothing at all. The Web UI now says so explicitly when the main service ignores the prompt.
The list is capped at 25 mirrored terms, because a long list makes the engine hallucinate those names into noisy audio. Your own terms are never capped.
Settings → Transcription → Transcript quality shows every mirrored term as a chip:
- Red chips are your own terms — always sent. Add one with the input field, remove it with the ×.
- Blue chips are mirrored from Home Assistant and clickable: filled = sent, outlined = not sent.
- Touching any blue chip switches from automatic (first 25 by priority) to hand-picked. Automatic selection returns to the default.
The counter shows how many are in play, and the collapsible block shows the exact prompt that goes out.
One rule per line:
fehlende Lichter -> Fehenlichter # replaces
Bad-Lightstrip ~> Bett-Lightstrip # annotates as an alternative
-> is a plain replacement and keeps HA's local intent matching working.
~> keeps what was heard and adds the alternative as a hint (see
Speaker Context).
Matching is case-insensitive on word boundaries, longer phrases win, #
starts a comment.
Rules always win over the automatic correction below — they run first.
The one that needs no rules and works with every backend. It maps spans of the transcript onto names Home Assistant actually has:
"schalte das Bad-Lightstrip ein" → "schalte das Bett-Lightstrip ein"
-
Candidates come from the full mirrored vocabulary — uncapped, since
a local index has no length budget — plus your own terms. Indexed by
Kölner Phonetik, a German phonetic code. That matters: "Bad" and
"Bett" both encode to
12, which is why the engine confuses them, and pure edit distance ranks them far apart. - Scoring blends sequence similarity, character-trigram overlap and a bonus when the phonetic codes match exactly.
- Two ways in. Either the score clears Minimum similarity (default 0.82), or — when the phonetic codes are identical — a bar 0.10 lower, because sounding the same is strong evidence of a mishearing rather than a coincidence.
- The margin decides ties. The winner must lead the runner-up by Minimum lead (default 0.10). If "Bett-Licht" and "Bad-Licht" are both almost equally close, picking one is a coin flip — so nothing is replaced. Same reasoning as the speaker margin gate.
Phonetics may propose, never decide: "Bad" and "Bett" sound identical, and only the entity list knows which one you own.
- text that already matches a known name — nothing to fix
- spans shorter than 4 characters
- common German words (there's a stop list), so "mach das Licht an" stays as it is
- anything an explicit dictionary rule already handled
The result is plain text. Unlike an inline [oder: …] marker it
doesn't break HA's local intent matcher — it helps it, because
"Bad-Lightstrip" matches no entity while "Bett-Lightstrip" does. And
because it runs on Murdock's side after transcription, it is independent
of which STT engine you use.
Every applied correction is counted. Under Recurring corrections the
frequent ones appear with a Make it a rule button, which writes an exact
-> entry into the dictionary and enables that tier.
Worth doing: an exact rule is faster, fully deterministic, and the fuzzy matcher never has to decide it again. It is deliberately a click and not an automatism.
Start with the defaults. If corrections you want aren't happening, lower Minimum similarity in 0.02 steps and watch the recognition log. If something gets replaced that shouldn't, raise Minimum lead first — that targets ambiguity specifically, rather than making everything stricter.