Repository navigation
Releases: BobMcGlobus/Murdock
Release list
Murdock 0.12.0 — Aufnahmen im Erkennungsprotokoll
The recognition log keeps the audio. Every logged utterance — blocked
ones included — now keeps two recordings for as long as it is among the
newest few: what the microphone sent, and the copy the speech-to-text
service received. They play inline in the log and can be downloaded. A
wrong transcript has two possible causes, bad audio or a misreading
engine, and they are indistinguishable from the text alone; comparing
the recording against the transcript settles it.
How many utterances keep their audio is set in the log itself (20 by
default, roughly 200 kB each, 0 to store nothing). Older utterances drop
their recordings, and clearing the log clears them too.
The Unknown tab is gone. The recognition log filters by outcome, has
the same "assign to speaker" action, and now carries the audio as well.
Voice clustering and bulk assignment went with the tab.
Murdock 0.11.1 — keine abgerissenen Transkripte mehr
Transcripts came back with holes and missing endings ("Schalte alle
Lich-") since 0.11. Home Assistant hands Murdock audio in chunks of about
10 ms, and 0.11 looked the main STT service up in the database on every
single one — around 700 SQLite queries per second of speech. On a small
VM that pushed the handler below real time; the add-on log showed 2.1 s
of audio arriving over a 5 s stream. The audio that could not be handed
over in time was lost before it reached Murdock, which also made normal
speech look whispered.
The settings the audio path needs are now read once when a session
starts, and the STT services are held in memory. A regression test
feeds 300 chunks through the real handler and fails on any query.
The AudioStop log line now reads e.g. 2.13s audio in 5.00s stream (43%), so a stream that loses audio is visible at a glance.
Murdock 0.11.0 — STT-Dienste mit Main, Fallback und Shadows
Speech-to-text engines are now services. Every engine — a Wyoming
server, an OpenAI-compatible API, Voxtral, a Home Assistant STT entity —
is set up once under Settings → Transcription with its own name,
address, key, model and optional timeout. Switching the engine that
answers is one click; nothing has to be retyped.
Each service takes roles:
- Main — exactly one; it answers Home Assistant.
- Fallbacks — an ordered chain, tried when the main fails. When a
service that needs the internet cannot even be reached, the remaining
internet fallbacks are skipped and the chain goes straight to the
local ones, so the offline fallback does not wait behind every dead
cloud timeout. Ask the fallbacks when the main heard nothing (on by
default) sends an empty answer down the chain as well; this replaces
"let the shadow answer when the primary heard nothing". - Shadows — any number, for benchmarking. They run only after the
answer has gone out, strictly one at a time, and wait while another
utterance is being transcribed, so they never slow the answer or skew
each other's timings. The recognition log lists every shadow's
transcript, time or error under the utterance.
A Test button checks a service with what is in the form: Wyoming
servers and Home Assistant are asked what they support (and whether your
language is among it); HTTP APIs get one second of silence, which is a
normal, billed request.
The dual transcript is removed. Existing settings are converted once
on first start: the backend becomes the main, the local fallback a
Wyoming fallback, the A/B shadow a shadow (and also a fallback if the
rescue was on). A stored-but-empty upstream address still means the
add-on's upstream_uri, as it always did. The add-on's STT options now
only seed the first service. Backups carry the services, keys included,
in stt_services.json.
Murdock 0.10.4 — Speichern-Knopf zurück, Transkripte vollständig
The transcription form had no Save button
So the speech-to-text backend could not be changed at all.
Folding the A/B shadow section away in 0.10.0 wrote its </details> at the wrong nesting level: instead of closing at the end of that section, it closed thousands of characters later — inside the MQTT card. The opening tag therefore swallowed the rest of the transcription form, an entire other form, and part of the next card, taking the only submit button with it. The tag counts balanced, which is why nothing complained.
Two tests now cover this: one fails on a submit button inside a collapsed <details>, and one on any <details> still open where a form ends — which is the shape a mis-nested closing tag actually takes.
The primary transcript was still being cut short
0.9.4 treated the first transcript after AudioStop as final. 0.10.3 found the same assumption in the one-shot helper behind the shadow. Both were still too generous: a streaming recogniser has the tail of the utterance in its buffer when the audio stops, and carries on emitting longer transcripts after AudioStop. Taking the first of those produced "Fahre alle Rollladen der gesam".
The newest transcript now wins, with the end of the stream inferred from a short silence. That wait is only spent on a server that has already shown itself to be streaming, by sending interim results while audio was still flowing — a one-shot engine such as faster-whisper never pays it.
Murdock 0.10.3 — Shadow-Transkripte nicht mehr abschneiden
The A/B shadow was cutting Wyoming transcripts short.
0.9.4 fixed this for the streaming reader in the handler. The same defect lived in the other place a Wyoming server is spoken to: the one-shot helper behind the A/B shadow and the local fallback, which returned the first Transcript it saw. A streaming recogniser sends one per firmed-up word, so the shadow column read "wann Avengers Doomsday rauskomm" where the cloud had "rauskommt?".
The marker used in the handler — that AudioStop has gone out — cannot discriminate here, because this helper always sends everything before it starts reading. And the protocol has no final flag to wait for. So the newest transcript is kept and the end of the stream is inferred from a short silence: long enough not to cut a server off mid-sentence, short enough to stay invisible on the fallback path. transcript-stop ends it immediately where a server sends one, and transcript-chunk is ignored outright, being partial by definition.
Only the shadow and the local fallback were affected. The primary upstream path has been correct since 0.9.4.
Murdock 0.10.2 — Testknopf für HA-STT-Entitäten
Telling you why the Home Assistant STT entity did not work.
Three things have to line up — the entity id, Home Assistant's URL and token, and a language the entity actually supports — and Home Assistant answers a wrong combination with a bare 415 or 404 and no explanation, which makes "it doesn't work" unfalsifiable.
A Test entity button now asks Home Assistant directly over GET /api/stt/{entity} and reports which part is missing, along with the entity's own supported languages and whether the configured one is among them. It compares on the primary subtag, since Home Assistant lists locales like de-DE while Murdock stores a bare de.
Like the upstream ping, it tests the value currently in the field rather than the saved one.
Murdock 0.10.1 — Threading-Fix und HA-Cloud-STT
The integration was writing entity state from a worker thread
Reported from a live install: Home Assistant logged ten warnings in a hundred minutes about a state write off the event loop — which it warns can crash it or corrupt data — plus twelve exceptions inside one second from a single dispatch.
The cause was a missing decorator. Home Assistant classifies a dispatcher target through HassJob: a coroutine runs on the loop, a function marked @callback runs on the loop, and anything else is treated as an executor job and run in a worker thread. All three _handle_update handlers were plain functions, so every speaker update wrote state from the wrong thread, once per entity — which is the twelve-in-a-second burst exactly.
They now carry @callback and check the entity is still attached before writing, since a dispatch can land while one is being torn down. A test parses the integration and fails on any dispatch handler that is not a callback; it parses rather than imports, because the integration needs Home Assistant and the test image deliberately does not carry it.
Integration version 0.3.1 — update it via HACS alongside the add-on.
Home Assistant Cloud transcription is now usable
Cloud STT is not a Wyoming service and has no API of its own, so it could not be reached the way the other backends are. But it is an ordinary stt.* entity, and Home Assistant exposes every such entity over POST /api/stt/{entity_id} — the endpoint its own frontend uses.
A new Home Assistant STT entity backend speaks it, which also opens up every other STT entity in the instance: a Whisper add-on, anything an integration provides. It needs the Home Assistant URL and token already configured under Home Assistant.
Pointing it at Murdock's own entity is refused — the request would come straight back here.
Murdock 0.10.0 — die Oberfläche, neu geschnitten
The interface, rearranged around what you came to do.
A new Overview opens the app, with four counts and a setup checklist ordered by dependency rather than importance — a transcription backend is useless without a speaker to recognise, and a speaker is useless if nothing reaches Home Assistant. Somebody working down the list never hits a step that cannot succeed yet, and a thin profile is named rather than left to guess at. The checklist disappears once every step is ticked.
The recognition log states what happened in a sentence — "Jonas was closest, but not close enough (0.32 against 0.30)" — and keeps d=, th=, verify= and the request breakdown behind a collapsed Details. Satellites read as their room name, with the entity id on hover.
One detail area per speaker instead of five buttons competing for the row. Opening it goes straight to the samples, because that is what people come for.
Collapsible settings cards work again. They targeted direct children of the settings tab, and the cards moved inside group containers some releases ago, so the selector had quietly stopped matching anything and every card rendered expanded. The first card of each group is open now and the rest are folded, per group rather than per page. The A/B shadow engine — nine fields you configure once and never look at again — is folded away too.
Two labels that were never there
uncertain-forwarded had been rendering as its own raw key in the log, and the outcomes added in 0.9.5 had never been given any. A test now derives the expected set from the OUTCOME_ constants and checks that both locales define the same keys — which immediately turned up a German string written into the English block, overwriting it.
Also in this release
Cancel phrases and a silence abort: a turn nobody started ends by itself, and a phrase like "Abbruch" drops one that should never have begun. Neither can shorten the recording — Home Assistant decides that — but both skip the transcript, the conversation agent and the spoken reply.
Murdock 0.9.5 — Fehlauslöser, UI-Politur, Stimmlage statt Emotion
A polish release.
Rejecting what wasn't a person
The liveness bar now rises with the room. Media tightening existed, but it only ever moved the verify threshold — which decides who spoke, not whether anyone did. That was backwards: a playing TV is exactly the moment a marginal liveness score is most likely to be the TV.
Cancel phrases drop a turn the wake word should never have started — a phone call, the television. Matching is deliberately narrow: the phrase must lead the utterance, so "kein Abbruch nötig" stays a normal request, and phonetic comparison catches a mangled "Abruch". Where the upstream sends interim results it is caught mid-sentence, before the command it was never meant to carry gets acted on. Off until configured.
Reading the interface
Satellites can be named — the existing active-satellite automation may send a name, persisted so a room stays named across restarts.
The 2-D map is cached. It re-embedded every stored sample on every view, which is why opening it felt like running a fresh enrolment. Keyed on a fingerprint of the enrolment data rather than a manual invalidation hook, so no future code path can forget to clear it.
Speakers show their whisper profile — whether they have a whisper voiceprint, or how many whispered samples are still missing.
Emotion out, voice style in
Emotion detection cost 706 ms per utterance against a total answer time of ~340 ms, needed a 356 MB opt-in model, and classified a mood rather than anything an automation could act on. It is removed.
Voice style replaces it and costs nothing: the whisper detector already measures the level, so quiet / normal / animated / whispered comes free. It is judged against a running median of that satellite's own recent utterances — "loud" has no absolute meaning across microphones and rooms — so a satellite calibrates itself within a handful of turns and needs no configuration. Published as sensor.murdock_voice_style.
The emotion database columns stay: dropping one in SQLite means rebuilding the table, and an old database is still readable with them present.
Murdock 0.9.4 — Streaming-Transkripte nicht mehr abschneiden
A streaming upstream had its transcripts cut off mid-word.
"Was ist die Hauptstadt von Niedersach". "Und wann findet die Maker-Fähr in Han". "Wir müssen nochmal einen Latenzte". Every utterance truncated at the end, while the cloud shadow on the same audio returned the full sentence.
The reader took the first Transcript event from upstream and hung up. A one-shot engine sends exactly one, after AudioStop, so that behaviour was indistinguishable from correct. A streaming recogniser sends one every time it firms up another word — and the first of those is the beginning of the sentence, captured while the user is still speaking.
A Transcript now counts as final only once AudioStop has gone upstream. Earlier ones are kept as the best-so-far and used when a session ends without a final result, so a dropped connection degrades to a partial answer rather than to none — which is also why an empty final no longer erases a good interim.
Relevant to anyone running Kroko, sherpa-onnx streaming, or any other Wyoming server that emits partial results.