Repository navigation
Releases: moazzamak/dsh-voice-input
Release list
v0.6.1 — the stylesheet has an owner
The harness client loader owns plugin styles by the data-plugin attribute:
claimStyles marks every <style> that LACKS it as belonging to whichever plugin
materialises next, and removeOwnedStyles(id) deletes every <style> whose
data-plugin equals an id when that entry is replaced or pruned. This sheet
carried only a private data-dsh-voice-input marker, so it looked unowned to the
loader: another plugin took ownership of it, and that plugin's first refresh or
prune deleted it. Losing the sheet strips the microphone row's styling and
leaves SVG shapes filling black, because a path with no fill: none fills. The
sheet is injected during apply, which is after the loader's claim pass, so
this plugin's own claim never saw it either. It now carries data-plugin and
data-plugin-css, and a guard means a second apply no longer stacks a second
sheet.
tests/client-styles.test.mjs is new and covers both halves of that: that the
sheet names its owner and is no longer claimable by whichever plugin materialises
next, and that repeated applies install exactly one sheet. It loads client.cjs
the way the harness does, through window.__ModuleLoader__, against a minimal
DOM that implements the two selectors which decide ownership. npm test now
runs it beside the other six test files.
Verification: npm test, 53 tests passing, up from 51.
v0.6.0 - remove the diagnostics now that the live view works
The live view works, so the diagnostics are gone.
Removed
Everything added while hunting the fault, from what the user sees:
- the per-tick status line — floor, gate, chunk count, decode state, draft size
- the peek chain traces —
skipped,waiting,sending,posted,got,
queued,stopped,reply had no text,failed,threw - the settled/pending counter written on every draft update
- the unconditional probe, which wrote a placeholder into the draft
- the build marker
Kept
Only what is about the user rather than about this code:
| Note | Meaning |
|---|---|
cleaned up |
the final pass polished the transcript |
| a transcription error | with the reason the host gave |
transcription cancelled |
the take was abandoned |
The muted-microphone hint survives: it is derived from the audio and answers a
question the user has.
Guarded
tests/live.test.mjs asserts those instrumentation strings stay gone from the
hand-written half of the bundle, so the noise cannot creep back in with the next
change. Scoped to the hand-written half because the inlined modules legitimately
describe such things in their own documentation.
51 tests pass.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.6.0Then restart the harness — the native window injects client bundles at
launch, so nothing changes until it restarts.
v0.5.2 - the live view does not polish (verified), and reports the draft size
A measurement, not a change — and one small fix so the question can be answered
from the UI.
The live view does not polish
Verified against the running host with deliberately filler-heavy audio, posted to
both routes:
| Route | Output |
|---|---|
| Peek (live view) | UMM, so, uh, like, can you please, you know, refactor the parser and then, UMM, run the tests. |
| Final (on stop) | Can you please refactor the parser and then run the tests? |
The peek returns the raw recogniser output with every filler intact, no
polished flag, in 356 ms. An LLM cleanup pass costs 800 ms or more. Only
the final pass polishes, which is what was wanted.
Both call sites of the cleanup are guarded by polish: false, so a peek
structurally cannot reach it.
Why the live text can still look tidy
- Whisper punctuates and capitalises by itself. Its output is never a bare
word stream, so it reads as tidied even untouched. condition_on_previous_textis on, and Whisper is trained on clean text, so
it tends to drop hesitation sounds rather than emit them.- The text shifts as the recording grows. Every peek re-reads the whole audio,
so an early mishearing can be corrected a second later. That looks like a
cleanup pass and costs nothing. - LocalAgreement holds words back until two passes agree, so what is shown is
the stable part.
The small fix
The fact that answers "is this being polished?" is how much text reached the
draft — and it was reported by the write itself, on a note the 50 ms status tick
erased within a frame. It now lives on the status line, which is recomputed every
tick and so cannot be erased:
listening — floor 0.0042 gate 0.0110 · 9 chunk(s) · decoding · draft 42+7
The last field is settled + pending characters. If it grows while the text stays
tidy, the tidiness is the recogniser, not a cleanup pass.
51 tests pass.
v0.5.1 - read the held ref through .current, or the peek never runs
That idle timer was the bug report. 1791247122.8s idle is Date.now() - 0 —
epoch — which says lastPeekAt was never set, and it is only set inside
flushPeek. So flushPeek was never reached, and the reason was two lines up.
The bug
requestPeek and flushPeek read held.recorder and held.chunks. held is a
React ref, so those are undefined: the first comparison is merely false, and
the next .length throws a TypeError. That throw landed inside
ondataavailable, which has no catch — so the handler died silently on every
chunk and no peek ever ran.
The status line read held.current.chunks correctly. That is what made the two
disagree in a way that took so long to see: one place counted 17 chunks while
the other never spoke, and neither reported an error.
It is the same mistake as the misspelled name that hid here before, one level
subtler — the name exists, so nothing warns, and only the value is wrong.
What now guards it
tests/bundle-identifiers.test.mjs fails on any ref read without .current,
alongside the existing check for ref-family names that are never declared.
Verified by reintroducing this exact bug: the test fails on it.
51 tests pass.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.5.1Then restart the harness — the native window injects client bundles at
launch, so nothing is live until it restarts.
v0.5.0 - LocalAgreement decides what the draft shows
The live view now uses the algorithm this problem already had.
What changed
LocalAgreement, from Turning Whisper into Real-Time Transcription System
(Macháček et al., IJCNLP 2023) — the reference whisper_streaming implementation.
A batch recogniser applied to a growing recording does not produce settled words.
Each pass re-reads everything it is given, so its tail is a guess the next pass
may revise. A draft built from those guesses visibly rewrites itself as you talk.
The fix is to show a word only once consecutive hypotheses agree on it: the
longest common word prefix of the last two hypotheses is settled, and everything
after it waits. Settled text never changes, so the draft only grows.
Also adopted from the reference
--min-chunk-size 1.0— a decode is asked for once about a second of audio
exists. Below that the recogniser has too little context to place a word among
its neighbours and returns nothing. An empty answer reads exactly like a
broken pipeline, and this is very likely what was happening.- The draft shows settled words, then the pending tail, with the caret where the
next word will land — so it is visible which part is promised and which is
still being decided.
One bug fixed along the way
Withdrawing the live text before the final pass removed only the settled words,
leaving the pending tail in the draft — which the final pass would then sit on
top of, so the user would have sent it twice.
Removed
The unconditional diagnostic probe. It answered its question: the composer does
accept text while a recording runs, so every remaining fault was downstream. Left
in place it would overwrite the live transcript on every recording.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.5.0Then restart the harness — the native window injects client bundles at
launch, so nothing is live until it restarts.
50 tests pass.
v0.4.3 - send a peek on any new audio, and time-bound it
The live view never asked for a transcript. Two faults, both in the throttle
rather than in the trigger.
The deadlock
A peek required two new chunks of audio before it would run. A recording
that hands over a single chunk never satisfied that — 1 - 0 < 2 is never true
— so the live view silently never asked for anything, while the final pass on
stop worked perfectly. That is exactly the reported behaviour: correct on stop,
nothing while ongoing.
The interval between peeks is the right knob for cost. Counting chunks only
re-created the deadlock for any recording that arrives as one piece. One new
chunk is now enough.
The freeze
A peek had no deadline. A decode grows with the recording, and one slow or
hanging decode left the in-flight guard set for the rest of the take, refusing
every later peek — a live view that stops updating after its first slow decode,
which looks exactly like one that never worked at all. Peeks are now abandoned
after 8 seconds.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.4.3Then restart the harness: the native window injects client bundles at launch,
so nothing is live until it restarts.
Verified
41 tests pass. The identifier test also learned to ignore comments — it was
failing the build for an explanatory comment that named the variable it hunts
for, which is noise, since a comment cannot throw.
v0.4.2 - submit a rolling window instead of waiting for silence
The live view no longer waits for silence. It submits a rolling window.
Why the change
The live view waited for the span detector to report a pause before sending
anything. That made it depend on an estimate of the room — and the estimate can
be wrong in a way nothing recovers from.
A hum louder than the gate is indistinguishable from a voice. The gate is
derived from a noise floor the detector measures, and that measurement can sit
below the room: a quiet moment seeds it, or the room is simply louder than the
quietest thing it ever heard. When that happens no silence is ever detected, no
span ever closes, and the user talks until they press stop with nothing returned
— no text, no error, no clue.
That is exactly what happened here, and the report was precise: "the hearing you
thing seems to not go away." A span that never closes.
What it does now
The audio is submitted on a timer, and the route answers with the transcript
of everything sent so far.
There is no overlap to detect, and that is the point. Each answer is a complete
reading of the audio rather than a fragment to be aligned with the previous one,
so the draft is always internally consistent — the words are not stitched from
pieces that could seam badly. A growing prefix removes the alignment problem
instead of solving it.
The span detector is no longer consulted about when to send. It still feeds the
muted-microphone hint, which reads the level directly rather than the span state,
and its known limitation is now pinned by a test that asserts it, so nobody
rediscovers it as a bug.
Cost, stated plainly
Each peek decodes the whole recording, so decoding grows with the length of the
recording. A peek is only asked for once two seconds of new audio exist, which is
a handful of decodes for a voice instruction and becomes noticeable for a long
monologue. The final pass over the finished recording remains authoritative and
still polishes.
Verified
41 tests pass. The behaviour changed, so the tests that asserted the old trigger
were rewritten rather than deleted, and the detector's limitation is documented
by a test rather than by a comment.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.4.2Then restart the harness: the native window injects client bundles at launch,
so the fix is not live until it restarts.
v0.4.1 - fix the live view: an undeclared ref made every peek throw
The live view never worked. This release fixes it.
The bug
flushPeek — the only function the live view runs — read heldRef.current. The
ref is declared as held. No heldRef exists anywhere in the bundle.
So the first line of that function threw ReferenceError: heldRef is not defined
on every call. The caller invoked it as void flushPeek(), which discards the
promise, so the rejection had no handler and no route to the user.
The result was a feature that produced no symptom of failure at all: no
request, no draft text, no caret, and no error message. A fully wired live view
was in exactly the state of one that had never been connected. That is what made
it survive four rounds of diagnosis — including ones that added instrumentation,
because the throw happened before any of it could report.
Why the tests missed it
Every test in the suite passed while the feature was completely broken. They all
asserted on what the source text contains, and a missing name is invisible to
that: the string heldRef was present, spelled consistently, and simply did not
exist.
What changed
- The ref is read from its declared name.
- Every
flushPeek()call now carries a rejection handler that writes the failure
into the composer. A floating promise is precisely what turned a hard
ReferenceErrorinto silence. tests/bundle-identifiers.test.mjsfails on any name ending inRefthat is
read without being declared, and on any barevoid flushPeek(). Verified by
reintroducing the typo — the test fails on it.
40 tests pass, 2 of them new.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.4.1
python "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-voice-input\python\setup.py"Restart the harness afterwards: the native window injects client bundles at
launch, so the fix is not live until it restarts.
v0.4.0 - the transcript appears while you are still speaking
Your words appear in the draft as you say them.
What was wrong
The draft stayed empty until you pressed stop. Three separate causes, each found by measuring rather than reading:
- Every decoded segment arrived in the same millisecond as the final frame. On a 12.36 s recording: first partial at 1988 ms, final at 1989 ms. A 1 ms window is not a live view.
- Nothing was transcribed while you spoke. The recorder posted one blob when you stopped, so there was no audio to decode yet.
- On the whisper.cpp engine there is exactly one segment, emitted at
done, because its HTTP surface exposes no per-segment progress at all.
What it does now
The recorder hands over audio every second, and a span detector decides when you have paused. At each pause the audio captured so far is decoded and the draft is rewritten, so words appear while you are still talking. A pulsing caret marks where the next word will land.
The pause is the point. Speech recognition given a recording that stops mid-word does not fail — it invents an ending for the half-word it was handed. Two cut points in one sentence produced "at least once" and "address it"; the speaker said neither. Cutting at measured silence is what stops the live view from putting words in your mouth.
The final pass still wins. When you stop, the full recording is transcribed and polished exactly as before, replacing everything. A word the live view got wrong is corrected on the way in.
Typing is never blocked. Whatever you type mid-transcription is kept, and the live text lands after it.
The silence detector
It learns the room instead of assuming it. The old rule was a fixed line (rms > 0.1), which any steady hum defeats permanently — a fridge, a fan, a laptop under load. This tracks the noise floor from the level the meter already computes, so it needs no extra microphone access, and it self-corrects when the room changes.
Two rules earn their place, both forced by failing tests:
- No span opens until a full window has settled. The very first loud ticks arrive while the floor is still a guess, and a span opened then freezes the learning that would fix it.
- While a span is open the floor may only be lowered, never raised. Raising it takes the gate with it, and the gate is what decides you have stopped talking.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.4.0
python "$HOME/.dsh/profiles/web/node_modules/dsh-voice-input/python/setup.py"On Windows (PowerShell):
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.4.0
python "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-voice-input\python\setup.py"backend: faster-whisper is required for the live view — whisper.cpp cannot stream. The bundled config layer sets it.
Verified
- 38 tests, 18 new. The detector's cases are the behaviour worth having: an utterance inside a constant hum, a breath that must not split a sentence, a click that must not cost a decode, a muted microphone that must not close the gate.
- The new route was driven end to end against a real recording and returned the spoken sentence verbatim, ~350 ms warm.
- The browser bundle loads with a
requirethat answers nothing butreact, which proves the bundled detector is the one that runs. - The live text is withdrawn in one operation before the final pass writes, so the draft cannot carry the same words twice.
The span detector lives in lib/vad.mjs and is spliced into the browser bundle by
tools/inline-vad.mjs. It is not required as a package subpath, because the
client-module loader answers only a bare package name or <id>/client — a
subpath require would have thrown while the bundle was evaluated and taken the
microphone button with it. A test compares the inlined copy against the module
line by line, so the two cannot drift.
Not yet verified
Everything above the bundle boundary was tested without a browser. The caret's
appearance, the timing of the draft rewrites on screen, and the feel of the peek
cadence have not been watched by a human. If the caret does not appear, that is
the first place to look.
v0.3.0 - the cleanup pass no longer buys hidden reasoning time
The cleanup pass stops buying time it cannot show you.
The problem
A model call that names no reasoning effort inherits the adapter's default, which is high on a deployment that configures none — and this plugin renders only text deltas, never reasoning deltas. So the cleanup deliberated where you could not see it, while the draft sat still. Worse, the deliberation raced the 15-second deadline: on a transcript of any length the slow path could end in no cleanup at all, with the raw recognizer text returned after a long wait.
Removing "um" and adding punctuation is mechanical work. The reasoning bought nothing.
What changed
- The cleanup asks the route not to reason (
reasoningEffort: 'off'). - A route that cannot express that — the pi-ai adapter refuses it for a model it reports as supporting only
off— is retried once with the deployment's own default, so the cleanup is never lost to an accelerator you did not ask for. - Cleanup text is accumulated per attempt, so a refused attempt cannot leave partial text behind for the retry to append to.
polishReasoningis validated at boot, so a typo cannot reach the runtime as one failed call per recording that reads like a provider fault.
Every failure path still returns the recognizer's own text.
Choosing what you get
polish |
Behaviour |
|---|---|
conservative (default) |
clean without spending reasoning tokens — the fast one |
conservative-reasoned |
let the cleanup think; the previous behaviour, for comparison |
off |
return the recognizer's text untouched |
| Key | Default | Meaning |
|---|---|---|
polishReasoning |
off |
one of off/low/high/max; ignored by conservative-reasoned |
Override the voice-input row in your profile's cordis.patch.yml. A row's config is replaced, never deep-merged, so restate every key you want.
Install
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.3.0
python "$HOME/.dsh/profiles/web/node_modules/dsh-voice-input/python/setup.py"On Windows (PowerShell):
dsh plugin --profile web add github:moazzamak/dsh-voice-input#v0.3.0
python "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-voice-input\python\setup.py"Restart DSH afterwards.
Verified
- 20 tests (
npm test), 7 of them new intests/polish.test.mjs, which drives a real transcriber over a scripted LLM service and asserts on the assembled request rather than on source text. - The retry, the length guard, the empty-cleanup guard, and
polish: offmaking no call at all are each covered.
The previous tests asserted that the polish code mentions the LLM service and the deadline — which this defect satisfied perfectly. Nothing in the source text said the request was slow; only the request does.
Not yet measured
The latency improvement is reasoned from the code path, not timed on a live deployment: the request shape is verified, the round-trip saving is not. If cleanup quality drops with reasoning off, polish: conservative-reasoned is the one-word revert.
Also in this release
v0.2.0,v0.2.1andv0.2.2are now tagged. Their code was already onmain; onlyv0.1.0had a tag, so the Releases page still advertised 0.1.0.- The README shows the current release.