Skip to content

Releases: jryang1997/dsh-hold-to-dictate

v2.1.0 — 边说边出字(实验)

Choose a tag to compare

@jryang1997 jryang1997 released this 02 Oct 01:45

录音时,真实识别文字现在可以直接出现在输入框;后续识别会修正文字,松开后整段定稿。继续复用 Harness 官方语音模块。

本次更新

  • 新增「边说边出字(实验)」开关,默认关闭。短录音约每秒启动一次识别,等待上一轮完成后才发下一轮;长录音降低刷新频率。
  • 取消会撤回未被修改的临时语音文字,并保留区间外的手写追加;直接编辑语音区间时保留编辑并明确提示。
  • 保护手动修改、取消后的迟到结果和立即重录;失败重试不会重复插入已经出现的文字。
  • 中英文 README 加入实操 GIF、两张设置截图和更新说明。

更新与开启

按照 README 更新说明 更新插件并刷新 Harness,再进入 设置 → 插件 → 按住说话,开启 边说边出字(实验)。保持官方语音输入模块启用,并确认本地识别模型已经准备好。

仍使用 v1.x 旧包 @jryang1997/dsh-composer-dictation 的用户,应先移除旧包,再安装 @jryang1997/dsh-hold-to-dictate,避免同时启用两份。本机设置会保留。

使用范围

实时模式仅对本地语音服务和不含引用芯片的纯文本草稿启用;云端服务或不支持的环境沿用松开识别。它通过官方完整音频接口滚动识别,不是模型原生逐字流。出字速度取决于模型启动和推理耗时,文字及标点可能随音频增长而修正。

验证

  • npm run check:语法、字典、界面状态、实时输入、取消、手动编辑、重试和兼容回退检查通过。
  • npm pack --dry-run 通过。
  • 真实 Chrome 音频采集、重采样、整段录音解码与官方 SenseVoice 模型验证通过;录音时出字、松开定稿及保留手写追加的取消均已实测。

完整版本记录 · English usage

v2.0.2 — restore hover hint and soft fades

Choose a tag to compare

@jryang1997 jryang1997 released this 02 Oct 00:43

鼠标移入输入框时,语音操作提示恢复显示,并以 320 毫秒缓缓显现、移出淡去。

修复

  • 模型名称较长时,根据工具行实际空隙显示提示,避免错误隐藏。
  • 鼠标移入时重新测量工具行;刷新或切换会话时,鼠标已在框内也能显示提示。
  • 完整动效、轻柔动效及系统减少动态效果设置均保留柔和渐显。
  • 转写保留标签和失败重试按钮调用当前有效的操作。

改进

  • 录音胶囊采用主题强调色、随音量变化的光晕、稳定计时及插入草稿后的短暂确认。
  • 更新项目图标及中英文安装说明。

验证:JavaScript 语法、双语字典、渲染与悬停回归检查、安装包预检均通过;本机 Harness 已确认提示恢复。

GitHub 安装不会自动更新。已安装的用户请按 README 的更新步骤重新安装并刷新页面;原 dsh-composer-dictation 安装请先移除旧包,再安装 dsh-hold-to-dictate,设置会保留。

完整版本记录

v2.0.1 — the settings rows describe themselves again

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 06:41

A bug the test suite could not have caught, found by a screenshot of the real page.

Fixed

  • The two hold-duration rows were printing their own key as the description. numberRow built its hint key by appending Hint to the label key, so it asked the dictionary for holdMsLabelHint while the dictionary says holdMsHint. A missing key renders as itself, so the settings page showed:

    Hold duration
    holdMsLabelHint              ← instead of "How long the press must stay still…"
    

    That shipped in 1.5.0 and survived 2.0.0. Both rows are fixed.

Why the tests missed it, and what now stops the next one

Every existing assertion passed, because none of them looked at what the page actually says. The 90-odd checks all verify behaviour — that a setting reaches the gesture, that a chord stops working when another is chosen — and behaviour was never broken.

So the guard added here is deliberately blunt: it walks the rendered settings tree and fails if any string a user is meant to read looks like a dictionary key. It cannot catch every wrong key, but it catches exactly this class, and this class is invisible until someone looks at the page.

How it was found

Not by reading the code — by rendering the plugin's real component against the host's real stylesheet in a headless browser and looking at the pixels. That harness was built for the promo video, and it paid for itself before the video existed.

Full changelog: 2.0.0...v2.0.1

v2.0.0 — dsh-hold-to-dictate

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 05:30

The name changed. Nothing about what the plugin does changed with it — this is a major version because the install identity did.

dsh-composer-dictation → dsh-hold-to-dictate

Why the old name had to go

"Composer" is DeepSeek Harness's own word for the input box, which is why it was chosen. But to anyone scanning GitHub it is also PHP's package manager, and a dozen other things. And "dictation" says what the plugin does without saying how you invoke it.

The new name carries both: the gesture and the result.

Why not hold-to-talk

It was the obvious name and it was taken twice over:

Where Who
npm dsh-hold-to-talk belongs to a different plugin (latest 0.1.3)
GitHub wangzhanchao883/dsh-hold-to-talk and Gammonmush803/dsh-hold-to-talk

Both are hold-to-talk dictation plugins for the same harness. A third repository carrying that slug would have been the least visible of the three, and would have read as a copy of the first. dsh-hold-to-dictate was clear on npm and GitHub when it was chosen.

Upgrading

dsh plugin remove @jryang1997/dsh-composer-dictation
dsh plugin add github:jryang1997/dsh-hold-to-dictate

Or, if you installed it through the desktop app: Plugins → remove the old row → add github:jryang1997/dsh-hold-to-dictate. GitHub redirects the old URL, so existing links keep working.

Your settings survive the move. That is what the unchanged storage key buys.

Two things deliberately did not follow the package name

  • The settings storage key stays dsh-composer-dictation.config. It is a storage key, not an identity — it is what your browser has already written your settings under. Renaming it to match the package would silently discard every value you had tuned, and buy nothing.
  • The dsh-htt- CSS prefix. htt stood for the project's first name, but it is an opaque namespace no user ever sees, and the changelog's own history refers to .dsh-htt-tool. Renaming it would have buried the real change under 279 lines of mechanical diff.

Both are commented where they live, because they otherwise read as oversights.

What you will notice

Nothing. No behaviour changed, no setting moved, no pixel shifted. The plugin list still calls it Hold to talk / 按住说话 — the package name is what GitHub sees, the title is what you read in Settings.

Full changelog: 1.5.2...v2.0.0

v1.5.2 — aligned to what is actually rendered

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 04:27

1.5.1 aligned the hover hint to a CSS rule that turned out not to be rendered by anything.

Fixed

  • The hint was given the wrong font weight. 1.5.1 copied 13px / 500 / 20px from the InputBar stylesheet's .RlGAzG_select rule. That rule is dead CSS — nothing in the host renders it, so it was the wrong thing to align to. The real model selector is ModelSelect's trigger, and it is 13px / 400 / 20px with normal tracking. The hint was therefore one weight too heavy. It is now 400 — which is what it had before 1.5.1, and what the control beside it actually uses. The 500 in that stylesheet belongs to the permission chip, and a chip is not body text.

  • Vertical alignment is now measured rather than assumed. 1.5.1 centred the hint on the tool row's box, and that row is padding: 2px 8px 6px — asymmetric. Every control in the row is centred on the content box instead, so a box-centred hint sat exactly 2 px lower than the model selector beside it: too small to name, large enough to read as "something is off".

    The hint is now centred on the trailing group's own measured box, so it lands on the same line box as the control it sits next to — whatever the row's padding happens to be, and without hard-coding a single number.

On getting the reference right

The lesson is worth writing down: a plausible-looking rule in the host's own stylesheet is not evidence that anything renders it. Both of these bugs came from reading CSS rather than from measuring the element on screen — which is exactly why the vertical fix now measures the neighbour instead of deriving its position.

Full changelog: 1.5.1...v1.5.2

v1.5.1 — a correction

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 04:11

A correction. 1.4.0 fixed a problem that did not exist.

Fixed

  • The duplicate microphone is gone. 1.4.0 added a focusable microphone button to the tool row on the grounds that a chord cannot be found by pressing Tab. That reasoning was wrong: this plugin cannot run without the official voice-input bundle, and that bundle already puts a focusable microphone button next to Send. The button was never the only keyboard entry — it was a second microphone a few centimetres from the first, which is exactly what it looked like on screen.

    The chord stays, and the honest limitation returns with it: the shortcut is documented rather than advertised, and the visible microphone belongs to the official plugin.

  • The hover hint now speaks the tool row's typography. It was 13px / 400 / 18px carrying its own letter-spacing; the host's model selector is 13px / 500 / 20px with none. Sitting between the mode chips and the model selector, it read as a foreign element that happened to be nearby rather than a line of the same row.

  • The microphone glyph is now the host's own, reproduced glyph-for-glyph: IconMicrophoneOutlineRegular from @deepseek-ai/dsh-client-ui-primitives — a stroked capsule and an arc, fill: none, stroke: currentColor, 1px in a 16×16 box. The previous one was a hand-drawn filled shape, which is why the two microphones never quite matched.

Removed

  • The button's component, its stylesheet, the three strings it needed, and the session store that carried state between it and the recording surface. The store existed solely so two seats could talk to each other; with one seat left, it has nothing to say.

Why a patch version

Nothing gained a capability here; a mistake made seven minutes earlier was corrected. SemVer's rule about removals assumes somebody depended on the thing.

Full changelog: 1.5.0...v1.5.1

v1.5.0 — built for a finger too

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 03:58

The gesture was built for a mouse and never tuned for a finger. The host does nothing about touch at all, so the plugin has to.

Added

  • A separate touch threshold. A touch press gets its own hold — 450 ms by default, configurable from 250 to 1200 — and a wider slop. A finger rolls, and a long press on a touch screen is also how a word gets selected, so the threshold has to be long enough that a deliberate hold is unmistakably deliberate.
  • Suppression of what a long press otherwise triggers. While a touch gesture is armed or recording, contextmenu and selectstart are cancelled in the capture phase. Neither is handled anywhere near the composer — there is no touch-action on the card, the editor or any ancestor, and the desktop shell raises a native context menu from the main process — so the collision was real rather than theoretical.
  • The press ring now draws over the hold the press actually uses, so the arc still finishes at the moment recording begins rather than 150 ms early.

Deliberately scoped

Suppression applies to touch only. Cancelling selectstart for the mouse would break dragging a selection out of the same card — so the tests assert both halves: an armed touch press must cancel, and a mouse press must not.

touch-action is not used, and that is a finding rather than an omission: it governs panning and zooming, not long-press selection.

Known limitation

Whether preventDefault() on the DOM contextmenu event reaches the desktop shell's main-process context-menu handler has not been confirmed on real hardware. This plugin has never been driven from a touch screen. The README says so in place of pretending otherwise.

Full changelog: 1.4.0...v1.5.0

v1.4.0 — something to press

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 03:51

The keyboard entry added in 1.2.0 was real but invisible: a chord cannot be found by pressing Tab. Now there is something to find.

Added

  • A microphone button in the composer's tool row, to the left of the model selector. conversation.input.left is an empty list slot inside that row, so what lands there is a real flex child — Tab reaches it, Enter activates it, and aria-keyshortcuts on it is the only place a keyboard user can learn the chord exists.
  • Its model is a toggle rather than a hold: press to start, press again to finish. Holding a key down with a button is an awkward thing to ask of a keyboard, and it is what the shipped microphone does. The chord stays the hold-shaped entry.

How the two seats talk

The recording surface lives in conversation.input.overlay and the button in conversation.input.left — different slots, different owners, so neither can pass the other a prop. A small session store carries exactly what they need to say to each other: one phase, three commands, and deliberately nothing more. It is not an event bus.

Changed

  • The known limitation is gone. "The keyboard shortcut is documented, not advertised" was accurate for exactly one release.
  • A new one replaces it, and it is real: the composer hides conversation.input.left while the shipped voice input's activity is expanded, so during an official recording the chord is the only way in.

Full changelog: 1.3.0...v1.4.0

v1.3.0 — settings, four of them

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 03:46

Settings → Plugins → Dictation now opens the bundle's own page.

Added

  • A settings page, registered into plugins.bundle.config and keyed by package name — the seat the manager offers a bundle for its own page. Four knobs, deliberately only four:

    Setting Why it is yours to change
    Hold duration (150–800 ms, default 300) Hands differ more here than at any other threshold
    Motion: full / calm Calm keeps the cross-fades and drops movement and scale — the same softening as prefers-reduced-motion, but chosen rather than signalled
    Hover hint on / off Some want the reminder, some find it noise
    Keyboard shortcut: off, Ctrl+Shift+Space, Ctrl+Shift+D, Ctrl+Shift+M, Ctrl+Alt+Space Shortcut conflicts are personal

    Deliberately not there: colours, materials, radii, the drag-to-discard distances and the waveform's physics — those are the design. The speech provider and language are absent for a different reason: this plugin sends neither, so the Host's Voice input settings govern. There is no theme setting because the plugin has no theme of its own; it reads every value from the Host's tokens and follows light and dark automatically.

Changed

  • Configuration became a module. createConfig owns the defaults, the clamping, the persistence and the change notification behind four methods — get / set / subscribe / reset. These were module constants read straight out of the gesture, and a page that writes them would otherwise have scattered storage calls across the whole effect.
  • Chord matching moved from a hard-coded Ctrl+Shift+Space to a catalogue matched on event.code, so a keyboard layout that moves the letters around cannot silently break the shortcut.
  • The press ring's duration is driven from configuration instead of being baked into the stylesheet.

A deliberate trade

Settings live in the browser's storage, not the DSH profile. That keeps the plugin free of any @deepseek-ai/dsh-* dependency — a wrong peer range makes DSH skip the entire bundle, silently, with nothing on screen to explain it. The cost is that settings do not travel to another machine, and that is written down in the README rather than hidden.

The Host's form primitives render text inputs only — no switch, no select, no stepper — so the controls here are built from the same theme tokens everything else uses.

Full changelog: 1.2.0...v1.3.0

v1.2.0 — a way in without a mouse

Choose a tag to compare

@jryang1997 jryang1997 released this 01 Oct 02:04

Making good on the three gaps the previous release left open: the gesture had no keyboard equivalent, a failure flashed past without offering a way out, and errors announced themselves as politely as a status update.

Added

  • A keyboard equivalent. Hold Ctrl+Shift+Space to record and release to transcribe — the same gesture, the same state machine, no pointer. It works anywhere in the app, so you do not have to focus the input box first, and Esc while still holding discards. The chord is deliberately awkward to hit by accident and never claims a keystroke unless it actually starts a recording.
  • Retry, in place. A failed transcription used to mean saying the whole sentence again. The failure card now re-sends the recording that is already captured, so a network hiccup or a provider error costs one click instead of one repetition.
  • A 1280×640 link preview card (docs/social-preview.png).

Changed

  • A failure is no longer a notice. It stays on screen until it is dismissed or retried, carries its own controls, clamps to two lines with the full text on hover, and uses role="alert" so assistive technology treats it as urgent. The transient notice keeps role="status", because nothing is being asked of the user there.
  • Saying nothing is no longer silent. A recording shorter than MIN_SECONDS used to drop out with no feedback at all — which a keyboard tap makes much easier to hit. It now says so.
  • A second entry point into begin() is guarded, so a hold and a chord can never race into two captures.

Known limitations

Full keyboard parity would need a focusable trigger, and the slot this plugin renders into does not offer one — the shortcut is therefore documented rather than advertised in the interface. That is stated plainly in the README.

Full changelog: 1.1.0...v1.2.0