Skip to content

Releases: NotWizard/Mouthpiece

v2.2.1: 更新提醒、自动下载开关、唤起即前台 / Update hint, auto-download toggle, front-on-open

Choose a tag to compare

@github-actions github-actions released this 30 Sep 03:28

更新有了自己的提醒方式,你还可以选择让它在后台静默完成;打开控制面板的体验也更顺畅了。

🎉 新功能

  • 侧边栏"有可用更新"提示:发现新版本后,侧边栏左下角出现带呼吸小圆点和版本号的常驻提示——关掉更新弹窗后依然在,点一下就跳到"权限与诊断"并重新弹出安装窗,再点一下即完成升级。Reduce Motion 开启时呼吸动画自动关闭。
  • 自动下载更新开关(默认关):在"通用 → 软件更新"里可以开启后台静默下载——发现新版本即自动下载,完成后弹窗提醒你安装。默认关闭,保持"跳过 / 安装并重启"两枚按钮的简洁弹窗。

✨ 改进

  • 控制面板唤起后立即置于前台:通过 Spotlight、Dock 图标或菜单栏打开控制面板时,窗口不再落在前一个应用的后面。

Updates now have their own reminder, you can let them happen silently in the background, and opening the control panel feels smoother.

🎉 New features

  • Sidebar "update available" hint: When a new version is found, a breathing dot and the version number appear at the sidebar's bottom-left — they stay even after you close Sparkle's dialog. Clicking jumps to Privacy & Diagnostics and reopens the install prompt; one more click finishes the upgrade. The breathing animation drops to a static dot under Reduce Motion.
  • Auto-download updates toggle (off by default): Enable it under General → Software Update to silently download new versions in the background and get a prompt when ready. Off by default, keeping the update dialog to a clean Skip / Install-and-Relaunch.

✨ Improvements

  • Control panel opens in front: When invoked via Spotlight, the Dock icon, or the menu bar, the panel window now appears immediately in front of the previously focused app instead of behind it.

v2.2.0: 听写兜底更稳、新增 OpenRouter 与 DeepSeek、词典全面生效 / Sturdier dictation, OpenRouter & DeepSeek, dictionary everywhere

Choose a tag to compare

@github-actions github-actions released this 29 Sep 10:21

听写失败有了多层兜底,个人词典对所有云服务生效,还多了两个新服务商可选。

✨ 改进

  • 听写不再轻易丢语音:百炼实时识别中断或失败时,自动改用一次性转写重试,仍失败再落到本地 Whisper——弱网或服务抖动时,说过的话不会无声消失。
  • 个人词典全服务生效:词库里的人名、术语现在对 Soniox、AssemblyAI、Mistral 的听写同样生效(此前仅百炼),换服务商不用重新攒词。
  • 文字整理会纠"听岔"的词:整理时结合整段上下文,把发音相近但明显不合语境的词改回正确写法;拿不准的词保持原样,不会过度改写。
  • 百炼模型升级:默认识别模型升级至 Qwen Audio 3.1 消息版,识别结果与此前一致、费用约降 62%;另可在设置中选择 3.1 流式版。已有设置自动升级,无需操作。
  • 切换服务商更省心:文字处理的模型名按服务商分别记忆——切走清空、切回自动恢复,不会再把一家的模型名带到另一家导致报错。
  • 默认识别模型更新:OpenAI、Soniox、AssemblyAI 的新装用户默认使用各家当前推荐的识别模型;已自定义的模型不受影响。

🎉 新功能

  • 新增 OpenRouter 听写服务:一个账号即可在 20 多家语音识别模型之间随意切换(默认为微软 MAI-Transcribe 2,中文表现优秀),无需逐家注册。
  • 新增 DeepSeek 文字处理服务:整理与翻译可选用 DeepSeek(默认使用其快速模型),价格低、中文表现出色;提供与百炼同款的深度思考开关,默认关闭以保持整理速度。

Dictation now falls back instead of failing, the personal dictionary works with every cloud provider, and two new providers join the lineup.

✨ Improvements

  • Your speech no longer vanishes on a hiccup: When Bailian's realtime recognition drops mid-dictation, the recording is automatically re-transcribed in one shot, falling back to local Whisper only if that fails too — network hiccups no longer swallow what you said.
  • Personal dictionary works everywhere: Names and terms from your dictionary now also boost Soniox, AssemblyAI, and Mistral dictations (previously Bailian only); switching providers no longer means starting your vocabulary over.
  • Cleanup fixes misheard words: Text cleanup now reads the whole transcript and corrects words that sound similar but clearly don't fit the context; uncertain words are left untouched, so nothing gets over-edited.
  • Bailian model upgrade: The default recognition model moves to Qwen Audio 3.1 Message — identical transcripts at roughly 62% lower cost — with the 3.1 streaming model also selectable. Existing settings upgrade automatically.
  • Smoother provider switching: The text-processing model is remembered per provider — switching away clears the field and switching back restores your model, so one vendor's model name never leaks into another's request.
  • Fresh defaults: New OpenAI, Soniox, and AssemblyAI installs default to each vendor's current recommended recognition model; existing custom models are untouched.

🎉 New features

  • OpenRouter dictation: One account, 20+ speech-recognition models to switch between at will (Microsoft MAI-Transcribe 2 by default, strong in Mandarin) — no per-vendor signups.
  • DeepSeek text processing: Cleanup and translation can now run on DeepSeek (its fast model by default) — inexpensive and strong in Chinese, with a deep-thinking toggle mirroring Bailian's that stays off by default to keep cleanup fast.

v2.1.6: 个人纠错学习、批量归纳、词库确认 / Personal corrections, batch learning, vocabulary review

Choose a tag to compare

@github-actions github-actions released this 12 Sep 09:31

语音输入后亲手改对的词,现在可以积累起来,帮助之后的听写识别你的个人用词。

🎉 新功能

  • 个人纠错学习:在词库页开启“从我的纠正中学习”,支持的输入框会收集听写后的修改。功能默认关闭,由你决定是否启用。
  • 整批归纳:默认累计 50 次转写后统一提取词语,也可调整次数或按天归纳。未修改的内容提供上下文,新词必须有人工纠正作为依据,不会每改一句就调用模型。
  • 确认后加入词库:建议按人名、产品、专业术语等分类,支持查看修改前后原句、编辑拼写、批量加入、忽略和恢复。
  • 手动补充纠正:可在历史中补充纠正稿;也可在原应用选中改好的听写片段,通过快捷操作核对并保存。

使用说明

批量归纳会将相关文本发送给已配置的智能处理服务,可能产生该服务的调用费用。失败时记录保留,可手动重试。

自动捕获取决于输入框的支持情况;终端等无法可靠定位的场景请使用手动补充。词语确认后才会参与后续处理,不会自动生成强制替换。清除学习记录不影响已有词库和原转写历史。

Words you correct after dictation can now accumulate into personal vocabulary for future dictations.

🎉 New features

  • Personal correction learning: Enable “Learn from my corrections” in Dictionary to collect edits in supported input fields. The feature is off by default.
  • Batch analysis: Analyze 50 dictations together by default, choose another batch size, or use daily analysis. Unchanged text provides context; new suggestions require a user correction. Individual edits do not trigger separate model requests.
  • Review before learning: Suggestions are grouped into names, products, technical terms and other categories. View the original correction, edit spelling, accept in bulk, ignore or restore suggestions.
  • Manual corrections: Add corrected text from History, or select the corrected dictation in its original app and use the review shortcut to save it.

Getting started

Batch analysis sends relevant text to your configured text-processing provider and may incur its usage charges. Failed batches remain available for retry.

Automatic capture depends on input-field support. Use manual corrections in terminals and other fields where reliable tracking is unavailable. Approved terms inform future processing without creating forced replacements. Clearing learning data preserves your existing dictionary and original history.

v2.1.5: 久置后听写不再假死

Choose a tag to compare

@github-actions github-actions released this 06 Sep 10:40

修复一个会让整个应用失去响应的严重问题。

🐛 Bug 修复

  • 修复应用长时间运行(尤其经历睡眠唤醒)后,第一次听写让整个应用卡住、只能强制退出的问题。旧录音设备在系统音频服务异常时无法在主线程上正常关闭,把整个应用一起拖住。现在设备关闭改在后台进行:即使系统音频再异常,也只会影响当次听写——15 秒后会提示「启动超时」,可以直接重试,应用本身始终可用。

v2.1.5: No more freeze after long idle

Fixes a serious issue that could make the whole app unresponsive.

🐛 Bug fixes

  • Fixed the first dictation after a long uptime (especially across sleep/wake) freezing the entire app so it had to be force-quit. The previous recording device could not shut down on the main thread while the system audio service was stuck, taking the whole app down with it. Device shutdown now happens in the background: even if the system audio service misbehaves again, only that one dictation is affected — after 15 seconds you get a "start timed out" notice and can simply retry, while the app itself stays responsive.

v2.1.4: 听写收尾更可靠

Choose a tag to compare

@github-actions github-actions released this 05 Sep 04:19

修复四个听写收尾环节的偶发问题,减少卡住和文字不完整的情况。

🐛 Bug 修复

  • 修复关闭「自动粘贴」时,开启文字整理或翻译的听写会一直卡在「处理中」的问题:文字实际已处理完,但结果不复制到剪贴板、不保存到历史,只能重新开始。现在按设置正常收尾。
  • 修复识别服务在录音初期直接报错后,应用一直停在错误状态、无法再次开始听写的问题。现在错误提示展示结束后会正常回到待命状态。
  • 修复文字整理过程中网络流被截断或报错时,把只生成了一半的文字当作最终结果的问题。现在检测到中断会回退为完整的原始转写,不会丢内容。
  • 修复文字处理超过 8 秒上限后仍在干等的问题。超时或按 Escape 取消现在都会立即生效并照常收尾。

v2.1.4: More reliable dictation wrap-up

Fixes four occasional issues in the dictation finish path, reducing stuck sessions and incomplete text.

🐛 Bug fixes

  • Fixed a freeze at "processing" when auto-paste was off and text cleanup or translation was enabled: the text was actually done, but it never reached the clipboard or history and the session had to be restarted. The session now finishes according to your settings.
  • Fixed the app getting stuck on the error state after the recognition service failed early in a recording, which blocked any new dictation. The error now clears after a brief display and the app is ready again.
  • Fixed half-written text being treated as the final result when the cleanup stream errored or was cut off mid-way. An interrupted stream now falls back to the complete raw transcript, so nothing is lost.
  • Fixed the 8-second text-processing limit not actually applying. Timeouts and Escape cancellation now take effect immediately and the session finishes as usual.

v2.1.3: 加强偶发卡死的排查能力

Choose a tag to compare

@github-actions github-actions released this 04 Sep 12:28

本次为排查用的小版本,功能没有任何变化。

✨ 改进

  • 增强了听写启动阶段的内部诊断记录。针对「应用长时间闲置后第一次听写偶发卡死、只能强制退出」的问题,新版本会在启动的每个环节留下更详细的排查痕迹;万一同类情况再次发生,凭日志即可快速定位卡住的具体环节,为彻底修复做好准备。

v2.1.3: Better diagnostics for a rare startup freeze

A small diagnostic release — no feature changes.

✨ Improvements

  • Dictation startup now records more detailed internal diagnostics. This targets an occasional freeze on the first dictation after the app has been idle for a long time, which previously forced a manual force-quit. If it ever happens again, the new traces pinpoint the exact step that got stuck, paving the way for a permanent fix.

v2.1.2: 文本整理更快、卡住会兜底

Choose a tag to compare

@github-actions github-actions released this 27 Aug 09:56

语音转文字之后的文本整理和翻译更快了,万一很慢也不会再让你干等。

✨ 改进

  • 语音转文字后的文本整理与翻译现在更快:结果边生成边返回,不再等整段响应生成完才显示。同样的整理,过去偶尔要等十几秒,现在通常几秒内就能完成。
  • 整理万一很慢也不会再卡住:超过 8 秒仍没结果,会自动改用未整理的原始转写并照常插入,不会让你一直等着。

🐛 Bug 修复

  • 常规设置里的「上次会话错误」卡片已移除。它对「没检测到语音」这类正常情况也会常驻一张橙色提示、还得手动点一次清除,几乎没有意义;会话出错的信息仍会在录音胶囊上闪现提示。

v2.1.2: Faster text cleanup, with a fallback when it's slow

Text cleanup and translation after dictation are faster, and a slow run no longer leaves you waiting.

✨ Improvements

  • Text cleanup and translation after dictation are faster: the result now streams in as it is generated instead of waiting for the whole response. A cleanup that occasionally took ten-plus seconds usually finishes in a few now.
  • A slow cleanup no longer makes you wait: if there is still no result after 8 seconds, the raw (un-cleaned) transcript is inserted as usual.

🐛 Bug fixes

  • The "Last Session Error" card in General settings is gone. It kept a persistent orange notice even for normal cases like "no speech detected" and made you clear it by hand for little benefit; session errors still flash on the recording capsule.

v2.1.1: 修复默认听写快捷键失效

Choose a tag to compare

@github-actions github-actions released this 25 Aug 03:13

紧急修复 v2.1.0 引入的回归:默认听写键(右 Command)按下毫无反应。

🐛 Bug 修复

  • 默认的按住说话快捷键——右 Command,以及一切「单修饰键」(左/右 Shift、Option、Control、Fn)——恢复可用。v2.1.0 为区分左右同位修饰键,给这类键的按下判定加了一道实时 CGEventSource.keyState 读取,但该读取在事件 tap 内对修饰键并不可靠,导致按键始终被判定为「未按下」、毫无反应;而普通组合键(如 Command+0)走的是另一条未受影响的路径,仍然正常。此版本将判定逻辑逐字回退到 v2.1.0 之前经过长期验证的实现。
  • 说明:v2.1.0 的自动化测试之所以没能拦住这个回归,是因为失败发生在实时事件 tap 的运行时行为上,而非可被单元测试覆盖的纯逻辑;对应的伪测试已删除。左右同位修饰键那个边角问题暂列为已知限制,后续用更稳妥的方式(读取事件自带的设备相关修饰位)修复。

v2.1.1: Fix the default dictation hotkey doing nothing

An urgent fix for a v2.1.0 regression: pressing the default dictation key (Right Command) did nothing.

🐛 Bug fixes

  • The default push-to-talk hotkey — Right Command, and every other modifier-only key (left/right Shift, Option, Control, Fn) — works again. To disambiguate the left and right keys that share one modifier bit, v2.1.0 added a live CGEventSource.keyState read to the press decision for these keys, but inside the event tap that read does not reliably reflect a modifier key's state, so the key was always judged "not pressed" and nothing happened; regular combo keys (e.g. Command+0) took a different, unaffected path and kept working. This release reverts the decision verbatim to the long-proven pre-v2.1.0 implementation.
  • Note: v2.1.0's automated tests did not catch this because the failure is in the live event-tap runtime behavior, not in unit-testable pure logic; the misleading test that gave false confidence has been removed. The left/right same-modifier edge case is a documented known limitation for now, to be fixed later by reading the device-specific modifier bit carried on the event itself.

v2.1.0: 一轮最高标准审计后的数据安全、隐私与健壮性加固

Choose a tag to compare

@github-actions github-actions released this 24 Aug 09:32

v2.1.0: 一轮最高标准审计后的数据安全、隐私与健壮性加固

这是一轮系统性代码审计后的集中修复版本:覆盖数据丢失、隐私泄漏、进程健壮性、发布链安全与无障碍打磨,共修复 29 项审计条目并新增回归测试(单元测试 116 → 160)。绝大多数用户无需任何操作即可获益;少数条目仅影响新保存的数据(见下)。

🔒 数据安全与隐私

  • 多句听写不再只保留最后一句:跨长停顿的首句音频不再被裁掉。
  • 实时听写(Bailian/Volcengine)在 socket 中途断开或收尾报错时,不再摧毁已采集的完整录音——有兜底时降级到本地/批量转写,无兜底时如实呈现原始错误。
  • 「允许本地兜底」对实时独占供应商真正生效:空转写不再过早判定「无语音」而丢弃录音。
  • Secure Input(密码框、sudo、部分登录框)下不再静默丢字:粘贴被系统丢弃时保留剪贴板供手动粘贴,并在胶囊明确提示。
  • 听写中途切换输入框,文字不再落进原来的框。
  • 转写文本写入剪贴板时标记为「机密」(org.nspasteboard.ConcealedType),Alfred / Paste / Raycast / Maccy / Pastebot 等剪贴板管理器不再长期留存听写内容。
  • 敏感应用(密码管理器、银行等)保护统一覆盖插入 / 历史 / 日志 / 云推理四条出口——密码框里的听写不再被上传云端或写入历史。
  • 删除历史 / 清空后不再在数据库与 -wal 里残留明文。
  • 保留策略(90 天 / 行数上限)现在有明确的界面告知。

🛠 健壮性

  • 实时事件改为单消费者顺序处理,杜绝迟到的 partial 覆盖已定稿的 final 造成转写错乱。
  • 停止 增加看门狗,异常挂起也能在有界时间内收尾,不再永久卡在「停止中」。
  • 慢连接冷启动不再静默丢失开头语音;五家实时供应商的收尾等待统一到一致的超时预算。
  • 本地模型服务器进程在崩溃/强退后会被回收,不再泄漏数 GB 内存;退出仅 SIGTERM 无效时升级到 SIGKILL。
  • 本地模型下载改为校验和 + 原子替换,失败不再删掉已有可用模型。
  • 供应商错误提示统一截断与脱敏,不再把上游 HTML 或回显的请求内容抛给用户。
  • 初始化失败不再假装「就绪」:进入可重试的降级态;退出前对在途听写、剪贴板恢复、设置写入做有界 flush。

⚡️ 性能与打磨

  • 听写时打开控制面板不再每帧全量重渲染(50Hz → 仅相关视图)。
  • 设置改动改为去抖写盘,滑块拖拽不再每帧重写整块设置。
  • 历史清理不再每次写入都跑一次全表反连接。
  • 词汇规则(避免词 / 替换规则)真正生效:作为推理提示软约束 + 收尾文本后处理。
  • 无障碍:胶囊 VoiceOver 播报覆盖全部相位;HUD 与控制面板尊重「减弱动态效果」「增强对比度」;小图标命中区达 44pt;引导页在拒绝辅助功能权限时可跳过。

🧰 发布链

  • CI 增加 pbxproj/project.yml 漂移门禁、appcast「本次条目」校验、以及测试执行/跳过下限门禁。
  • 发布产物校验断言 bundled 模型运行时与 MediaRemote 适配器确实在 .app 内且架构匹配;DMG 打包改用唯一卷名与相对路径校验和。

v2.1.0: Data-safety, privacy, and robustness hardening after a top-standard audit

This is a consolidated fix release following a systematic code audit: it covers data loss, privacy leaks, process robustness, release-pipeline safety, and accessibility polish — 29 audit items fixed with regression tests added (unit tests 116 → 160). Nearly all users benefit with no action required; a few items affect only newly-saved data (see below).

🔒 Data safety & privacy

  • Multi-utterance dictation no longer keeps only the last sentence — the first utterance across a long pause is no longer trimmed away.
  • Realtime dictation (Bailian/Volcengine) no longer destroys the captured recording when the socket dies mid-session or finalize errors — it degrades to local/batch transcription when a fallback exists, and surfaces the real error otherwise.
  • "Allow local fallback" now actually works for realtime-only providers: an empty transcript no longer prematurely reports "no speech" and discards the recording.
  • Under Secure Input (password fields, sudo, some login forms) text is no longer silently dropped — when the paste is rejected the transcript is kept on the clipboard for a manual paste and the capsule explains why.
  • Switching text fields mid-dictation no longer lands the text in the old field.
  • Transcripts written to the clipboard are marked concealed (org.nspasteboard.ConcealedType), so clipboard managers (Alfred / Paste / Raycast / Maccy / Pastebot) no longer retain dictation content.
  • Sensitive-app protection (password managers, banking, etc.) now covers all four sinks — insertion / history / logs / cloud reasoning — so a password dictated into a protected app is neither uploaded to the cloud nor stored in history.
  • Clearing history no longer leaves plaintext behind in the database or the -wal sidecar.
  • The retention policy (90 days / row cap) is now disclosed in the UI.

🛠 Robustness

  • Realtime events are now processed by a single ordered consumer, so a late partial can no longer overwrite a finalized result.
  • stop() gains a watchdog so a hung teardown still finalizes within a bounded time instead of parking in "stopping" forever.
  • Slow cold-start connects no longer silently drop leading audio; the finalize wait across all five realtime providers now shares one consistent timeout budget.
  • Local model-server processes are reaped after a crash/force-quit instead of leaking multiple GB; quit escalates SIGTERM → SIGKILL when a child ignores SIGTERM.
  • Local model downloads now verify a checksum and swap atomically, so a failed download no longer deletes an existing good model.
  • Provider errors are uniformly truncated and sanitized, so upstream HTML or echoed request content no longer reaches the user.
  • A failed init no longer pretends to be "ready" — it enters a retryable degraded state; quit performs a bounded flush of in-flight dictation, clipboard restore, and settings writes.

⚡️ Performance & polish

  • Opening the control panel during dictation no longer re-renders the whole panel every frame (50 Hz → only the relevant views).
  • Settings changes are debounced to disk, so a slider drag no longer rewrites the entire settings blob every frame.
  • History pruning no longer runs a full-table anti-join on every single write.
  • Vocabulary rules (avoided terms / replacement rules) now actually take effect: as a soft reasoning-prompt constraint plus final-text post-processing.
  • Accessibility: the capsule's VoiceOver announcements cover every phase; the HUD and control panel honour Reduce Motion and Increase Contrast; small icon buttons meet the 44 pt hit-target minimum; onboarding can be advanced when Accessibility permission is declined.

🧰 Release pipeline

  • CI gains a pbxproj/project.yml drift gate, an appcast "current release present" check, and a test execution/skip floor gate.
  • Release artifact verification asserts the bundled model runtimes and the MediaRemote adapter are actually inside the .app with matching architecture; DMG packaging switched to a unique volume name and relative-path checksums.

v2.0.10: 关闭程序坞图标的行为更可靠

Choose a tag to compare

@github-actions github-actions released this 21 Aug 07:46

修复一个"在程序坞中显示"设置失效的问题,并顺带清理了历史遗留代码。

🐛 Bug 修复

  • 修复了关闭"在程序坞中显示"后,从聚焦搜索/访达/启动台重新打开控制面板,会导致程序坞图标重新出现并一直不消失的问题。现在无论从哪里唤醒控制面板,都会正确遵循你的设置。

🧹 内部清理

  • 移除了早期 Electron 版本的设置迁移代码及其依赖(老用户均已完成迁移)。本地已下载的语音模型不受影响,无需重新下载。

v2.0.10: More Reliable "Hide Dock Icon" Behavior

Fixes a case where the "Show in Dock" setting stopped taking effect, plus some legacy cleanup.

🐛 Bug fixes

  • Fixed the Dock icon reappearing for good after reopening the control panel from Spotlight/Finder/Launchpad while "Show in Dock" was off. Waking the control panel from anywhere now respects your setting.

🧹 Internal cleanup

  • Removed the legacy Electron settings-migration code and its dependency (all users have already migrated). Locally downloaded speech models are unaffected and are not re-downloaded.