Skip to content

Releases: DragonKingIO/Cadenza-voice

Cadenza 1.2.0 (early preview)

Choose a tag to compare

@DragonKingIO DragonKingIO released this 09 Oct 15:38
7012f3a

Early preview. Cadenza 1.2.0 adds voice translation with its own shortcut, more cloud speech and text recognition services, two more on-device speech models, and quieter speech handling. The package runs on Apple silicon and Intel Macs, macOS 14 or later.

中文说明见下方。

What's new

  • Voice translation has its own shortcut and page (Settings → Translation). Ordinary dictation no longer translates. If you had a language chosen, set a translate shortcut to keep translating.
  • Shortcuts: "Hold to talk" and "Tap to start and stop" are two rows, each with its own switch and key. The Hold/Tap choice is gone.
  • More cloud speech services: OpenAI, Groq, Google Cloud, Microsoft Azure, AssemblyAI, ElevenLabs, and any service that copies the OpenAI transcription API. Each one has a setup guide on the website.
  • More cloud text recognition: Microsoft Azure AI Vision and Mistral OCR. For the same company, one key is shared between speech recognition, text recognition and My AI models.
  • Two more on-device speech models (experimental): Paraformer and Qwen3-ASR.
  • Quiet speech: local models read soft recordings after noise reduction, when a steady room noise is close to the voice.
  • Vocabulary packs, text tidying, AI polish, and "My AI models" for polishing and translation.

Full list: CHANGELOG.md

Install

  1. Download Cadenza-1.2.0-macos-universal.dmg, open it and drag Cadenza to Applications. (The zip Cadenza-1.2.0-macos-universal.zip holds the same app, for people who prefer it.)
  2. Open Cadenza from Applications. macOS blocks the first launch because the app is not notarized: open System Settings → Privacy & Security, scroll to the message about the app and choose Open Anyway.
  3. Allow Microphone, Accessibility and Input Monitoring when asked (Screen & System Audio Recording only if you use screenshots).

Requires macOS 14 or later, on an Apple silicon or Intel Mac.

Upgrading from 1.1.0: macOS asks for the permissions again after an update. If dictation stops typing after the update, turn Accessibility off and on again for Cadenza in System Settings → Privacy & Security, then restart the app.

Verify

The checksum file covers the disk image and the zip.

shasum -a 256 -c SHA256SUMS.txt

Third-party licenses (sherpa-onnx, ONNX Runtime and the code they contain) are inside the app, in Contents/Resources/Licenses.

Known limitations

  • Tested by CI, not by hand, on Intel and on macOS 14 and 15. The self-test suites run on GitHub's machines: Apple silicon with macOS 14, 15 and 26, and Intel with macOS 15 and 26. Nobody has used this build by hand on an Intel Mac or on macOS 14 or 15 yet. Please report what you find.
  • The new cloud services were not tested with real accounts. OpenAI, Groq, Google Cloud, Microsoft Azure, AssemblyAI, ElevenLabs, Microsoft Azure AI Vision and Mistral OCR were checked with scripted replies (request shapes follow each vendor's documentation). Only iFLYTEK and Deepgram were tested online, with synthesized speech. Feedback from real accounts is very welcome.
  • Not notarized; macOS asks for the permissions again after each update.
  • Compatibility with every app is not verified.
  • Interface in English and Simplified Chinese only.

中文

早期预览。 随言 1.2.0 增加了有独立快捷键的语音翻译、更多云端语音识别和文字识别服务、两个新的本机语音模型,以及对轻声的处理。安装包支持 Apple 芯片和 Intel 的 Mac,要求 macOS 14 及以上。

这一版有什么

  • 语音翻译有自己的快捷键和页面(设置 → 翻译)。普通听写不再翻译。原来选了语言的,请设置一个翻译快捷键,才会继续翻译。
  • 快捷键:“按住说话”和“点按开始和结束”各有一行,各自有开关和按键。“触发方式”的选择已去掉。
  • 更多云端语音识别:OpenAI、Groq、Google Cloud、Microsoft Azure、AssemblyAI、ElevenLabs,以及任何兼容 OpenAI 转录接口的服务。网站上每一项都有配置教程。
  • 更多云端文字识别:Microsoft Azure AI Vision、Mistral OCR。同一家公司的密钥在语音识别、文字识别和“我的 AI 模型”之间共用,只需填写一次。
  • 两个新的本机语音模型(实验性):Paraformer 和 Qwen3-ASR。
  • 轻声识别:本机模型在环境噪声接近人声时做降噪处理。
  • 词库、文字整理、AI 润色,以及用于润色和翻译的“我的 AI 模型”。

完整列表见 CHANGELOG.md。

安装

  1. 下载 Cadenza-1.2.0-macos-universal.dmg,打开后把随言拖进“应用程序”。(zip Cadenza-1.2.0-macos-universal.zip 里是同一个应用,给喜欢压缩包的人。)
  2. 从“应用程序”里打开随言。因为没有公证,macOS 第一次会拦截:到 系统设置 → 隐私与安全性,下拉找到关于这个应用的提示,点 仍要打开。
  3. 按提示允许麦克风、辅助功能和输入监控(使用截图时还需要“录屏与系统录音”)。

要求 macOS 14 及以上,Apple 芯片或 Intel 的 Mac 都可以。

从 1.1.0 升级:更新后 macOS 会再次询问权限。如果更新后听写无法输入文字,请到 系统设置 → 隐私与安全性,把随言的辅助功能关掉再打开,然后重启应用。

校验

校验文件同时包含 .dmg 和 .zip。

shasum -a 256 -c SHA256SUMS.txt

第三方许可证(sherpa-onnx、ONNX Runtime 及其包含的代码)在应用内的 Contents/Resources/Licenses。

已知限制

  • Intel 和 macOS 14、15 目前只经过 CI 测试,没有人亲手用过。 自测套件在 GitHub 的机器上运行:Apple 芯片的 macOS 14、15、26,以及 Intel 的 macOS 15、26。欢迎反馈你遇到的情况。
  • 新增的云端服务没有用真实账号测试。 OpenAI、Groq、Google Cloud、Microsoft Azure、AssemblyAI、ElevenLabs、Microsoft Azure AI Vision 和 Mistral OCR 只用模拟回复测试过(请求格式按各家文档)。只有讯飞和 Deepgram 联网测试过(用合成语音)。欢迎真实账号的反馈。
  • 没有公证;每次更新后 macOS 会再次询问权限。
  • 没有验证所有应用的兼容性。
  • 界面只有英文和简体中文。

Cadenza 1.1.0 (early preview)

Choose a tag to compare

@DragonKingIO DragonKingIO released this 07 Oct 14:57
a4bb4cd

Early preview. Cadenza now runs on Intel Macs as well as Apple silicon, and on macOS 14 or later (1.0.0 needed macOS 26).

中文说明见下方。

What's new

  • One package for both chips, and macOS 14 as the minimum system. Where an interface exists only on newer systems, the app uses a plain alternative.
  • Recording bar: see-through glass in light mode, and it matches the Settings preview in dark mode. An optional character style shows a short animation instead of the waveform (the animations are used with the author's permission; they are not covered by the MIT license, see THIRD_PARTY_NOTICES.md).
  • On-device text recognition models for screenshots: download a PP-OCR model set (PP-OCRv5 mobile is recommended) in Settings → Text Recognition.
  • Cancelling a model download no longer leaves an empty staging folder behind.

Full list: CHANGELOG.md

Install

  1. Download Cadenza-1.1.0-macos-universal.zip and unzip it.
  2. Open the app. macOS blocks the first launch because the app is not notarized: open System Settings → Privacy & Security, scroll to the message about the app and choose Open Anyway.
  3. Allow Microphone, Accessibility and Input Monitoring when asked (Screen & System Audio Recording only if you use screenshots).

Requires macOS 14 or later, on an Apple silicon or Intel Mac.

Verify

shasum -a 256 -c SHA256SUMS.txt

Third-party licenses (sherpa-onnx, ONNX Runtime and the code they contain) are inside the app, in Contents/Resources/Licenses.

Known limitations

  • Not yet verified on a real Intel Mac, or on macOS 14 and 15. It is tested on Apple silicon with macOS 26, and the Intel half of the package was run under Rosetta (Apple Vision text recognition cannot run there). Please report what you find.
  • Not notarized; macOS asks for the permissions again after each update.
  • Compatibility with every app is not verified. Of the cloud providers, only iFLYTEK and Deepgram were tested online (with synthesized speech).
  • Interface in English and Simplified Chinese only.

中文

早期预览。 随言现在除了 Apple 芯片,也支持 Intel Mac,最低系统从 macOS 26 降到 macOS 14。

这一版有什么

  • 一个安装包同时支持两种芯片,最低 macOS 14。在较旧系统上没有的界面效果,会改用普通样式。
  • 录音条:浅色模式下是透明玻璃,深色模式下和设置里的预览一致;可选的角色样式会用一小段动画代替波形(动画经作者许可使用,不在 MIT 许可范围内,见 THIRD_PARTY_NOTICES.md)。
  • 本地文字识别模型(用于截图):在 设置 → 文字识别 里下载 PP-OCR 模型(推荐 PP-OCRv5 mobile)。
  • 取消模型下载后不再留下空的暂存目录。

安装

  1. 下载 Cadenza-1.1.0-macos-universal.zip 并解压。
  2. 打开应用。因为没有公证,macOS 第一次会拦截:到 系统设置 → 隐私与安全性,下拉找到关于这个应用的提示,点 仍要打开。
  3. 按提示允许麦克风、辅助功能和输入监控(使用截图时还需要“录屏与系统录音”)。

要求 macOS 14 及以上,Apple 芯片或 Intel 的 Mac 都可以。校验:shasum -a 256 -c SHA256SUMS.txt。第三方许可证在应用内的 Contents/Resources/Licenses。

已知限制

  • 还没有在 Intel 真机以及 macOS 14、15 上验证。 目前在 Apple 芯片、macOS 26 上测试;Intel 部分只在 Rosetta 下运行过(Apple Vision 文字识别在 Rosetta 下无法运行)。欢迎反馈你遇到的情况。
  • 没有公证;每次更新后 macOS 会再次询问权限。
  • 没有验证所有应用的兼容性;云端服务里只有讯飞和 Deepgram 联网测试过(用合成语音)。
  • 界面只有英文和简体中文。

Cadenza 1.0.0 (early preview)

Choose a tag to compare

@DragonKingIO DragonKingIO released this 07 Oct 08:02

Early preview. The first public build of Cadenza: open-source voice input for macOS. It is used every day on the maintainer's Mac but has not been tested on many setups yet. Bug reports are very welcome.

中文说明见下方。

What's new

  • Voice input. Hold Left Option (or Right Option, another shortcut, or tap to start / tap to stop), speak, and the text is typed at your cursor. If it cannot be typed, it is kept for you to copy.
  • Local recognition with SenseVoice (recommended), FireRedASR2 and Parakeet (both experimental), downloaded inside the app and checked against their SHA-256. "Compare models with my voice" shows which one understands you best.
  • Your own cloud service, optionally: iFLYTEK, Volcengine, Tencent Cloud, Alibaba Cloud, Baidu and Deepgram, with your own keys. Audio is sent only after you agree, provider by provider.
  • Screenshots and text recognition (new, still being tested): capture, mark up, pin, and copy the text or QR code in a picture.
  • Local API for your own programs and hardware, off by default, loopback only, one revocable token per device.
  • Private by design: no account, no analytics, no crash reports; a "never go online" switch. Update checks only when you ask, or weekly if you agree.

Full list: CHANGELOG.md

Install

  1. Download Cadenza-1.0.0-macos-arm64.zip and unzip it.
  2. Open the app. macOS blocks the first launch because the app is not notarized: open System Settings → Privacy & Security, scroll to the message about the app and choose Open Anyway.
  3. Allow Microphone, Accessibility and Input Monitoring when asked (Screen & System Audio Recording only if you use screenshots).

Requires macOS 26 or later on Apple silicon.

Verify

shasum -a 256 -c SHA256SUMS.txt

Third-party licenses (sherpa-onnx, ONNX Runtime and the code they contain) are inside the app, in Contents/Resources/Licenses.

Known limitations

  • Not notarized; macOS asks for the permissions again after each update.
  • Compatibility with every app is not verified. Of the cloud providers, only iFLYTEK and Deepgram were tested online (with synthesized speech).
  • Interface in English and Simplified Chinese only.
  • If you used a development build: the internal identifiers changed to Cadenza. The first launch moves your settings, models and keys; grant the permissions again.

中文

早期预览。 随言的第一个公开版本:开源的 macOS 语音输入。维护者每天都在用,但还没在很多环境里测试过,欢迎反馈问题。

这一版有什么

  • 语音输入:按住左 Option(也可以改成右 Option、其他快捷键,或点一下开始、再点一下结束)说话,文字写到光标处;写不进去时会保留下来方便复制。
  • 本地识别:SenseVoice(推荐)、FireRedASR2、Parakeet(后两个为实验性),在应用内下载并校验 SHA-256;“用我的声音比较模型”告诉你哪个最听得懂你。
  • 云端服务,自带账号,可选:讯飞、火山引擎、腾讯云、阿里云、百度、Deepgram;只有你对该服务商同意之后才会发送音频。
  • 截图与文字识别(新功能,仍在测试):截图、标注、贴在屏幕上,取出图里的文字和二维码。
  • 本地接口:给你自己的程序和硬件用,默认关闭,只监听本机,每个设备一个可撤销的令牌。
  • 隐私是设计前提:没有账号、没有统计、没有崩溃上报,有“永不联网”开关;只有你主动检查,或同意每周检查时,才会查看更新。

安装

  1. 下载 Cadenza-1.0.0-macos-arm64.zip 并解压。
  2. 打开应用。因为没有公证,macOS 第一次会拦截:到 系统设置 → 隐私与安全性,下拉找到关于这个应用的提示,点 仍要打开。
  3. 按提示允许麦克风、辅助功能和输入监控(使用截图时还需要“录屏与系统录音”)。

要求 macOS 26 及以上、Apple 芯片。校验:shasum -a 256 -c SHA256SUMS.txt

第三方许可证(sherpa-onnx、ONNX Runtime 及其包含的代码)在应用内的 Contents/Resources/Licenses。

已知限制

  • 没有公证;每次更新后 macOS 会再次询问权限。
  • 没有验证所有应用的兼容性;云端服务里只有讯飞和 Deepgram 联网测试过(用合成语音)。
  • 界面只有英文和简体中文。
  • 用过开发版的话:内部标识已改为 Cadenza,首次启动会搬移设置、模型和密钥,需要重新授予权限。