Skip to content

Talkty 1.3.3: MAI-Transcribe 2 and faster cloud transcription

Choose a tag to compare

@v2matosevic v2matosevic released this 15 Sep 22:27

Talkty 1.3.3 makes Microsoft's MAI-Transcribe 2 the recommended cloud model and makes cloud transcription much faster.

Changes

  • New recommended cloud model: MAI-Transcribe 2, via OpenRouter. It is the cheapest cloud option ($0.10 per hour of audio), fast, covers 60 languages and leaves out filler words like "um". It does not support Croatian or Serbian; pick another cloud model or local Whisper for those.
  • Faster cloud transcription. Recordings upload as compressed MP3 instead of WAV, about a fifth of the size. In testing, a 2.4-minute recording went from 10 to 26 seconds down to about 3, and short dictations took roughly half as long. The connection also warms up while you are still speaking.
  • Your vocabulary now reaches MAI-Transcribe 2 as spelling hints: the first 50 words in Settings > Vocabulary, with the ones you added first.
  • A cloud recording with no speech in it now says "No speech was detected" instead of showing an error.
  • GPT-4o Mini Transcribe is no longer the recommended cloud model, and Qwen3 ASR Flash is no longer labelled the cheapest.

Install or upgrade

Download TalktySetup-1.3.3.exe and run it. Existing settings, history, models and an installed CUDA pack are kept during an upgrade. Vulkan GPU support is included; the optional NVIDIA CUDA pack remains available from Settings.

Windows 10/11, x64. Local transcription remains the default; cloud transcription and Prompting are opt-in and need an OpenRouter API key. MIT-licensed source is included with this release.

Verification

120 tests passed in Release. All 488 installed payload files were verified during installation. Live checks against OpenRouter: all six cloud models accepted the MP3 uploads, and MP3 produced the same words as WAV on test recordings (2 of 302 words differed on a 2.4-minute clip, one of them a correction). Speed and accuracy were measured with synthesized English speech; natural dictation and Croatian speech were not measured.

Installer SHA-256:
6cb7b29020c268214eb26b90f9c4ff25f0fbf0e4724ccf547f95fe8c9357c95e