Repository navigation
Releases: v2matosevic/Talkty
Release list
Talkty 1.5.0: clearer dictation, searchable history and ADE drafts
Talkty 1.5.0 improves the dictation window and removes avoidable work while recording.
- Clearer microphone action, recording state and selected local/cloud model.
- A larger output-destination control with accurate cursor, clipboard, history-only or ADE draft feedback.
- Search history by final text or original spoken transcription, with distinct no-match and empty-history messages.
- Pause checks use the existing audio buffer rather than copying it. A focused 1,000-check benchmark eliminated 51.2 MB of temporary allocations. This does not establish a faster overall transcription distribution.
- Obsolete early recognition cancels when speech resumes. Full recordings, the selected model and exact final-audio matching remain intact. An early cloud request may already have been accepted and billed when cancelled.
- Optional dictation to a named Hephaestus ADE agent draft. The complete instruction is staged for review without replacing the clipboard or automatically sending a turn. Unverified handoffs stay encrypted locally and retry with their original destination and operation ID.
- Command delivery reconciles operation receipts after lost responses rather than blindly repeating an action.
- Updated privacy and command-integration documentation, including the installer's ADE data-sharing disclosure.
Direct draft delivery needs the matching protocol-v1 ADE capture receiver. ADE 0.0.108 does not include that receiver; ordinary Talkty dictation remains available. Durable command receipts require the matching Hermes update. This installer does not install ADE, Athena Desktop or Hermes.
Upgrade by running TalktySetup-1.5.0.exe over the existing installation. Settings, history, recovery files, downloaded models and an existing CUDA pack are preserved. Existing all-users installations may require Windows administrator approval. The installer is not code signed. The optional CUDA pack is unchanged.
The committed source passed the full 405-test Release suite and exact-source GitHub CI. Production WPF layouts were rendered and inspected at the minimum and larger window sizes. Clean-machine installation and live microphone/desktop-input behavior are not established by these checks. No private recording or paid model request was used for packaging verification.
Compiled source: 8698fb6c13e27df5ffffba01d6d026f99ecabbd5.
Installer SHA-256: a0b7380a96388a338e5e7c1563d70177464c1cfebe3fae12eea464edac3a8421.
Release and installation evidence · Performance and appearance · Privacy
Talkty 1.4.0
Talkty 1.4.0 reduces waiting while preserving full-recording context and the selected transcription model.
- Start transcription during natural pauses; use the early result only if final audio matches exactly. Configure it in Settings > Behavior. Cloud can bill one extra early pass, including cancelled recordings.
- Faster CUDA attention. Two English fixtures on an RTX 3060 Ti measured 13–24% lower median decoding time; cloud timing varied. These are bounded measurements, not a guarantee.
- Fixed MAI cloud requests failing with provider HTTP 400 while retaining native Opus, vocabulary hints and clean output.
- Restored the opt-in Prompting switch in Settings > Cloud & Prompting. It stays off by default and adds a paid text request and delay when enabled.
- Optional prompt fidelity checks and advanced prompt planning. Planning defaults off; fidelity Record only can make a bounded paid text request after a generated prompt is delivered.
- Optional command mode with a separate shortcut, live local preview, service discovery and results on the recording pill. A compatible local command service is required and is not included.
- Removed the blocking command-preview stop, improved cancellation/shutdown ownership, and verified the uninstaller's Keep/Remove data choices using isolated fixtures.
Upgrading preserves settings, history, recovery recordings, downloaded models and the optional CUDA pack. Standard installation includes Vulkan and native Opus; CUDA remains a separate optional download.
389 local tests passed. The source, exact package, Windows installation and CI evidence are documented in the release record. See latency measurements and limits and the command integration requirements.
Talkty 1.3.5: cloud recovery and selectable backup
Cloud transcription failures no longer make your recording disappear.
- Failed cloud recordings stay encrypted on your PC, survive app restarts, and appear in Talkty with Retry and Discard controls. A new recording never replaces an older failed take.
- Choose an automatic backup model in Settings > Cloud & Prompting. Qwen3 ASR Flash is the default backup; Off and other cloud models are available. Your primary model stays unchanged.
- Temporary provider errors switch directly to a compatible backup instead of waiting and sending another request to the same busy provider. Authentication, insufficient balance, cancellation and no-speech results do not trigger a backup. Backup requests cost per use.
- Added Qwen3 ASR 1.7B as a separate cloud option. Qwen Flash keeps MP3 uploads for accuracy; MAI keeps native Opus, vocabulary hints and clean transcription.
- Retrying from Talkty copies the recovered transcript when clipboard output is enabled, without auto-pasting into an old window. The recording stays recoverable if saving history fails.
- Upgrading preserves settings, transcription history, downloaded models, recovery recordings and the optional CUDA pack.
The default Qwen Flash backup supports the catalog's 11 languages, including English, but not Croatian or Serbian. Unsupported explicit languages and identical primary/backup selections are skipped. Previously overwritten recordings cannot be recovered retroactively. No new speed or accuracy guarantee is claimed.
Validation: 143 automated tests passed locally. Full behavior and verification: https://github.com/v2matosevic/Talkty/blob/main/docs/RELEASE-1.3.5.md
Talkty 1.3.4
MAI-Transcribe 2 now uploads less audio data and starts preparing the connection earlier.
- Native Opus compression made requests 56–57% smaller on two test recordings, with identical transcripts. Smaller uploads help reduce sensitivity to a slow uplink; this is not a claim that transcription is 56–57% faster.
- Connection and encoder warm-up run concurrently and start when the cloud model becomes ready.
- Your vocabulary hints and clean transcription style are preserved. Other cloud models keep MP3, and MP3/WAV fallback remains available.
- Detailed timings make remaining cloud delays easier to diagnose.
- Upgrading keeps settings, history, downloaded models and the existing optional CUDA pack.
Network and provider response times can still vary. The local build's three real dictations reached paste in 1.25, 0.62 and 0.57 seconds; that small sample does not establish long-term consistency.
Talkty 1.3.3: MAI-Transcribe 2 and faster cloud transcription
Talkty 1.3.3 makes Microsoft's MAI-Transcribe 2 the recommended cloud model and makes cloud transcription much faster.
Changes
- New recommended cloud model: MAI-Transcribe 2, via OpenRouter. It is the cheapest cloud option ($0.10 per hour of audio), fast, covers 60 languages and leaves out filler words like "um". It does not support Croatian or Serbian; pick another cloud model or local Whisper for those.
- Faster cloud transcription. Recordings upload as compressed MP3 instead of WAV, about a fifth of the size. In testing, a 2.4-minute recording went from 10 to 26 seconds down to about 3, and short dictations took roughly half as long. The connection also warms up while you are still speaking.
- Your vocabulary now reaches MAI-Transcribe 2 as spelling hints: the first 50 words in Settings > Vocabulary, with the ones you added first.
- A cloud recording with no speech in it now says "No speech was detected" instead of showing an error.
- GPT-4o Mini Transcribe is no longer the recommended cloud model, and Qwen3 ASR Flash is no longer labelled the cheapest.
Install or upgrade
Download TalktySetup-1.3.3.exe and run it. Existing settings, history, models and an installed CUDA pack are kept during an upgrade. Vulkan GPU support is included; the optional NVIDIA CUDA pack remains available from Settings.
Windows 10/11, x64. Local transcription remains the default; cloud transcription and Prompting are opt-in and need an OpenRouter API key. MIT-licensed source is included with this release.
Verification
120 tests passed in Release. All 488 installed payload files were verified during installation. Live checks against OpenRouter: all six cloud models accepted the MP3 uploads, and MP3 produced the same words as WAV on test recordings (2 of 302 words differed on a 2.4-minute clip, one of them a correction). Speed and accuracy were measured with synthesized English speech; natural dictation and Croatian speech were not measured.
Installer SHA-256:
6cb7b29020c268214eb26b90f9c4ff25f0fbf0e4724ccf547f95fe8c9357c95e
Talkty 1.3.2: vocabulary, recording reliability, and desktop improvements
Talkty 1.3.2 brings the recording reliability, vocabulary, and desktop improvements developed since 1.3.0.
Changes
- Words added in Settings now reach local Whisper transcription in English, with personal terms prioritized over the starter vocabulary. Changes apply to the next recording and survive idle model reloads.
- Cancelling recording, transcription, or Prompting stops pending output. Repeated recordings no longer mix late audio from an earlier recording.
- Clearer recording and cancellation controls, keyboard-accessible history, better resizing, and more useful transcription feedback.
- History saves stay in order. Audio meter updates remain responsive when the window is busy, and microphone testing cannot overlap dictation.
- Silent recordings and empty results leave the clipboard alone. Microphone and clipboard failures are visible.
Install or upgrade
Download TalktySetup-1.3.2.exe and run it. Existing settings, history, models, and an installed CUDA pack are retained during an upgrade. Vulkan GPU support is included; the optional NVIDIA CUDA pack remains available from Settings.
Windows 10/11, x64. Local transcription remains the default; cloud transcription and Prompting are opt-in. MIT-licensed source is included with this release.
Verification
99 tests passed in Release. All 488 installed payload files were verified during installation. Before publication, all 489 publish files (including the debug symbol file excluded from the installer) matched the recorded manifest; 64 source-document checksums and the compiled WPF resources matched current source. Recognition accuracy on a reference recording corpus and mixed-monitor native interaction were not measured.
Installer SHA-256:
2a54025499ce41f20702303f41394564973fb65c0bbb4f2e3acdc419fd69c2f8
Talkty 1.3.0 — 88% smaller, races fixed, honest auto-paste
The installer went from 484 MB to 58 MB, and two long-standing reliability bugs found in real usage logs are gone.
Smaller
- CUDA is no longer bundled. GPU acceleration still works out of the box — on NVIDIA, AMD and Intel GPUs — via the bundled Vulkan backend. NVIDIA users who want maximum speed can install the optional CUDA pack from Settings → Behavior (a one-time ~450 MB download with resume support, then one restart).
- Upgrading keeps your CUDA. If a previous version installed the CUDA files, they stay in place and Talkty keeps using them — nothing to redo.
- Dropped 128 MB of Linux/ARM binaries that could never load on Windows anyway.
More reliable
- Fixed: a dictation could fail right after the app sat idle. After the 15-minute memory unload, the first transcription could hit an internal race ("Cannot dispose while processing") and get eaten. The engine now builds with your language up front and properly serializes its warmup — verified against weeks of real logs.
- Fixed: volume ducking errored on nearly every recording. A garbage-collection bug killed the smooth fade-restore mid-fade (an instant fallback masked it). The fade now survives, and the log spam is gone.
- Auto-paste tells you when it can't paste — and why. Apps running as administrator silently swallow synthetic keystrokes (a Windows security rule); Talkty now detects this properly and shows a toast instead of failing silently. Your text is always on the clipboard either way.
New (both opt-in, in Settings → Behavior)
- Recording pill near the text cursor — the pill can follow where you're typing instead of sitting at the bottom of the screen. Bottom-center of your active monitor stays the default.
- Restore previous clipboard after auto-paste — dictating no longer eats what you had copied. Talkty puts your old clipboard back after the paste lands (and never clobbers a newer copy).
Under the hood
- CUDA pack installs are atomic (a crash mid-install can never leave a corrupted DLL) and downloads resume where they left off.
- App restart after installing the pack is race-free.
- Startup GPU detection no longer logs an error on every launch; dependency security bump (SharpCompress 1.0).
Upgrade: just run the installer over your existing installation — settings, history, models and CUDA files stay put.
CUDA runtime pack (cu13)
CUDA runtime pack for Talkty — not an app release.
Since v1.3.0 the Talkty installer no longer bundles the CUDA runtime (it was ~75% of the download). GPU acceleration works out of the box on NVIDIA, AMD and Intel GPUs via the bundled Vulkan backend; this optional pack gives NVIDIA GPUs maximum transcription speed via CUDA.
You normally don't download this file yourself — Talkty offers it in Settings → Behavior when it detects an NVIDIA GPU, downloads it with resume support, installs it, and asks for a restart.
Contents: cuBLAS / CUDA runtime DLLs (CUDA 13.1) + whisper.cpp CUDA backend, laid out exactly as the app expects (cublas*/cudart* in the app root, ggml CUDA natives in runtimes/cuda/win-x64).
This tag is version-independent: app releases point here so the ~450MB asset is uploaded once, not per release.
Talkty 1.2.2 — Settings redesign
A ground-up redesign of Settings plus a round of accuracy and quality-of-life fixes. (This release bundles the unreleased 1.2.1 UI work — full details for both are in the CHANGELOG.)
Settings, redesigned
- Left navigation with six clear sections — Model, Audio & Language, Vocabulary, Cloud & Prompting, Behavior, Hotkey — replaces one long scroll. The window is now resizable.
- Text replacements are finally editable in the app: one
misheard => correctrule per line under Vocabulary, with a Reset to defaults. (If you relied on the oldcloud => Claudedefault removed in 1.2.0, re-add it here.) - Your OpenRouter API key is masked by default, with an eye toggle to reveal it.
- Dark themed scrollbars, text boxes, and dialogs everywhere — no more white OS controls breaking the theme.
History
- Hover an entry for copy and delete buttons; a Clear action wipes the list.
Accuracy & reliability
- Punctuation cleanup no longer corrupts abbreviations — "e.g. the config" stays "e.g. the config".
- Cloud transcription retries once on transient errors (rate limits, gateway hiccups).
- The early streamed clipboard copy now matches the final text's cleanup.
Discoverability
- Hover the recording pill for the "esc cancels" hint.
- Onboarding fixed: visible step numbers, current model recommendations, and tips for auto-paste, ESC, and the Prompting sparkle.
Upgrading: run the installer over your existing install — settings are kept.
Talkty 1.2.0 — Lighter, more accurate, more honest
The app now behaves like a good tray citizen and tells you when something goes wrong.
Lighter on your PC
- The speech model's memory is freed after 15 minutes of inactivity. A large model no longer holds 1.5–3 GB of RAM around the clock while Talkty sits in the tray. It reloads automatically while you speak, so the next recording works exactly as before. Toggle: Settings → Behavior → "Free memory when idle" (on by default).
- Faster startup (GPU detection no longer blocks the launch path) and much less disk activity from logging.
More accurate transcription
- Removed the default
cloud→Claudeandsequel→SQLreplacements — they rewrote legitimate speech ("AWS cloud" became "AWS Claude"). Re-add them in your vocabulary if you want them. - The English coding-vocabulary prompt is no longer applied when transcribing other languages or using auto-detect — it was biasing non-English decoding toward English.
- Whisper's temperature fallback is re-enabled, so a decode stuck in a repetition loop recovers instead of producing garbage.
Prompting hardening
- Long dictations are never silently cut off: the output ceiling was raised 4× and a truncated result escalates to the next model instead of shipping incomplete.
- An invalid or out-of-credits OpenRouter key fails fast with a clear notification instead of silently trying all four models.
- The recording pill shows a purple Prompting… state while the AI expands your dictation.
You always know what happened
- Failed transcriptions and Prompting fallbacks now notify you — including as a tray balloon when the window is hidden.
- Fixed the main window showing red "Model Not Loaded" while a model was actually loading.
- Fixed: saving Settings silently reset custom text replacements to defaults on the next launch.
Full details in the CHANGELOG. Upgrading: run the installer over your existing install — settings are kept.