Releases: Budanovvv/Dictate
Release list
Dictate 3.2.3
The pill used to pick the screen under the mouse cursor — a proxy for
"where the user is" that breaks with two monitors: typing into PyCharm
on one display with the cursor parked over Chrome on the other put the
pill on the wrong screen.
Ask the WindowServer for the frontmost app's front normal window and use
its screen; the mouse remains the fallback. A CGWindowList query rather
than AX into the target app, so a busy or lazily built accessibility
tree can never stall show().
Full Changelog: v3.2.2...v3.2.3
Dictate 3.2.2
Dictate 3.2.2 is a day of fixes from three live calls, all in the meetings window.
-
The Sources list holds still. During a recording, the sidebar's platform rows — Google Meet, Zoom — could swap places every few seconds as the transcript grew. Ties are now ordered by name, so the list never dances again.
-
Cleaning up voices takes one rename, not three. Voice separation goes by sound, and sound can split one person in two — plus a collective "Call" bucket for lines no voice could be pinned to. Now, naming a voice mid-call no longer switches off the automatic merge of its splinters, and the moment every voice of the call side answers to one name, the collective lines follow it on their own.
-
The outline's three levels are three levels again. A meeting's contents quietly defaulted to the sparsest cut, and the Fewer / Standard / More control could end up switching between three identical views of it. The default is the middle again, and each level now shows its own depth.
Download Dictate-3.2.2.dmg below — the -update file is the in-app updater's payload, not for manual installs.
Full Changelog: v3.2.1...v3.2.2
Dictate 3.2.1
Dictate 3.2.1 fixes what 3.2 shipped broken, and starts telling the truth in three places it used to stay silent.
-
The transcript's ⋯ menu works again. In 3.2 the three dots above a transcript did nothing — the menu refused to open while the app was not frontmost, which for a menu-bar app is almost always. It now opens the way the status-bar menu always has.
-
Calls file under their platform. The moment the detection card names a call — Google Meet, Zoom — that name now travels into the recording. Before, every call landed in "Other browser calls" no matter what the card had just said. The transcript's summary also no longer trips over the platform note inside the file.
-
Voice separation says what it is. Splitting a call into voices goes by sound, and sound can show one person as two. The first multi-voice transcript now says so, once — click a name to rename a voice, give two voices the same name to merge them — and a small info mark by the Voices header keeps the explanation reachable later. Summaries and answers rely on what was said, not who said it.
-
You can always tell what you are running. About and Settings now show the exact build (commit count, hash, dirty marker), a downloaded update announces itself in the menu instead of waiting invisibly, and "Check for updates" no longer answers "you're up to date" while a newer version already sits staged on disk.
Download Dictate-3.2.1.dmg below — the -update file is the in-app updater's payload, not for manual installs.
Full Changelog: v3.2...v3.2.1
Dictate 3.2
Dictate 3.2 puts you in charge of what the app is allowed to do — and gives your meetings an agent.
Everything about meetings is now opt-in. Four independent switches — notice calls, record the call audio, separate the voices, let the model read your meetings — all off on a fresh install. The meetings window opens on a setup page that says what each one adds; turn on one, all three, or none. Existing installations keep everything on, exactly as before.
Ask becomes your Agent. It has its own place at the top of the sidebar and says plainly what it has: "Has read all 38 meetings", "Reading 12 of 38…", or "Off — has read nothing yet". Ask it anything across everything you have recorded; every answer names the meeting and the moment it came from.
Honesty instead of silence. A transcript recorded with capabilities off says exactly what it lacks and what can still be done: a summary can be written once, on request, without turning anything on; the other side of an unrecorded call is stated as unrecoverable rather than papered over. If macOS denies the microphone, the switches say they are waiting instead of pretending to work.
Windows behave. Settings opens centred over the meetings window, on whatever Space you are on. The manual update check answers with a quiet card instead of a modal alert that could freeze the app. Recording always starts as the small pill — from the call card or the menu — and the full window opens only by your hand.
Under the hood: the whole app compiles clean under Swift strict concurrency; call detection probes left the main thread, so a busy browser can no longer stall dictation; and a long list of review findings is fixed, from a leaked sleep-blocker to a level meter that finally breathes with your voice.
Download Dictate-3.2.dmg below — the -update file is the in-app updater's payload, not for manual installs.
Full Changelog: v3.1...v3.2
Dictate 3.1
Dictate 3.1 makes recording a call something you no longer have to remember.
Dictate notices your calls. The moment a recognisable call starts — Zoom, Teams, FaceTime, Webex, or Google Meet in a browser tab — a small card drops from the top of the screen: Record or Not this one. It never records on its own; the offer is automatic, the recording is always your click. The card follows you across Spaces and full-screen apps, waits until you answer, and dismisses itself only if the call itself ends. Declining a call stays declined for that call.
Recording starts small. Accepting the card starts the compact recording pill, not the full window — you are on a call, and the pill now follows you across displays too (drag it somewhere and it stays put).
A forgotten recording ends itself. If the call's app has released the microphone and the air has stayed quiet, the recording stops about two minutes after everyone left; with no recognisable call around at all, ten silent minutes end it. While a call is live it never stops itself, however long the silence — a quiet hold is not an ending. The transcript's last line says which rule stopped it.
A quick Record button now sits at the top of the meeting list for the calls Dictate can't detect — and deleting a meeting works like deleting a conversation: right-click → Delete, with Undo and ⌘Z, and the file goes to the Trash rather than to nothing.
Download Dictate-3.1.dmg below — the -update file is the in-app updater's payload, not for manual installs.
Full Changelog: v3.0...v3.1
Dictate 3.0
Dictate 3.0 is the release where the whole app grew into its design — every surface was rebuilt against one visual system, and the product now does four things, all on your Mac.
A meeting library, not just transcripts. Every recording gets a title and a one-line summary written by a local model, a clickable outline of the moments that mattered (two levels for ordinary meetings, three past forty minutes), per-voice speaking time with inline rename, tags, stars, and the platform the call ran on. Search covers every spoken word, title, summary and outline line.
Ask — an optional agent over your archive. Connect Claude or ChatGPT with your own API key and ask questions across everything you've recorded; the agent lists, searches and reads your transcripts itself, and every answer names the meetings it drew from. Off by default, and the only feature that ever talks to a server — the recordings never leave your Mac.
Titles and summaries in the meeting's own language. A Russian meeting gets a Russian summary; the earlier English-always behavior is gone.
A calmer, designed surface everywhere. One design system across the meetings window, dictation overlay, recording pill, menu bar and onboarding; light and dark themes with a hand-tuned dark palette; adjustable transcript text size; resizable and collapsible panels; a rebuilt onboarding that walks the model download, keys and permissions honestly.
Quality under the hood. A second voice that is really a fragment of the first now merges back into it; a call channel's language is pinned once it is stable, so bilingual calls stop flip-flopping; the recording HUD levels breathe instead of twitching.
Download Dictate-3.0.dmg below — the -update file is the in-app updater's payload, not for manual installs.
Full Changelog: v2.6.1...v3.0
Dictate 2.6.1
A meeting that has just been named can be acted on. Naming a
transcript renames its file, and the window went on pointing at the old
path — so Show in Finder revealed nothing at all, and Rename and Move to
Trash were aimed at a file that no longer existed. The window now
follows the transcript to its new name.
The debug log keeps the beginning of a long session. It dropped
itself and started over once it grew past a couple of megabytes, which
one long meeting is enough to do — so the part worth reading, the start,
was the part that disappeared. It now sets the full file aside and
starts a fresh one, and does it later.
Full Changelog: v2.6.0...v2.6.1
Dictate 2.6.0
Meetings you can read, search, and get out of the way of.
One person stays one person. The model that finds speaker changes can only
see ten seconds at a time, so every longer window was cut into pieces — and the
last piece, padded out with silence, was embedded as a stranger. On a real
forty-minute call that split one voice into two near-equal "speakers" across
70% of windows, which is why a monologue arrived in the transcript alternating
between two names. Windows are now capped at what the model actually sees.
Measured on the same recording: 6% instead of 70%, and the talking split out
85/14/1 — which is what the meeting actually sounded like.
And when a sentence does get torn, the words decide who owns it. A line
that stops mid-phrase followed by one that opens in lower case is one person
being interrupted by the clock, not two people talking, and it is now filed
that way.
A recording can get out of the way. Fold the transcript into a pill — level
meter, clock, stop — with the yellow button or by closing the window; close the
pill and the menu bar is all that is left, where the mark breathes for as long
as the meeting runs. The menu brings the transcript back at any point.
Copy anything. Every turn has a copy button and a right-click menu — the
words alone, or the line with speaker and time — and ⌘C works in the window
without stealing focus from your call. Selecting text no longer fights the
transcript: it stops auto-scrolling the moment you scroll up or start a
selection, and a button brings you back to the newest line.
A transcript that reads like a conversation. Recognition cuts audio into
windows, and every window used to become its own line, so sentences broke
mid-phrase and hums stood alone as if they were turns. They are now joined back
into paragraphs — 1204 lines became 693 in a typical archive — while the file
on disk keeps every word exactly as it was recorded.
Contents and summaries. Each meeting gets a one-line summary of what it was
about and a table of contents with timestamps, written into the Markdown file
itself, so the folder is useful without the app.
Search that understands meaning. Ask for "legal" and you also get the
meeting where nobody said the word. Type a question in any language; it is
matched against what the meetings were about, and clicking a result opens the
transcript at that moment.
Titles and summaries are written on your Mac by a model you can choose to
download — 2.5 GB, one time, Apple Silicon and Intel alike. It replaces the
system model that only exists on macOS 26 and refuses to summarize anything it
deems sensitive, which in practice meant medical and legal conversations came
back blank. Nothing leaves the machine either way.
The library opens instantly. It used to read every transcript before
showing anything, on the thread that draws the app — and in an iCloud-synced
Documents folder a read can mean a download. Opening it as a call started
froze everything for sixteen seconds and cost the first twenty seconds of the
meeting.
The consent reminder appears once, before the first meeting you record,
instead of before every one.
Full Changelog: v2.4.0...v2.6.0
Dictate 2.4.0
Meeting transcripts: Dictate now transcribes your calls, entirely on
your Mac. Start a transcript from the menu bar during a browser call
and both sides are captured - your voice from the microphone, everyone
else from system audio - into a live Markdown file. Nothing leaves
your machine, and you confirm before every session starts.
-
Speakers are told apart automatically: participants get Speaker 1,
Speaker 2... labels via local voice diarization, and you can rename
anyone with a click - in the live session or in an old transcript. -
The transcript window is now a meetings library: past meetings
grouped by day, full-text search across words and speakers, and the
live session pinned on top while recording. -
Finished meetings name themselves: on macOS 26 an on-device Apple
model titles the transcript from what was said (translating first when
the meeting's language needs it), and files are named like
"2026-08-10 09.17 - Release planning.md" so Finder sorting stays
chronological. -
The live window feels alive: a gray in-progress line while a phrase
is still forming, level meters for both channels, and an honest
warm-up state instead of silence. -
Phrase boundaries are cut by the same neural voice detector as
dictation, so breaths and background noise never reach the recognizer,
and the live preview never delays the actual transcript.
Also fixed: the onboarding "Preparing the model" step now retries by
itself instead of spinning forever if the first attempt fails on a
fresh install.
Full Changelog: v2.3.6...v2.4.0
Dictate 2.3.6
Dictation is now pipelined: press the key while the previous phrase is
still being recognized and recording starts instantly - no more
swallowed presses between thoughts. Results always paste in the order
you spoke them, and with a quick rhythm the previous phrase lands in
your document while you're already dictating the next one, so the
recognition wait effectively disappears.
Switching between the plain and translate keys is seamless: pressing
one while holding the other hands the capture over on the spot, even
when the presses overlap by a few milliseconds, and the microphone
stays hot through the switch - not a syllable is lost at the boundary.
Esc now has clear scope: while recording it cancels just that capture,
otherwise it abandons everything still being recognized.
Also fixed: a quick accidental tap can no longer starve the next
recording's start, and a too-early verdict no longer shows a spurious
"didn't catch that" after a fraction-of-a-second tap.
Full Changelog: v2.3.5...v2.3.6