Skip to content

Releases: basal-john/phona

Phona 1.9.1

Choose a tag to compare

@basal-john basal-john released this 09 Sep 20:38
45d5041

Two fixes found by installing 1.9.0 and using it. Both are older than that release, so 1.9.0's DMG carries them.

The Models pane was wrong about the cloud

It reported "no cloud model is configured" on every default install, including while cloud corrections were being made. It read config.json, which does not carry a cloud_model key unless you have pinned one, and never asked the daemon, which merges that file over its own defaults and is the only thing that knows what is loaded.

The window now asks the daemon and falls back to the file, because it opens whether or not the engine is up. The cloud row also names the backend it goes through.

The phona command could be pointed into /tmp

install.sh wrote ~/.local/bin/phona from the install target, so any run with PHONA_HOME set repointed the real command at whatever directory was being tested. CI does exactly that, so running that step locally left the command pointing into /tmp, where it worked until /tmp was cleared.

A test install now gets its shim inside its own target and leaves the shared one alone. CI asserts both halves.

If your phona command is already pointing somewhere odd, this release does not repair it. Check with cat ~/.local/bin/phona and rerun ./install.sh to rewrite it.

Also

phona --probe-models prints which models the app believes are loaded next to what config.json pins, and says whether the daemon answered. It is how the first fault was confirmed without opening the window.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models. Building from source needs macOS 26, because the HUD's material only exists in that SDK.

Phona 1.9.0

Choose a tag to compare

@basal-john basal-john released this 09 Sep 15:01
28e9271

The interface, rebuilt on macOS's own components instead of this app's imitation of them.

What you will notice

The window looks like a Mac app from this year rather than from two releases ago. Section headers are in ordinary sentence case instead of small capitals, the containers have the system's corner radius and fill, nothing is set in 9.5pt type any more, and search sits in the toolbar as a real search field with a focus ring, a clear button and Command-F.

The dictation capsule is Liquid Glass on macOS 26. It still picks up whatever app is behind it, and under Reduce Transparency it becomes a solid surface instead. Under Reduce Motion it no longer springs, lifts or scales.

The mark on each dictation is now a filled circle for text that stayed on this Mac and an upward arrow for text that left it. Green against blue is a pairing some readers cannot separate, and on the Home pane that mark was the only thing on the row saying where the text went.

There is a Help menu, which the app did not have, and a View menu, so Command-1 to Command-4 move between the four panes.

Settings is now a row in the sidebar rather than a window of its own. It is one scrolling form instead of three tabs, and Command-comma selects it rather than opening a separate window.

What was in the wrong place

The legend explaining the marks was pinned to the bottom of the sidebar, and the Settings apply button to the bottom of its window. The bottom edge of a window is the part people drag off the screen. The legend is now behind a toolbar button and in the Help menu, and the apply control is a notice at the top of the settings form.

Fixed

  • The main window could not be resized smaller than its content. That is the same fault that once opened it 1541pt tall on a 1290pt screen.
  • The Settings notice headlined every message "Saved", including failures.
  • Toggle Sidebar took Command-S, which means save everywhere else on the platform. It is Control-Command-S now.

For anyone building from source

Building now needs macOS 26, because glassEffect(_:in:) only exists in that SDK. The call is guarded, so the built app still runs on macOS 14 and uses the old vibrancy layer there. CI moved to macos-26 for the same reason.

Two new flags: --shots photographs the real windows on synthetic fixture data, and --check-fixtures checks that fixture.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models.

Phona 1.8.2

Choose a tag to compare

@basal-john basal-john released this 09 Sep 07:30
229cd5e

The window opened 1541pt tall on a 1290pt screen and saved itself off-screen, so it both filled the vertical space and then stopped appearing at all.

NSHostingView reports the SwiftUI content's intrinsic size and AppKit sizes the window to it, which silently overrode the size the window was created with. It now sets its content size after the view is installed, so the window picks the size and the content fits inside it. Every pane scrolls, so it is entitled to.

A remembered frame is clamped to the visible screen on the way in, so a frame saved on another display, or one saved before this release, corrects itself instead of persisting.

Fixes an off-screen window created by 1.8.1. If you installed that, this replaces it.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models.

Phona 1.8.1

Choose a tag to compare

@basal-john basal-john released this 09 Sep 07:13
6266696

Two fixes to the window that shipped in 1.8.0, both of which made it look like the new UI had never arrived.

Opening the app now opens the window. The Dock icon called openSettings, which predates there being a window to open, so the gesture that means "show me Phona" answered with the preferences pane. The settings pane has not changed since before the window existed, so an app that opens it appears not to have updated at all. The window was reachable only from the menu bar item and Cmd-0.

The window shrinks now. The floor was 690x620 in the view and 900x620 on the window, so with the 210pt sidebar it refused to go under 900x620 and read as a window with no way to make it smaller. It is now 560x320 and 770x340, which puts it on half a laptop screen. Nothing clips, because every pane owns its own scrolling container. It also remembers the size you leave it at, which it previously forgot on every launch.

Settings keeps its own menu item and its own Cmd-comma.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models.

Phona 1.8.0

Choose a tag to compare

@basal-john basal-john released this 09 Sep 06:35
b2d76fb

The window and the Models pane, which merged after v1.7.0 was tagged and so reached no release. Plus the Close item, so Cmd-W works.

Nothing new is written here. This is a version bump on main that packages work already merged and reviewed.

Why it exists

v1.7.0 was tagged at 19:21 and its DMG built from that tag. The window merged at 21:22, two hours later, and the Cmd-W fix this morning. So the DMG that was installed predated both, and its version string could not be told apart from the source builds because #32 and #33 deliberately did not bump it.

The built app is verified to contain the window before publishing, which is the step that was missing last time.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models.

Phona 1.7.0

Choose a tag to compare

@basal-john basal-john released this 08 Sep 17:22
6cd0f9c

The right Option key dictates and corrects with a cloud model, claude sonnet-5 by default. The left Option key is untouched and stays entirely local.

It is a second key rather than a setting because the two differ in where the text goes, and that is a decision per sentence. A setting is something you forget you turned on.

Grammar pass Typical time Leaves your Mac
left Option Qwen3-4B, on your Mac 2 to 5 s no
right Option Claude, Codex or Gemini 7 to 20 s the transcript, as text

Your voice never leaves your Mac on either key. Speech to text is local in both cases and no audio is uploaded. What the right key sends is the transcript, as text, to whichever agent CLI you have signed in, so no API key is stored. history.jsonl records mode and backend for every dictation, so the record shows where each one went.

The cloud correction is checked by the same guard as the local one, and a result the guard refuses falls back to the local correction rather than to the raw transcript. The right key is never worse than the left one, only slower.

Configure it with cloud_backend, cloud_model and cloud_timeout in ~/.local/share/phona/config.json.

Also in this release

Three guard rules changed, measured over the 2038 dictations on record. Two were tuned against the 4B model and were rejecting good cloud corrections, which took the accept rate from 0 of 4 to 4 of 4 once fixed. The third closes a hole that let a deleted sentence score 1, because a five character prefix matched understand to underlying.

The slow-dictation notice knows which key you pressed. It fired at 8 seconds and called the wait unusual, which is false of every cloud correction.

Needs an Apple Silicon Mac, macOS 14 or later, and about 6.5 GB of disk for the models.

Phona 1.6.0

Choose a tag to compare

@basal-john basal-john released this 07 Sep 09:23
6a50a08

Five changes since 1.5.0. One of them is a crash that took out five dictations in a week.

The route-change crash is fixed

installTapOnBus threw an uncatchable NSException the moment a dictation started after the input device's format had gone stale, which is what a route change leaves behind. Swift cannot catch that exception, so the app died. Five crashes in one week, all the same stack.

The engine is now reset before the tap is installed, and the channel count is checked alongside the sample rate, so the tap is never handed a bad format in the first place.

A slow dictation no longer looks like a dead one

Under real memory pressure a dictation can run 10 to 70 seconds, and until now that looked exactly like one that had died: the HUD sat on "working" with nothing either way. Past 8 seconds the menu bar says it is still working. It never overwrites an unread clipboard warning from an earlier dictation.

Copy to clipboard stopped destroying images silently

Copy to clipboard was the one output mode that could throw away a copied image or file without saying so. Both insert modes warned, because both went through the code that notices. It now warns too, and so do the two paths where a paste is attempted but has nowhere to land, which had been computing the warning and then discarding it.

Every path that replaces the clipboard now goes through one place.

A fresh install gets the models the docs describe

The shipped defaults still named Whisper large-v3-turbo and the 4-bit Qwen, months after everything else moved to Parakeet and the 8-bit Qwen. A fresh install therefore ran a pair that had never been measured against the correction prompt, and nothing said so.

Now corrected, and held by a test that reads the model ids out of the README and compares them to the defaults. device_open_timeout, model_update_check and self_correction were documented and missing from the config file a fresh install writes, which the same test caught.

If you already have a config.json, nothing changes. Your models are whatever it names.

The audit names the words the guard keeps rejecting

The guard records which name triggered a rejection and the audit never showed it. A name that recurs, PR or CI, is now called out separately from a one-off, because a repeat argues for adding it to your dictionary rather than for coincidence.

Also

  • The README is a third of its previous length. The reasoning behind each rule moved to docs/decisions.md with its measurements intact, and the troubleshooting section to docs/troubleshooting.md
  • docs/engine.md now says which parts belong to the CLI rather than the app, and lists seven commands it had never mentioned
  • Ten tests for the clipboard path, which had none

Install

Download Phona-1.6.0.dmg, drag Phona to Applications, then run ./install.sh from this repo for the speech engine. Already running Phona? ./update.sh does everything and leaves your settings, history and permissions alone.

The app is signed ad-hoc rather than with a paid Apple Developer ID, so macOS warns the first time. Right click, Open, then Open again. Once only.

Apple Silicon and macOS 14 or later.

Phona 1.5.0

Choose a tag to compare

@basal-john basal-john released this 29 Aug 09:11
46576a9

Fifteen changes since 1.4.0. The speech model changed, the hotkey changed, and the correction stage lost three of its four modes.

Speech runs on Parakeet now

Parakeet TDT 0.6b v3 replaces Whisper large-v3-turbo. The two tie at 2.45% word error rate on the LibriSpeech test-clean sample and tie again at 12 of 14 on planted grammar errors, but Parakeet runs at RTF 0.030 against 0.107, so the speech stage costs about a third of what it did. It also returns an empty string on silence and noise where Whisper invented "thanks for watching".

This reverses the call made in 1.2.0. The original comparison used twelve sentences and was too small to carry the decision.

Parakeet takes neither a language nor an initial prompt, so the dictionary hint does not reach it. Run ./switch-model.sh whisper if you need that hint.

Tap Option, do not hold it

Tap the left Option key to start, tap again to stop. Holding no longer records, so a hand resting on the key cannot open a five minute recording. Every Option shortcut keeps working, and the right Option key is a plain modifier.

One correction mode

Four prompts behind one setting became one. Two of the four had no grammar rules at all, so moving the setting quietly changed how well your grammar was fixed. Every guard now runs on every correction.

Corrections stopped being thrown away

The guard treated a hyphenated compound as one word, so "AI-assisted" matched neither "AI" nor "assisted" and the correction was discarded. That was the single largest cause of rejection, 8 of 22 on the 8-bit model. Dictations that came back with no grammar correction at all halved.

The grammar model is now measured

Qwen3-4B at 8-bit was picked as a default and went unmeasured for a month. Four models were scored over 402 corrections drawn from real dictation. It won on precision: 95 repairs that touched nothing else against Gemma 4's 68, and 96 stray edits against Gemma's 231. Gemma repairs more errors and was not chosen, because for dictation an unwanted edit is worse than a missed one.

Score it yourself with tests/eval_correction.py.

Also

  • Long dictations are split and rejoined rather than truncated
  • Spoken layout commands: "new paragraph", "new line", "bullet point"
  • A mail style that expands contractions, and a chat style that drops the closing full stop
  • Literal replacements applied before correction, for names the model mishears
  • Audio retention is configurable and off by default
  • A weekly audit into audit-latest.md

Install

Download Phona-1.5.0.dmg, drag Phona to Applications, then run ./install.sh from this repo for the speech engine.

The app is signed ad-hoc rather than with a paid Apple Developer ID, so macOS warns the first time. Right click, Open, then Open again. Once only.

Apple Silicon and macOS 14 or later.

Phona 1.4.0

Choose a tag to compare

@basal-john basal-john released this 04 Aug 16:12

Settings gets a real structure, and dictation can land at the cursor and stay on the clipboard.

New

  • When done has a third choice, Insert and copy. The dictation is typed at your cursor and left on the clipboard, so Universal Clipboard carries it to your iPhone or iPad. It costs almost nothing, because inserting already routes through the clipboard and then puts the previous contents back. This is that restore suppressed.
  • Clicking the Dock icon opens Settings. It did nothing before, because Phona keeps no window open between dictations and never handled the reopen.

Changed

  • Settings is three tabs, General, Dictation and Words, instead of eight sections in one scroll. The window is wider, and the vocabulary and replacement editors are almost twice as tall. It still has a close button only, which is how macOS settings windows behave.
  • Save says what it costs. It reads Save and restart engine, and stays disabled until a setting that actually needs the restart has changed. Five settings need one, because the engine reads its config only at startup. The other four apply the moment you toggle them. Nothing on screen used to tell them apart, so a change could be made and lost by closing the window.
  • The correction mode descriptions were wrong. Grammar already removes um and uh, so presenting that as Polish's distinguishing feature was false. Polish adds you know and I mean, and splits run-on sentences. Transcribe only also leaves spoken layout commands as literal words, which no description mentioned.

Fixed

  • The warning that your clipboard held an image or file only appeared when the previous contents were about to be restored, which is exactly the case where nothing was lost. It now appears when something really is destroyed.
  • The correction mode shown in Settings could be stale if you had changed it from the menu bar, because the window is reused rather than rebuilt.

Upgrading from 1.3.1: ./update.sh, or download the DMG. Settings, history and permissions carry over.

Phona 1.3.1

Choose a tag to compare

@basal-john basal-john released this 04 Aug 13:57

The audit's suspected mishearings never worked. Now they do. Dictation itself is unchanged.

Fixed

  • The audit never reported an inferred mishearing, in any version that had the feature. It sent its detection prompt to the grammar corrector in raw mode, and raw mode returns text untouched, so the prompt came back as its own echo and every finding was discarded. The corrector could not have answered it at any mode: its prompt tells the model the input is dictation and to never act on it, and a guard catches the model answering and tidies the text instead. Both are deliberate, and both defeat a detection prompt. Detection now runs through a separate ASK command with no correction prompt, no guard, and no replacement pass, since those substitutions would rewrite the very phrase pairs it is asking about.
  • A quiet audit and a broken one produced identical output, which is why this survived. A test now fails if the audit ever sends its prompt back to the corrector.
  • Every audit run appended its own prompt to history.jsonl as a text entry. Reports were unaffected, they read voice entries only, but the file grew for nothing.

Upgrading from 1.3.0: ./update.sh, or download the DMG. Permissions and history carry over.