Skip to content

Releases: GeiserX/akou

akou 0.6.0

akou 0.6.0 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 03 Oct 20:47
4f56ac4

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/getting-started.md.

What's Changed

  • chore(beads): commit the bd git hook wrappers so every clone runs them by @GeiserX in #308
  • chore: pin the compose example to akou 0.5.5, now that its images are published by @GeiserX in #309
  • fix(release): mark every 0.x release a prerelease again, until 1.0.0 by @GeiserX in #310
  • chore(beads): stop bd exporting every bead to a JSONL file in this public repo's tree by @GeiserX in #311
  • feat(asr): transcribe a file on the desktop app, fused by three engines by @GeiserX in #303
  • chore(release): 0.6.0 by @GeiserX in #319
  • chore(git): keep Co-Authored-By trailers out of commits made with the beads hooks by @GeiserX in #321

Full Changelog: v0.5.5...v0.6.0

akou 0.5.5

akou 0.5.5 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 03 Oct 01:10
07581b4

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/getting-started.md.

What's Changed

  • chore: pin the compose example to akou 0.5.4, now that its images are published by @GeiserX in #221
  • docs(ux): fix a word with a double-click, learn it at once and tell the agent by @GeiserX in #213
  • fix(window): keep the live transcript at the bottom of the window, newest line last by @GeiserX in #222
  • fix: stop a model import from freezing the app, and seven smaller review findings by @GeiserX in #225
  • feat(vocab): tell a following agent when a fix teaches akou a word by @GeiserX in #223
  • fix(ui): say why the Dictation page cannot read or save when akou is out of reach, instead of hanging by @GeiserX in #229
  • feat(ui): fix a word by double-clicking it, and rename or forget what a fix learned by @GeiserX in #224
  • fix(cli): restart an app that stopped answering, so a recording starts in seconds by @GeiserX in #226
  • docs: make every diagram readable in the dark theme by @GeiserX in #230
  • fix(app): end an app whose thread is stuck, and keep a log that says so by @GeiserX in #227
  • docs: the roadmap and design say what ships today and what waits by @GeiserX in #289
  • fix(cli): say "akou has quit" only once every akou process is gone by @GeiserX in #228
  • fix(ui): say Back to the end on a saved call, where nothing is live by @GeiserX in #235
  • chore(beads): keep bead text out of the public repo, tracker moves to private Gitea by @GeiserX in #304
  • chore(beads): point the tracker at private Gitea so it never pushes to this public repo by @GeiserX in #305
  • test: stop two dictation tests reading what is still being written by @GeiserX in #232
  • fix(models): let two pulls of one model write its file once by @GeiserX in #234
  • fix(helpers): stop a write to a helper that just exited from failing the run by @GeiserX in #236
  • feat(jobs): tell a client when each word was said, how sure the engine was, and what was lost by @GeiserX in #238
  • fix(cli): exit 3 whenever there is no call to act on, and 130 on Ctrl-C by @GeiserX in #239
  • fix(asr): keep dictation ids unique in one millisecond and drop late review cancels by @GeiserX in #240
  • fix(server): let a client trust the job feed: no ghost events, cancels it can match, a reset it can see by @GeiserX in #242
  • feat(server): run Qwen for jobs that name no model wherever it is downloaded by @GeiserX in #243
  • fix(desktop): stop the app dying on a Mac with no display when a call starts by @GeiserX in #244
  • fix(server): fail a diarize job that has no speaker helper instead of dropping the labels by @GeiserX in #245
  • feat(ui): pick a job's model, language and dictation engine by name on both pages by @GeiserX in #246
  • feat(server): let each job bound its auto language, so one server can serve clients that speak different languages by @GeiserX in #247
  • fix(vocab): stop a word correcting its own inflections, and import archive glossaries whole by @GeiserX in #248
  • feat(api): list every error code in the OpenAPI file, so clients can rely on them by @GeiserX in #250
  • feat(mcp): let an agent transcribe a file on its machine in one tool call by @GeiserX in #251
  • fix(dictation): stop a tapped key session after silence, keep Enter a send in a Shift chord, and take Mouse4 as the key by @GeiserX in #252
  • feat(dictation): name engines and apps in words on the Dictation page by @GeiserX in #258
  • docs: fail the check when the settings reference drifts from the registry by @GeiserX in #259
  • test(ui): save every parity screen in both themes so a window change can be reviewed from CI by @GeiserX in #261
  • docs(server): put a reverse proxy and a long-recording Mac setup on the server page by @GeiserX in #271
  • feat(dictation): make the draft key, fix last and paste last keys actually work by @GeiserX in #277
  • ci: turn the nightly green with real baselines, and stop long dictations losing words in Qwen by @GeiserX in #278
  • feat(window): draw the player bar as b2, with a mark on the scrubber for each note taken in the part by @GeiserX in #282
  • docs(design): stop calling unbuilt CLI commands, window features and the shared tray icon built by @GeiserX in #288
  • docs: say what speech engines akou runs today: Nemotron live, Qwen final, Parakeet as fallback by @GeiserX in #290
  • ci: let the repeat dispatch that proves the flake fixes go green on the helper legs by @GeiserX in #300
  • chore(release): 0.5.5, published as a full release now that only a version like 0.6.0-rc.1 is a prerelease by @GeiserX in #306

Full Changelog: v0.5.4...v0.5.5

akou 0.5.4

akou 0.5.4 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 01 Oct 05:51
47f0c85

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/getting-started.md.

What's Changed

  • docs(skill): stop reminding about consent from the skill too, as the window no longer does by @GeiserX in #211
  • feat(live): give a read up to 30 s for the second pass, so Qwen finishes a 3 min backlog before the agent reads by @GeiserX in #212
  • chore: pin the compose example to akou 0.5.3, now that its images are published by @GeiserX in #214
  • fix(models): stop a model that does not fit the call from looking impossible to download by @GeiserX in #215
  • feat(final): show how far the final pass is, in the window and akou status by @GeiserX in #217
  • feat(models): show the whole catalog on the Models page, and add six streaming Nemotron chunk sizes by @GeiserX in #216
  • feat(final): run Qwen for the final transcript whenever it is downloaded, with no Parakeet needed by @GeiserX in #218
  • feat(models): stop requiring Parakeet on a Mac that runs Nemotron live and Qwen after the call by @GeiserX in #219
  • chore(release): 0.5.4 by @GeiserX in #220

Full Changelog: v0.5.3...v0.5.4

akou 0.5.3

akou 0.5.3 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 30 Sep 21:02
43aff82

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/getting-started.md.

What's Changed

  • chore: pin the compose example to akou 0.5.2, now that its images are published by @GeiserX in #197
  • docs: publish the docs as a site at https://geiserx.github.io/akou/ by @GeiserX in #190
  • docs(ux): give the live menu two slots, each taking a model you can add in place by @GeiserX in #198
  • docs: show the real app on the site and the README by @GeiserX in #201
  • feat(live): pick the live model by name, with the second pass as a separate choice and interval by @GeiserX in #199
  • fix(indicator): make the recording popup's level bars move and remove the extra space past Stop by @GeiserX in #202
  • fix(window): stop showing a consent reminder row every time a call starts by @GeiserX in #204
  • feat(live): offer Parakeet as a second pass, cutting live errors 16 % in English and 46 % in Spanish by @GeiserX in #200
  • feat(window): pick the live model and the second pass in one panel, and add a model from it by @GeiserX in #203
  • feat(live): run the second pass before an agent reads the call, so it reads corrected lines by @GeiserX in #206
  • feat(dictation): show everything you say in the pill, scrolling past eight lines by @GeiserX in #205
  • fix(window): name the 1120 ms Nemotron by its wait, not by a steadiness nobody measured by @GeiserX in #209
  • feat(dictation): type with the live model by default, with no wait for a model to load by @GeiserX in #207
  • fix(dictate): make the dictation key work once Accessibility is granted, without a restart by @GeiserX in #208
  • chore(release): 0.5.3 by @GeiserX in #210

Full Changelog: v0.5.2...v0.5.3

akou 0.5.2

akou 0.5.2 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 30 Sep 09:27
1946c0f

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/getting-started.md.

What's Changed

  • feat(window): make Settings a page of plain rows you leave from the sidebar, not a dialog of config keys by @GeiserX in #178
  • chore: pin the compose example to akou 0.5.1, now that its images are published by @GeiserX in #181
  • feat(window): make Models a page of plain facts you leave like Settings, and let a download be cancelled by @GeiserX in #182
  • feat(window): pick the live transcription model in the Record row, where the template was by @GeiserX in #184
  • docs: a README that says what akou is in three sentences and how to install it by @GeiserX in #188
  • feat(window): hide Enhance and Find misheard words, so the notes pane is just your notes by @GeiserX in #185
  • feat(window): make Dictation a page of plain rows you leave from the sidebar, not a dialog of config keys by @GeiserX in #189
  • feat(window): let the workspace be switched and added with a click, and stop asking for a notes template by @GeiserX in #183
  • feat(live): review live lines with Qwen once a minute, 80 % fewer requests at the same WER, and on Automatic when the Mac allows by @GeiserX in #191
  • feat(vocab): fix a misheard word once on its line and the whole call reads it right by @GeiserX in #186
  • feat(start): let an agent follow the call already recording, and search the call when no assistant is set up by @GeiserX in #192
  • feat(assistant): keep your API key in the macOS Keychain, and pick Claude Code, a key, a local model or none in Settings by @GeiserX in #194
  • feat(window): make Words and History pages under Dictation, so no settings dialog is left by @GeiserX in #193
  • feat(window): set up only what akou is for on first run: calls, dictation or both by @GeiserX in #195
  • feat(tray): show the akou mark with a red dot while a call records, with no text by @GeiserX in #187
  • chore(release): 0.5.2, so every setting is a page of plain words and first run asks only what you need by @GeiserX in #196

Full Changelog: v0.5.1...v0.5.2

akou 0.5.1

akou 0.5.1 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 29 Sep 19:43
5fcd2f3

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

What's Changed

  • feat(live): rewrite each live line once, with Qwen alone, for 4.75 % WER on FLEURS English by @GeiserX in #171
  • chore: add the security policy and maintenance files every GeiserX repo carries by @GeiserX in #172
  • fix(ui): let every dialog close from its top corner, Escape or a click outside by @GeiserX in #175
  • docs(ux): the Settings, Dictation and Models pages as rows in the main window (direction A) by @GeiserX in #174
  • ci(stale): let the stale job save its progress and stop promising reopen-by-comment by @GeiserX in #176
  • feat(window): let the macOS window run to the top edge, with no grey title bar by @GeiserX in #177
  • fix(final): stop each final pass from keeping 2.7 GB of models, which filled the disk with swap by @GeiserX in #179
  • chore(release): 0.5.1, so a final pass no longer fills the disk with swap by @GeiserX in #180

Full Changelog: v0.5.0...v0.5.1

akou 0.5.0

akou 0.5.0 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 29 Sep 13:30
3445a75

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

What's Changed

  • feat(ui): find a call by workspace, day or title, and see whether akou can record by @GeiserX in #138
  • docs(ux): the dictation pill becomes an island at the top of the display by @GeiserX in #139
  • feat(ui): record from one composer row with a round red Record, not a header of debug chips by @GeiserX in #140
  • docs(dictation): show the words as you speak by default, and make streaming partials a P0 by @GeiserX in #141
  • chore(beads): what a three-day batch run against a 190-hour archive found missing by @GeiserX in #143
  • feat(dictation): make the pill the black island at the top, and the draft box a sheet dropped from it by @GeiserX in #142
  • feat(ui): keep Ask and the note input on screen together, instead of tabs that hide one behind the other by @GeiserX in #144
  • feat(dictation): show the words as you speak on the island, and switch the language from it by @GeiserX in #145
  • feat(ui): keep the accent for the one primary action, and show the player only when a call has a recording by @GeiserX in #146
  • fix(brand): run the k's arm into its stem, so no rounded lump shows under it by @GeiserX in #147
  • feat(dictation): hear a dictation when the island is off, and let the Dictation page reach the running helper by @GeiserX in #148
  • docs(ux): make the window docs describe the sidebar shell that shipped, and stop saying ready while nothing can record by @GeiserX in #149
  • feat(final): run the final pass on the calls the app records, by decoding their Opus parts in the helper by @GeiserX in #150
  • feat(dictation): make the Dictation page's mic meter move, and say when reading the field waits for Accessibility by @GeiserX in #151
  • feat(calls): let a call be renamed at any time, live or saved, from every door by @GeiserX in #152
  • feat(jobs): let a transcription started over the API carry a name, and find it by that name by @GeiserX in #154
  • feat(dictation): say on the island why the dictation key does nothing, and start again once Accessibility is back by @GeiserX in #153
  • feat(live): write the live transcript with streaming Nemotron, so words show sooner and are never taken back by @GeiserX in #156
  • feat(live): let the user choose what writes the live transcript, per machine and per call by @GeiserX in #158
  • feat(dictation): land dictated words with the spaces and case the sentence needs, and follow each app's own rule by @GeiserX in #157
  • feat(asr): fuse several engines' words into one line by confidence vote, matching the benchmark on every unit by @GeiserX in #159
  • feat(dictation): hear your languages without asking, fix a wrong one with a click, and keep a failed dictation's words by @GeiserX in #160
  • feat(live): rewrite each utterance during the call with Parakeet, then Qwen, cutting live WER by about a quarter by @GeiserX in #161
  • feat(dictation): offer the other engine's word under an unsure one, and show which language a dictation went in by @GeiserX in #162
  • feat(live): stop the GPU running all call by default; auto uses streaming Nemotron and the upgrade is opt-in by @GeiserX in #165
  • fix(window): show a call started from the CLI or an agent even after another call was clicked by @GeiserX in #164
  • feat(dictation): show the island the moment the key goes down, on the display you dictate into by @GeiserX in #163
  • feat(dictation): copy the words when Accessibility is refused, and keep an app's insert rule in every draft box by @GeiserX in #167
  • fix(asr): stop Qwen inventing words on silence, and check its accuracy and memory every night by @GeiserX in #166
  • chore(beads): put the lanes' notes and closures in the tracked export by @GeiserX in #168
  • chore(release): 0.5.0, a new window, live dictation words, and the final pass on every call by @GeiserX in #169

Full Changelog: v0.4.0...v0.5.0

akou 0.4.0

akou 0.4.0 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 28 Sep 06:29

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

What's Changed

  • chore: pin the compose example to akou 0.3.0 and record the 0.3.0 image tags on the GPU bead by @GeiserX in #72
  • docs(ux): spec dictation so akou types what you say into any app and learns from your fixes by @GeiserX in #70
  • chore(beads): track the dictation spec as an epic with one bead per item, so the lanes can start by @GeiserX in #73
  • fix(server): stop a reused Idempotency-Key from returning a job made with other options by @GeiserX in #74
  • chore(beads): record that #74 fixed the idempotency, s? label and install.md server bugs by @GeiserX in #77
  • fix(asr): one llama.cpp build table, so a Mac on cpu runs best and /v1/server stops claiming a GPU on the CPU by @GeiserX in #76
  • fix(server): speaker labels no longer add "Yeah." at turn edges by @GeiserX in #75
  • chore(beads): record that #76 fixed the llama.cpp build table review, and track its four small follow-ups by @GeiserX in #78
  • chore(beads): close the diarization filler and similar-voices beads fixed and measured in #75 by @GeiserX in #79
  • fix(server): a retry with keywords reordered or the language in another case keeps its job by @GeiserX in #82
  • feat(dictation): answer a dictation at once on a busy server, and send it there without the audio crossing the internet in the clear by @GeiserX in #80
  • feat(capture): akou-capture dictate runs a whole dictation on fakes, so no test ever presses a real key by @GeiserX in #83
  • feat(models): see how each model behaves, download or delete it from the UI, and drop unused ones in the app too by @GeiserX in #81
  • chore(beads): close the Models page bead shipped in #81 and track its three review follow-ups by @GeiserX in #86
  • feat(ui): give dictation a pill, a draft box and a settings page before the session lands by @GeiserX in #84
  • feat(dictation): a remote that is down no longer loses the dictation, and dictated text can go through your provider first by @GeiserX in #87
  • feat(dictation): decode a dictation on the loaded model and read the helper's own lines, so akou dictate clip.wav prints a word by @GeiserX in #85
  • feat(capture): dictation pastes only where it began, never into a password field, and gives the clipboard back after the app reads it by @GeiserX in #88
  • feat(dictation): send the audio to the remote while the key is held, so a 20 s dictation answers as fast as a 3 s one by @GeiserX in #90
  • fix(capture): make the helper's tests compile on main again by @GeiserX in #91
  • feat(ui): find, retry and fix past dictations, and bind a dictation key by pressing it by @GeiserX in #89
  • test(ui): stop the narrow Models page test from failing every open PR by @GeiserX in #95
  • feat(capture): the key tap never waits on an insert, and a fix typed in the app comes back to learn from by @GeiserX in #93
  • fix(dictation): a remote dictation that fails at the press no longer holds the rest of the hold in memory by @GeiserX in #92
  • feat(dictation): set dictation up from its settings, and a new dictation key works at once or is refused where you set it by @GeiserX in #94
  • feat(ui): a per-app rules editor, and the grant setup the dictation switch runs once akou reports its grants by @GeiserX in #96
  • feat(capture): a real Mac hears the dictation key, and the grant it needs is measured by @GeiserX in #97
  • feat(ui): choose the dictation microphone from a list, and hear dictation when the pill is off by @GeiserX in #99
  • feat(dictation): offer to learn only real mishearings, and keep dictation fixes out of call transcripts by @GeiserX in #98
  • feat(ui): teach dictation your own words and replacements without touching calls, and show live words in the pill once allowed by @GeiserX in #101
  • feat(capture): dictation on a Mac now pastes and types into the app instead of always failing by @GeiserX in #100
  • feat(dictation): start a dictation from the tray or a script, and delete dictations for good by @GeiserX in #102
  • feat(capture): dictate on the built-in mic instead of a Bluetooth headset, and keep Fn from opening the emoji picker by @GeiserX in #103
  • feat(dictation): dictate with Qwen where a GPU runs it, and fall back to Parakeet instead of hanging by @GeiserX in #105
  • feat(dictation): stop room noise and filler words from being typed, and stream dictation events by @GeiserX in #106
  • fix(ui): stop the dictionary wiping a word's note, decode and pending review when you add a form by @GeiserX in #104
  • test(dictation): catch app-helper drift in CI on every OS, and let spoken marks become punctuation by @GeiserX in #107
  • feat(dictation): keep each dictation's audio so Retry can decode it again on another engine by @GeiserX in #108
  • feat(dictation): show the dictation pill without taking the keyboard from the app you dictate into by @GeiserX in #109
  • feat(dictation): keep a dictation the app could not take in the draft box, and learn the word you fix there by @GeiserX in #110
  • fix(dictation): keep dictated words out of the app log when Learn or Undo fails by @GeiserX in #111
  • chore(beads): record the dictation wave's closes, follow-ups and flaky tests from #104 to #110 by @GeiserX in #112
  • chore(beads): keep the owner's deployment detail out of the public tracker by @GeiserX in #115
  • feat(ui): save dictation words and replacements for real, so "example dot com" types example.com by @GeiserX in #117
  • feat(capture): a Linux paste waits for the target's read, and the first syllable is proven to reach the text by @GeiserX in #114
  • feat(dictation): dictation.format now changes what gets typed, and a remote dictation uploads during the hold instead of at release by @GeiserX in #113
  • test(ui): prove the draft box and its learn chip work end to end, so page and main side cannot drift apart by @GeiserX in #118
  • feat(dictation): let Enter send, Shift+Enter draft and Escape cancel while a dictation runs by @GeiserX in #116
  • feat(ui): name a per-app rule's app by dictating into it, instead of typing its bundle id by @GeiserX in #120
  • feat(ui): let the Dictation page test the remote akou and say when it is down, so a bad key or dead remote shows before the first press by @GeiserX in #121
  • feat(dictation): a forgotten dictation stops by itself, "send it" sends, and the words you fixed can be reviewed by @GeiserX in #122
  • fix(dictation): keep dictated words out of the log when formatting fails, and prove the dictation lane on the shipped server by @GeiserX in #124
  • feat(ui): list the words fixed while dictating under Words to review, so a let-go fix can still be learned or taken back by @GeiserX in #123
  • feat(capture): dictation on Windows hears its key and inserts into the app instead of exiting by @GeiserX in #119
  • chore(beads): track the gaps a real archive re-transcription found in best jobs by @GeiserX in #126
  • feat(dictation): check a fix against the dictation's own audio before offering to learn it by @GeiserX in #125
  • feat(dictation): learn a word you fix in the app itself, not only in the draft box by @GeiserX in #128
  • feat(capture): dictation on Linux hears its key and records the mic instead of exiting by @GeiserX in #127
  • fix(dictation): stop pressing into a GPU a final pass holds, and warm best again once it is free by @GeiserX in #129
  • feat(capture): a dictation pauses the music that is playing and gives back only what it paused by @GeiserX in #130
  • feat(capture): dict...
Read more

akou 0.3.0

akou 0.3.0 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 26 Sep 20:15

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

What's Changed

  • feat(server): let a 52,000-file backlog drain: jobs in parallel, priority first, and 429 with Retry-After when the queue is full by @GeiserX in #62
  • feat(asr): run the best preset on Qwen3-ASR, on the Mac's GPU natively or any box's CPU by @GeiserX in #67
  • chore(beads): close SV-P1 and SV-T6 now that drumsergio/akou:0.2.1 is published, and pin the compose example to it by @GeiserX in #68
  • feat(server): find the GPU the box has, Intel, AMD, NVIDIA or Apple, and ship images that can use it by @GeiserX in #65
  • feat(server): let one akou send its jobs to another, so a Mac mini's GPU can serve the archive by @GeiserX in #63
  • chore(beads): close the backlog, best, offload and Mac server beads that shipped in #62, #63 and #67, and note what the GPU work still needs by @GeiserX in #69
  • chore(release): 0.3.0, the best preset on the GPU the box has by @GeiserX in #71

Full Changelog: v0.2.1...v0.3.0

akou 0.2.1

akou 0.2.1 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 26 Sep 18:01

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

0.2.1 is the first published 0.2 release. The v0.2.0 tag built but never published an image or a release, so everything listed under 0.2.0 in the changelog ships here. The server image is drumsergio/akou:0.2.1, for linux/amd64 and linux/arm64. There is no latest tag.

Known limitations

  • Unsigned macOS build. The first open needs a manual step, and macOS may ask for the microphone and system audio again after an update. See docs/install.md.
  • macOS only as an app. The release ships the macOS app (Apple Silicon), the CLI and the server image. There is still no packaged desktop app for Windows or Linux. The Linux and Windows CLI archives manage models, the skill and the settings, but cannot record. The new linux-arm64 archive has not been run on a Raspberry Pi yet.
  • Server mode and the image are new in this release. CI builds the image on amd64 and arm64 and transcribes a spoken sentence in each. Nobody has run it for long on a real server yet. The design and what is still missing are in docs/ux/SERVER.md. A container refuses to start until you set AKOU_BEHIND_PROXY=true and put a reverse proxy with TLS in front of it, because akou has no TLS of its own. See docs/install.md.
  • The server's web page is partly built. The Models page shows the models' state and a download button, but not each model's size, last use, deletion date or a Delete button. There is no preset picker yet, and the Jobs page polls twice a second instead of following the event feed.
  • One engine, on the CPU. Only the fast preset has an engine. lite, best and fusion are refused, and akou does not choose by hardware yet. The image uses no GPU. Results carry no word times or confidences: words is empty and both confidence fields are null.
  • Transcribing a file needs server mode. The single-file CLI's akou serve answers the API but carries no speech engine, and says so when it starts. On a Mac, use the image or a source checkout.
  • The window is tested in Chromium, but the app draws it in WebKit. CI runs the window tests in headless Chromium. The WebKit run is manual while some of its tests still fail there, so the new player, line menu and indicator are not tested in the app's own webview.
  • Drift between two clocks is not measured. One recording can take the mic and the call from two devices with separate clocks. How far they drift apart over an hour has not been measured on real hardware yet. See docs/gates/M0-results.md.
  • The 1 s rebuild has not been seen on a real device. It is proven in simulated capture, but no call audio died during the hour-long run on a real Mac.
  • Large first download. The speech and speaker models are about 3.0 GB, downloaded on first run.

Full notes: CHANGELOG.md. Everything since the last release: v0.1.0...v0.2.1

What's Changed

  • docs(brand): give akou a logo and a README banner that says what it is by @GeiserX in #59
  • chore(beads): record why v0.2.0 published no image, so Telegram-Archive does not pin it by @GeiserX in #64
  • feat(desktop): show the akou mark in the tray, the Dock and the browser tab by @GeiserX in #61
  • chore(release): 0.2.1, publish the server image as drumsergio/akou because 0.2.0's geiserx namespace does not exist by @GeiserX in #66

Full Changelog: v0.2.0...v0.2.1