Skip to content

akou 0.2.1

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 26 Sep 18:01
· 247 commits to main since this release

This build is not signed by Apple. The first time you open akou, macOS refuses it. On macOS 14, Control-click akou in Applications, choose Open, then Open again. On macOS 15 and later, open akou once, then go to System Settings, Privacy & Security, and click Open Anyway. After an update macOS may ask for the microphone and system audio again. The speech models are downloaded on first run. See docs/install.md.

0.2.1 is the first published 0.2 release. The v0.2.0 tag built but never published an image or a release, so everything listed under 0.2.0 in the changelog ships here. The server image is drumsergio/akou:0.2.1, for linux/amd64 and linux/arm64. There is no latest tag.

Known limitations

  • Unsigned macOS build. The first open needs a manual step, and macOS may ask for the microphone and system audio again after an update. See docs/install.md.
  • macOS only as an app. The release ships the macOS app (Apple Silicon), the CLI and the server image. There is still no packaged desktop app for Windows or Linux. The Linux and Windows CLI archives manage models, the skill and the settings, but cannot record. The new linux-arm64 archive has not been run on a Raspberry Pi yet.
  • Server mode and the image are new in this release. CI builds the image on amd64 and arm64 and transcribes a spoken sentence in each. Nobody has run it for long on a real server yet. The design and what is still missing are in docs/ux/SERVER.md. A container refuses to start until you set AKOU_BEHIND_PROXY=true and put a reverse proxy with TLS in front of it, because akou has no TLS of its own. See docs/install.md.
  • The server's web page is partly built. The Models page shows the models' state and a download button, but not each model's size, last use, deletion date or a Delete button. There is no preset picker yet, and the Jobs page polls twice a second instead of following the event feed.
  • One engine, on the CPU. Only the fast preset has an engine. lite, best and fusion are refused, and akou does not choose by hardware yet. The image uses no GPU. Results carry no word times or confidences: words is empty and both confidence fields are null.
  • Transcribing a file needs server mode. The single-file CLI's akou serve answers the API but carries no speech engine, and says so when it starts. On a Mac, use the image or a source checkout.
  • The window is tested in Chromium, but the app draws it in WebKit. CI runs the window tests in headless Chromium. The WebKit run is manual while some of its tests still fail there, so the new player, line menu and indicator are not tested in the app's own webview.
  • Drift between two clocks is not measured. One recording can take the mic and the call from two devices with separate clocks. How far they drift apart over an hour has not been measured on real hardware yet. See docs/gates/M0-results.md.
  • The 1 s rebuild has not been seen on a real device. It is proven in simulated capture, but no call audio died during the hour-long run on a real Mac.
  • Large first download. The speech and speaker models are about 3.0 GB, downloaded on first run.

Full notes: CHANGELOG.md. Everything since the last release: v0.1.0...v0.2.1

What's Changed

  • docs(brand): give akou a logo and a README banner that says what it is by @GeiserX in #59
  • chore(beads): record why v0.2.0 published no image, so Telegram-Archive does not pin it by @GeiserX in #64
  • feat(desktop): show the akou mark in the tray, the Dock and the browser tab by @GeiserX in #61
  • chore(release): 0.2.1, publish the server image as drumsergio/akou because 0.2.0's geiserx namespace does not exist by @GeiserX in #66

Full Changelog: v0.2.0...v0.2.1