Skip to content

v0.7.0 — Voice input with Whisper on the device

Choose a tag to compare

@Maxren2 Maxren2 released this 29 Sep 15:06
· 4 commits to main since this release

Voice input with Whisper, on the device (#2)

  • Settings → Voice input is a choice per device: Keyboard dictation (as before) or Whisper on this device.
  • With Whisper, Quick entry gets a Speak button. Say the day ("36.52, light bleeding, no cramps, headache, tired"), tap Stop, and the form fills in. What was heard is shown so you can correct it.
  • Speech is recognised in the browser (OpenAI's Whisper, run with ONNX Runtime WebAssembly in a background worker). The audio and the text never leave the device.
  • The model is downloaded once from your own Ebbwell server (never from the internet), only by signed-in users, and kept in the browser. Settings can download it ahead of time (on Wi-Fi) or remove it.

Configuration:

  • VOICE_MODEL=base (default, ~79 MB, more accurate, especially in Arabic), tiny (~43 MB, faster on old phones) or off. Both models are in the image, which is about 140 MB larger.
  • The microphone needs HTTPS.
  • New or changed security headers: Cross-Origin-Embedder-Policy: require-corp (lets the recogniser use several threads), 'wasm-unsafe-eval' in the CSP (WebAssembly only, not JavaScript eval), Permissions-Policy: microphone=(self). If your reverse proxy rewrites security headers, keep these.

Upgrade: change the image to ghcr.io/maxren2/ebbwell:0.7.0. No configuration change is needed. Voice input stays on the keyboard until someone picks Whisper in Settings.

See docs/DEPLOY.md § 13.