v0.7.0 — Voice input with Whisper on the device
Voice input with Whisper, on the device (#2)
- Settings → Voice input is a choice per device: Keyboard dictation (as before) or Whisper on this device.
- With Whisper, Quick entry gets a Speak button. Say the day ("36.52, light bleeding, no cramps, headache, tired"), tap Stop, and the form fills in. What was heard is shown so you can correct it.
- Speech is recognised in the browser (OpenAI's Whisper, run with ONNX Runtime WebAssembly in a background worker). The audio and the text never leave the device.
- The model is downloaded once from your own Ebbwell server (never from the internet), only by signed-in users, and kept in the browser. Settings can download it ahead of time (on Wi-Fi) or remove it.
Configuration:
VOICE_MODEL=base(default, ~79 MB, more accurate, especially in Arabic),tiny(~43 MB, faster on old phones) oroff. Both models are in the image, which is about 140 MB larger.- The microphone needs HTTPS.
- New or changed security headers:
Cross-Origin-Embedder-Policy: require-corp(lets the recogniser use several threads),'wasm-unsafe-eval'in the CSP (WebAssembly only, not JavaScripteval),Permissions-Policy: microphone=(self). If your reverse proxy rewrites security headers, keep these.
Upgrade: change the image to ghcr.io/maxren2/ebbwell:0.7.0. No configuration change is needed. Voice input stays on the keyboard until someone picks Whisper in Settings.
See docs/DEPLOY.md § 13.