Skip to content

Releases: VocaHQ/vocalinux

Nightly Build - 0.15.0.dev20260809+0685305

Choose a tag to compare

@github-actions github-actions released this 09 Aug 03:49
0685305

Nightly Build 0.15.0.dev20260809+0685305

This is a nightly development build. It may contain bugs or incomplete features. Use at your own risk.


What's in This Build

Commit Details:

  • SHA: 0685305
  • Message: chore(license): migrate from GPL-3.0 to AGPL-3.0 (#660)
  • Author: Jatin K Malik
  • Date: 2026-08-08 13:10:37 -0700

Changes Since Last Nightly

  • 0685305 chore(license): migrate from GPL-3.0 to AGPL-3.0 (#660)
  • c337d8f feat(shortcuts): default to hold Right Alt (push-to-talk) for new installs (#648)
  • 06e8809 feat(updates): notify when a newer GitHub release is available (#645)

Installation

Quick Install (Recommended)

Download the wheel file below and install:

pip install vocalinux-0.15.0.dev20260809+0685305-py3-none-any.whl

From PyPI (Stable)

For the latest stable release:

pip install vocalinux

Using Install Script

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh | bash

AppImage (no install, no root)

Download the .AppImage matching your CPU (x86_64 or aarch64) below:

chmod +x Vocalinux-0.15.0.dev20260809+0685305-x86_64.AppImage
./Vocalinux-0.15.0.dev20260809+0685305-x86_64.AppImage

Requires xdotool/wtype/ydotool on the host for text injection, same as the PyPI install path.


Build Information

Property Value
Version 0.15.0.dev20260809+0685305
Base Version 0.15.0
Build Date 20260809
Python 3.9+
OS Linux (X11 & Wayland)

Files

  • vocalinux-0.15.0.dev20260809+0685305-py3-none-any.whl - Python wheel (recommended)
  • vocalinux-0.15.0.dev20260809+0685305.tar.gz - Source distribution
  • Vocalinux-0.15.0.dev20260809+0685305-x86_64.AppImage - Self-contained, no install (Intel/AMD)
  • Vocalinux-0.15.0.dev20260809+0685305-aarch64.AppImage - Self-contained, no install (ARM64, added shortly after this release is published)

Links


Only the 3 most recent nightly builds are kept. This build will be automatically deleted when newer nightlies are created.

Generated on 2026-08-09 03:49:22 UTC

Nightly Build - 0.15.0.dev20260808+4a1d3ed

Choose a tag to compare

@github-actions github-actions released this 08 Aug 03:41
4a1d3ed

Nightly Build 0.15.0.dev20260808+4a1d3ed

This is a nightly development build. It may contain bugs or incomplete features. Use at your own risk.


What's in This Build

Commit Details:

  • SHA: 4a1d3ed
  • Message: fix(clipboard): text-only reads and safe overlapping restore (#646)
  • Author: Jatin K Malik
  • Date: 2026-08-06 00:01:44 -0700

Changes Since Last Nightly

No new commits since last nightly


Installation

Quick Install (Recommended)

Download the wheel file below and install:

pip install vocalinux-0.15.0.dev20260808+4a1d3ed-py3-none-any.whl

From PyPI (Stable)

For the latest stable release:

pip install vocalinux

Using Install Script

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh | bash

AppImage (no install, no root)

Download the .AppImage matching your CPU (x86_64 or aarch64) below:

chmod +x Vocalinux-0.15.0.dev20260808+4a1d3ed-x86_64.AppImage
./Vocalinux-0.15.0.dev20260808+4a1d3ed-x86_64.AppImage

Requires xdotool/wtype/ydotool on the host for text injection, same as the PyPI install path.


Build Information

Property Value
Version 0.15.0.dev20260808+4a1d3ed
Base Version 0.15.0
Build Date 20260808
Python 3.9+
OS Linux (X11 & Wayland)

Files

  • vocalinux-0.15.0.dev20260808+4a1d3ed-py3-none-any.whl - Python wheel (recommended)
  • vocalinux-0.15.0.dev20260808+4a1d3ed.tar.gz - Source distribution
  • Vocalinux-0.15.0.dev20260808+4a1d3ed-x86_64.AppImage - Self-contained, no install (Intel/AMD)
  • Vocalinux-0.15.0.dev20260808+4a1d3ed-aarch64.AppImage - Self-contained, no install (ARM64, added shortly after this release is published)

Links


Only the 3 most recent nightly builds are kept. This build will be automatically deleted when newer nightlies are created.

Generated on 2026-08-08 03:41:02 UTC

Nightly Build - 0.15.0.dev20260807+4a1d3ed

Choose a tag to compare

@github-actions github-actions released this 07 Aug 04:33
4a1d3ed

Nightly Build 0.15.0.dev20260807+4a1d3ed

This is a nightly development build. It may contain bugs or incomplete features. Use at your own risk.


What's in This Build

Commit Details:

  • SHA: 4a1d3ed
  • Message: fix(clipboard): text-only reads and safe overlapping restore (#646)
  • Author: Jatin K Malik
  • Date: 2026-08-06 00:01:44 -0700

Changes Since Last Nightly

  • 4a1d3ed fix(clipboard): text-only reads and safe overlapping restore (#646)

Installation

Quick Install (Recommended)

Download the wheel file below and install:

pip install vocalinux-0.15.0.dev20260807+4a1d3ed-py3-none-any.whl

From PyPI (Stable)

For the latest stable release:

pip install vocalinux

Using Install Script

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh | bash

AppImage (no install, no root)

Download the .AppImage matching your CPU (x86_64 or aarch64) below:

chmod +x Vocalinux-0.15.0.dev20260807+4a1d3ed-x86_64.AppImage
./Vocalinux-0.15.0.dev20260807+4a1d3ed-x86_64.AppImage

Requires xdotool/wtype/ydotool on the host for text injection, same as the PyPI install path.


Build Information

Property Value
Version 0.15.0.dev20260807+4a1d3ed
Base Version 0.15.0
Build Date 20260807
Python 3.9+
OS Linux (X11 & Wayland)

Files

  • vocalinux-0.15.0.dev20260807+4a1d3ed-py3-none-any.whl - Python wheel (recommended)
  • vocalinux-0.15.0.dev20260807+4a1d3ed.tar.gz - Source distribution
  • Vocalinux-0.15.0.dev20260807+4a1d3ed-x86_64.AppImage - Self-contained, no install (Intel/AMD)
  • Vocalinux-0.15.0.dev20260807+4a1d3ed-aarch64.AppImage - Self-contained, no install (ARM64, added shortly after this release is published)

Links


Only the 3 most recent nightly builds are kept. This build will be automatically deleted when newer nightlies are created.

Generated on 2026-08-07 04:33:06 UTC

Nightly Build - 0.15.0.dev20260806+a15baa8

Choose a tag to compare

@github-actions github-actions released this 06 Aug 05:24
a15baa8

Nightly Build 0.15.0.dev20260806+a15baa8

This is a nightly development build. It may contain bugs or incomplete features. Use at your own risk.


What's in This Build

Commit Details:

  • SHA: a15baa8
  • Message: fix: restore clipboard after ydotool clipboard-paste injection (#588)
  • Author: Mohamed Hashim
  • Date: 2026-08-05 22:08:13 -0700

Changes Since Last Nightly

  • a15baa8 fix: restore clipboard after ydotool clipboard-paste injection (#588)
  • 67dde78 docs(user-guide): note CUDA device 0 limit for dual NVIDIA (#644)
  • 5d06d21 fix(whisper): use CUDA device 0 when pywhispercpp is CUDA-backed (#636)
  • 7a5ddfd fix(ibus): restore engine after register_component teardown (#558) (#643)

Installation

Quick Install (Recommended)

Download the wheel file below and install:

pip install vocalinux-0.15.0.dev20260806+a15baa8-py3-none-any.whl

From PyPI (Stable)

For the latest stable release:

pip install vocalinux

Using Install Script

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh | bash

AppImage (no install, no root)

Download the .AppImage matching your CPU (x86_64 or aarch64) below:

chmod +x Vocalinux-0.15.0.dev20260806+a15baa8-x86_64.AppImage
./Vocalinux-0.15.0.dev20260806+a15baa8-x86_64.AppImage

Requires xdotool/wtype/ydotool on the host for text injection, same as the PyPI install path.


Build Information

Property Value
Version 0.15.0.dev20260806+a15baa8
Base Version 0.15.0
Build Date 20260806
Python 3.9+
OS Linux (X11 & Wayland)

Files

  • vocalinux-0.15.0.dev20260806+a15baa8-py3-none-any.whl - Python wheel (recommended)
  • vocalinux-0.15.0.dev20260806+a15baa8.tar.gz - Source distribution
  • Vocalinux-0.15.0.dev20260806+a15baa8-x86_64.AppImage - Self-contained, no install (Intel/AMD)
  • Vocalinux-0.15.0.dev20260806+a15baa8-aarch64.AppImage - Self-contained, no install (ARM64, added shortly after this release is published)

Links


Only the 3 most recent nightly builds are kept. This build will be automatically deleted when newer nightlies are created.

Generated on 2026-08-06 05:24:07 UTC

v0.15.0

Choose a tag to compare

@github-actions github-actions released this 02 Aug 22:29
fbb8123

Vocalinux v0.15.0

image image image

TL;DR:

Vocalinux 0.15.0 is a stable minor release on the 0.14 packaging line. Settings move to a searchable sidebar with dictation controls in the footer. GitHub Releases ship AppImages. The speech-language list is much larger (Hungarian included). Continuous dictation capitalizes sentences and leaves a space between utterances. There are new power and GPU controls, better Wayland IBus when ibus-wayland is present, and a fix for Bluetooth mic crashes during PortAudio probing.


Highlights

Area Description
Settings Sidebar + search instead of seven tabs; status, mic level, Test Dictation, and Close stay in the sidebar footer
AppImage x86_64 and aarch64 builds attached to the GitHub Release
Languages About 33 selectable speech languages plus Auto-detect (including Hungarian); VOSK only lists languages with official models
Dictation polish Capitalize after . / ! / ?; trailing space after each completed utterance
Auto-pause / keep-alive Unload the model while configured apps run, or after an idle timeout
Vulkan devices Prefer a discrete GPU when one exists; override under Advanced settings
ibus-wayland Prefer IBus on previously “unbridged” compositors when ibus-wayland is running
CLI vocalinux --version
Bluetooth mics Capture open/probe path no longer heap-corrupts on SCO/HFP devices

New Features

  • Settings: searchable sidebar instead of seven notebook tabs. (#601 by @jatinkrmalik)
  • Settings: recognition status, mic level, Test Dictation, and Close live in the sidebar footer so they stay visible while you change pages. (#618 by @jatinkrmalik)
  • Packaging: AppImage for x86_64 and aarch64, plus smoke-test packaging fixes. (#573, #602 by @jatinkrmalik)
  • Languages: expand the speech catalog (Hungarian and many more). CLI --language uses the same list as Settings. Official Alphacephei VOSK models are wired where they exist; Whisper-only languages stay out of the VOSK dropdown. (#616 by @jatinkrmalik, fixes #565)
  • Dictation: capitalize at the start and after sentence-ending punctuation. (#554 by @jatinkrmalik, closes #553)
  • Dictation: append a trailing space after each completed transcription so the next session does not glue onto the previous sentence. (#608 by @jatinkrmalik, fixes #605)
  • Auto-pause competing apps and idle model keep-alive unload. (#592 by @jatinkrmalik, closes #445, #591)
  • Vulkan: auto-select a discrete GPU and allow a manual device in Advanced settings. (#590 by @jatinkrmalik, closes #589)
  • Wayland: use IBus when ibus-wayland is running, including on compositors previously treated as unbridged. (#614 by @eiseleb47, closes #607)
  • CLI: vocalinux --version. (#563 by @AnayGarodia, closes #555)
  • Website: marketing redesign and a languages page with honest per-engine badges. (#582, #616 by @jatinkrmalik)

Bug Fixes

  • Settings: Custom Shortcut entry / Record / Set show again after the sidebar settings refactor. (#619 by @jatinkrmalik)
  • Languages: English (India) (en-in) maps to Whisper code en for whisper.cpp / Whisper / remote API. VOSK still uses en-in. (#617 by @jatinkrmalik)
  • Audio: stop Bluetooth mic probing from corrupting the heap. One PortAudio open per capture session, stop-before-close, and no stereo probe on mono-only devices. Fixes crashes on SCO/HFP headsets such as PipeWire “Bluetooth internal capture”. (#599 by @jatinkrmalik, fixes #567 reported by @HaleTom)
  • IBus: keep engine teardown correct when parent do_destroy fails. (#613 by @eiseleb47, fixes #606)
  • Settings UI: flatten info notices so helper text matches the rest of the dialog. (#615 by @jatinkrmalik)
  • KDE Plasma Wayland: skip unbridged IBus when ibus-wayland is not present so injection does not fail quietly. (#577 by @jatinkrmalik, fixes #574)
  • xdotool: keep input focus after injection. (#564 by @AnayGarodia, fixes #549)
  • Uninstall: remove the IBus data dir, stop the app via PID file, and remove ~/.local/bin launchers reliably. (#597 by @jatinkrmalik, #569 by @webenefits)
  • Installer: install both libgirepository 1.0 and 2.0 when the distro has them. (#583 by @jatinkrmalik, fixes #571)
  • AUR: depend on virtual python-pywhispercpp; move clipboard and wtype tools to optdepends. (#579 by @Asseel-Naji, #586 by @jatinkrmalik)
  • UI: first-run dialog response without Gtk.Dialog.do_response. (#580 by @jatinkrmalik, fixes #566)
  • Vosk: add Italian and English-India rows to the medium/large model tables. (#551 by @AnayGarodia, fixes #550)
  • Docs: fix ydotool service setup guidance. (#560 by @AnayGarodia, fixes #557)

Docs / maintenance

  • Multi-distro tray icon FAQ. (#584)
  • robots.txt no longer blocks indexable pages. (#610)
  • Website CI lint / action warning cleanup. (#611)
  • README badge layout tweaks. (#542, #544, #545)
  • Next.js dependency bump. (#587)

Thanks

Thank you @eiseleb47 for the IBus teardown fix and for teaching Vocalinux about ibus-wayland (#613, #614).

Thank you @AnayGarodia for --version, the xdotool focus fix, the Vosk model table entries, and the ydotool docs correction (#563, #564, #551, #560).

Thank you @webenefits for making uninstall remove the ~/.local/bin wrappers (#569).

Thank you @Asseel-Naji for the AUR python-pywhispercpp dependency fix (#579).

Thank you @HaleTom for reporting the Bluetooth mic malloc crash (#567).

Thanks to everyone who filed bugs and feature requests on the way to 0.15 and to everyone who tested builds.


Install / Upgrade

Recommended (install.sh)

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

AppImage

Download Vocalinux-0.15.0-x86_64.AppImage or Vocalinux-0.15.0-aarch64.AppImage from this release:

chmod +x Vocalinux-0.15.0-x86_64.AppImage
./Vocalinux-0.15.0-x86_64.AppImage

Host text-injection tools (xdotool, wtype, or ydotool) are still required, same as a PyPI install. FUSE is needed to run the AppImage on most hosts.

Arch Linux (AUR)

yay -S vocalinux

PyPI

pip install -U vocalinux

Flatpak

Local build notes live under packaging/flatpak/. Flathub may still trail this tag.

Upgrade details: docs/UPDATE.md



AppImage hotfix (2026-08-03)

Rebuilt and replaced Vocalinux-0.15.0-x86_64.AppImage and Vocalinux-0.15.0-aarch64.AppImage on this release.

The original 0.15.0 AppImages could fail to start on non-Debian distros (e.g. openSUSE) with a false "missing GTK3 / AppIndicator" error. Please re-download the AppImage if you hit that.

Details: #637

Full changelog: v0.14.2...v0.15.0

v0.14.2

Choose a tag to compare

@github-actions github-actions released this 18 Jul 01:43
3d65e5c

Vocalinux v0.14.2

Vocalinux 0.14.2 is a stability patch on the 0.14 series. Feature set is the same as 0.14.x: Flatpak packaging, AUR, configurable hotkeys, FunASR/SenseVoice remote API, layout-aware shortcuts, and the Wayland/IBus reliability work from this line. This release fixes IBus reliability bugs and makes the settings dialog fit normal monitors.


Highlights

Area Description
IBus engine launch Engine process starts after the Flatpak XDG path import change instead of dying with ImportError and falling back to ydotool/clipboard paste
First dictation Cold start on GNOME Wayland no longer drops the first utterance of a session (wait for FocusIn before commit)
Settings dialog Notebook tabs scroll so the window fits 1080p and similar displays; wheel events from unfocused widgets reach the tab scroller

Bug Fixes

  • IBus: Restore engine process launch after Flatpak XDG path import so the engine no longer dies with ImportError and falls back to ydotool/clipboard paste. (#534 by @jatinkrmalik; also reported as #539 by @dbyentzen)
  • IBus: Wait for FocusIn before commit on scoped injection so the first dictation of a session is not dropped on GNOME Wayland. (#533 by @jatinkrmalik, fixes #523 reported by @BlackEyedHatMan)
  • Settings UI: Make settings notebook tabs scrollable so the dialog fits the monitor instead of growing taller than the screen. (#538 by @BlackEyedHatMan)
  • Settings UI: Forward wheel events from unfocused combos/spins to the tab scroller and clean up nested Advanced ScrolledWindow shadows. (#541 by @jatinkrmalik)

CI / packaging (not user-facing)


Thanks

Thank you @BlackEyedHatMan for reporting the first-dictation drop (#523) and for shipping the settings dialog scroll fix (#538). Clear bug report plus a solid PR made this patch better.

Thank you @dbyentzen for reporting the IBus engine process / missing component registration failure on X11 (#539). That write-up helped confirm the same class of failure fixed in #534.

Thanks to everyone testing the 0.14 line and filing bugs. Reports like these are how the app gets better.


Install / Upgrade

Recommended (install.sh)

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

From the release tag

cd vocalinux
git fetch origin
git checkout v0.14.2
./install.sh

Arch Linux (AUR)

yay -S vocalinux

PyPI

pip install -U vocalinux

Flatpak

Flathub submission is still in review. Local build: see packaging/flatpak/. When Flathub lands: flatpak update com.vocalinux.Vocalinux.

Upgrade details: docs/UPDATE.md


Full changelog: v0.14.1...v0.14.2

v0.14.1

Choose a tag to compare

@jatinkrmalik jatinkrmalik released this 17 Jul 02:00
1783331

Vocalinux v0.14.1

Vocalinux 0.14.1 is a packaging and reliability release (stable tag of 0.14.1-beta for Flathub). You can install via Flatpak and AUR, and several installer and Wayland paste bugs that showed up right after 0.14.0 are fixed.


Highlights

Feature Description
Flatpak packaging Universal Flatpak on GNOME Platform 50 with whisper.cpp, global hotkeys, and Wayland paste (ydotool + wl-copy)
AUR package Official Arch packaging with CI publish
Layout-aware hotkeys Combo shortcuts respect non-US keyboard layouts
Installer + paste fixes Broken launch wrappers, sg on newer Ubuntu/Debian, ydotool 0.1.x vs 1.x Ctrl+V, Flatpak ydotoold socket probe

New Features

  • Flatpak packaging. GNOME Platform 50, whisper.cpp (Vulkan when available), evdev global hotkeys, and text injection via wl-copy + ydotool (xdotool fallback). Build from packaging/flatpak/. Flathub submission is in progress. (#484 by @jatinkrmalik, closes #167)
  • AUR release package. PKGBUILD and CI publish path for Arch Linux. (#518 by @jatinkrmalik; feature request #116 by @nxtkofi)

Bug Fixes

  • Hotkeys: Layout-aware combo keys so modifier+letter shortcuts work on non-US layouts. (#514 by @jatinkrmalik, fixes #513 reported by @laurentpayot)
  • Installer: Expand EXEC_CMD correctly in ~/.local/bin wrappers so vocalinux no longer fails with a missing command after install.
  • Installer: Honor absolute --venv-dir paths (no nested repo/home/.../venv).
  • Installer: Avoid requiring sg when it is missing on Ubuntu 26.04 / Debian 13. (#524 by @jatinkrmalik, fixes #522 reported by @OpenUAS)
  • Text injection: Treat XIM none as unset so injection does not skip useful backends. (#512 by @jatinkrmalik, fixes #511 reported by @canis-lupus03)
  • Wayland paste (ydotool): Use ctrl+v on distro ydotool 0.1.x and keycode chords on ydotool 1.x (Flatpak pins 1.0.4). Stops host installs typing garbage like 2442 instead of pasting.
  • Wayland paste (Flatpak): Probe ydotoold with a real connect on the Unix datagram socket, drop only dead sockets, and start the daemon when needed. Fixes paste exit status 2 after a previous session.
  • Models: HTTP timeouts and clearer errors when Hugging Face is slow or returns HTML; HF downloads use ?download=true.
  • Web: Dependabot npm fixes in the website lockfile. (#515 by @jatinkrmalik; also #510 by @dependabot)

Docs

  • Refreshed v0.14 UI screenshots and website gallery. (#521 by @jatinkrmalik)
  • README Star History and related polish.

Since 0.14.1-beta

  • Flathub domain verification: empty well-known file at /.well-known/org.flathub.VerifiedApps.txt on vocalinux.com (#530 by @jatinkrmalik)
  • Stable release tag for Flathub: drop the -beta AppStream/version suffix so Flathub stable lint accepts the build (#532 by @jatinkrmalik)

Install / Upgrade

Recommended (install.sh)

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

From the release tag

cd vocalinux
git fetch origin
git checkout v0.14.1
./install.sh

Arch Linux (AUR)

yay -S vocalinux

Flatpak (local build; Flathub pending review)

flatpak install flathub org.gnome.Platform//50 org.gnome.Sdk//50
git clone https://github.com/jatinkrmalik/vocalinux.git && cd vocalinux
git checkout v0.14.1
flatpak-builder --user --force-clean --install build-dir packaging/flatpak/com.vocalinux.Vocalinux.yml
flatpak run com.vocalinux.Vocalinux

Notes

  • Flatpak and host (install.sh) installs should not be run at the same time; they can fight over the ydotool socket and the mic.
  • First-run model downloads need network access to Hugging Face. If HF is down, retry later or keep using a model already under ~/.local/share/vocalinux/models/.

Thanks

Thanks to everyone who contributed to this release:

Pull Requests:

Issue reporters:

Full changelog: v0.14.0-beta...v0.14.1

v0.14.0-beta

Choose a tag to compare

@github-actions github-actions released this 14 Jul 06:44
8a085b5

Vocalinux v0.14.0-beta

Vocalinux 0.14.0-beta adds configurable keyboard shortcuts and FunASR/SenseVoice support for the Remote API engine. It also fixes several Wayland text-injection problems, an audio crash on startup, and whisper.cpp using too many CPU cores on hybrid laptops.


Highlights

Feature Description
Configurable Shortcuts Bind any modifier combination to a key: Alt+R, Ctrl+Shift+V, or Super+F10
FunASR / SenseVoice Remote API Remote-API engine now supports FunASR and SenseVoice models via OpenAI-compatible endpoints
GNOME Wayland IBus Reliability Text injection works again on GNOME Wayland with bare xkb layouts and engine restore fallbacks are fixed
Audio Crash Fix Recording no longer crashes when the system audio device index changes between sessions
Hybrid-CPU Efficiency whisper.cpp no longer defaults to all cores on hybrid Intel/AMD processors

New Features

  • Configurable modifier+key hotkeys. The Settings dialog now lets you set custom shortcuts using any combination of Ctrl, Alt, Shift, and Super plus a letter/number key. The legacy defaults still work, and you can now bind combinations like Alt+R or Ctrl+Shift+V. (#493 by @subsy)
  • Remote API FunASR/SenseVoice support. OpenAI-compatible remote endpoints can specify FunASR/SenseVoice model names (e.g. sensevoice) and return richer response shapes; SenseVoice metadata labels are stripped before text injection. (#468, closes #469 reported by @LauraGPT)

Bug Fixes

  • GNOME Wayland/IBus: Restore text injection when only a bare xkb engine is configured; the engine restore fallback now picks the correct IM engine instead of silently dropping text. (#506 fixes #504 reported by @tubemakersc1-tech, #500 fixes #497 reported by @ynoguti)
  • KDE Wayland/IBus: Restore the KDE Plasma Wayland IBus text-injection path that was regressed in recent compositor-detection changes. (#502 fixes #501 reported by @rightpad)
  • Wayland injection: Wait for held modifiers (Ctrl/Alt/Shift/Super) to release before injecting text, preventing accidental shortcut triggers and garbled output on modifier-heavy workflows. (#494 by @subsy)
  • Shortcuts UI: Keep preset and custom shortcut selection exclusive. Selecting a preset now clears the custom field, and setting a custom combo selects the "Custom Shortcut" preset. (#509)
  • whisper.cpp: Stop defaulting to all CPU cores on hybrid processors (Intel Performance + Efficient cores), which caused UI lag and excess battery drain. (#492 by @only21mil)
  • Audio: Fix a crash on recording start when the selected audio device index no longer matches the current system enumeration. (#499 closes #498 reported by @devnix)
  • Installer: Include xsel as a fallback for the Wayland clipboard path when xclip is unavailable. (#496 fixes #477 reported by @fanofjop)

Improvements

  • Code style. Removed an outdated long comment about whisper.cpp default thread counts. (#505)
  • CI. Added a codecov.yml with a 0.5% project coverage threshold so docs-only and version-bump PRs are not blocked by minor coverage fluctuations.

Install / Upgrade

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

Or update from source:

cd vocalinux
git fetch origin
git checkout v0.14.0-beta
./install.sh

Thanks

Thanks to everyone who contributed to this release:

Pull Requests:

Issue Reporters:

Detailed changelogs and nightly builds are always available on the releases page.

Full changelog: v0.13.0-beta...v0.14.0-beta

v0.13.0-beta

Choose a tag to compare

@github-actions github-actions released this 02 Jul 00:35
63c1053

This release is all about making dictation feel solid on Wayland, plus a much smarter whisper.cpp model picker. If you're on a wlroots-style compositor (COSMIC, sway, Hyprland, niri), a non-US keyboard layout, or you've ever watched text silently fail to appear, this one is for you.

A lot of what's here came straight from people filing clear bug reports and sending fixes. Thank you. 🙏

✨ New features

  • Guided whisper.cpp model selection. The Settings dialog now splits whisper.cpp into a Model Size and a Specialization picker, so you can choose English-only, quantized (Q5/Q8), Large v3 Turbo, or legacy large variants, with language-aware recommendations and hover hints. You can also pass exact model IDs like medium.en-q5_0 or large-v3-turbo straight to --model. (#465, closes #461, thanks @rafaeln)

🐛 Reliability fixes (mostly Wayland)

  • Silent text drops on wlroots/COSMIC compositors. Vocalinux now uses wtype/ydotool instead of IBus on compositors that don't bridge IBus to native apps (COSMIC, sway, Hyprland, and similar), so injection stops silently doing nothing. (#486, closes #485, thanks @bbbboother)
  • IBus false-positive that swallowed text. On Wayland we now require a real IM engine before trusting IBus, so a bare xkb:* layout (including on GNOME/Mutter) no longer reports success while inserting nothing. (#491, closes #478, thanks @subsy)
  • Keyboard layout flipping to us after dictation. We no longer run setxkbmap on Wayland, where it was quietly rewriting the XWayland layout and leaving XWayland/Electron apps on the wrong keys. (#490, closes #474, thanks @bcs-23, and @fsioni and @karottenreibe for tracing the cause)
  • Garbled output on non-US layouts (AZERTY/QWERTZ/Dvorak) and a clipboard-copy hang. ydotool now pastes through the clipboard, which is layout-independent and Unicode-safe. (#480, thanks @fsioni)
  • Faster ydotool typing. An explicit --key-delay noticeably cuts injection latency. (#488, closes #482, thanks @subsy)
  • Hotplugged keyboards now work. The evdev shortcut backend rescans, so a keyboard plugged in after startup is picked up and reconnects recover. (#467, closes #431, thanks @LuigiKraken)
  • Dictation spacing preserved. Spacing between speech segments separated by a pause is kept, so words no longer run together after a silence. (#464)
  • KDE Plasma Wayland guidance. Vocalinux detects KDE Plasma Wayland and points you to enable IBus Wayland when wtype injection fails, both in the installer and in the logs. (#466, closes #463, thanks @cyb3rko)

🔧 Under the hood

  • Refreshed website docs with dedicated pages for the Remote API engine, Silero VAD, advanced whisper.cpp settings, and desktop reliability. (#470)
  • Automatic PR labeling by changed files. (#473)

🚀 Try it

New install or update (recommended, it handles system deps, models, and desktop integration):

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh | bash

From source:

git clone https://github.com/jatinkrmalik/vocalinux.git
cd vocalinux && ./install.sh

Already running a source checkout:

cd vocalinux && git fetch origin && git checkout v0.13.0-beta && ./install.sh

To try the new model picker, open Settings, go to Speech Engine, and choose a Model Size plus a Specialization. Or from the CLI: vocalinux --model large-v3-turbo.

💚 Thank you

Vocalinux keeps getting better because of the people who use it and push back on it. Big thanks to the contributors who sent code this cycle, @fsioni and @bbbboother, and to everyone whose reports and feature requests turned into these fixes: @rafaeln, @cyb3rko, @LuigiKraken, @subsy, @bcs-23, and @karottenreibe. This is a project I rely on every single day, and it genuinely wouldn't be this good without you.

If something breaks, or you have an idea for what would make Vocalinux better, please open an issue or start a discussion. Feedback and suggestions are always welcome.

Full changelog: v0.12.0-beta...v0.13.0-beta

v0.12.0-beta

Choose a tag to compare

@github-actions github-actions released this 07 Jun 07:23
748d318

Vocalinux v0.12.0-beta: Remote API Recognition, Silero VAD, and Reliability Hardening

New remote transcription backend, neural voice activity detection, and thread safety improvements across the board.


If you are new here, Vocalinux is a free, open-source (GPL-3.0) voice dictation system for Linux. It runs 100% offline, works on both X11 and Wayland, and lets you control your computer and dictate text using just your voice. Under the hood, it uses whisper.cpp (the default engine), OpenAI Whisper, VOSK, or now a Remote API for speech recognition.

Today we are shipping v0.12.0-beta, a release that expands where your dictation can run and improves the core recognition pipeline. Here is what changed.


What's New

Remote API Engine: Dictate to Your Own Server

The headline feature in this release is the new Remote API speech recognition engine. If you have a self-hosted whisper.cpp server, an OpenAI-compatible API endpoint, or any compatible remote transcription service, you can now point Vocalinux at it.

The Remote API engine supports two formats:

  • whisper.cpp server: the /inference endpoint format
  • OpenAI-compatible: the /v1/audio/transcriptions endpoint format

Configuration lives in the Speech Engine tab alongside your existing local engines. Switch between local and remote with a single dropdown. The engine includes HTTP connection pooling, suspend/resume support, and comprehensive documentation.

This feature came from community requests for cloud-based dictation while keeping Vocalinux's privacy-first philosophy. You control where your audio goes.

Silero VAD: Smarter Speech Detection

We replaced the simple amplitude-threshold voice activity detection with Silero VAD, a neural network model specifically trained for speech detection.

The bundled Silero VAD v5 model (via ONNX Runtime) provides:

  • More accurate speech/non-speech classification: no more false triggers from keyboard clicks or background noise
  • Silence buffer dropping: recordings that contain only silence are now discarded entirely
  • Existing sensitivity control: the familiar vad_sensitivity (1-5) slider maps to Silero thresholds internally

Silero VAD activates automatically when onnxruntime is available. Without it, Vocalinux falls back to the amplitude-based VAD. Enable it with:

pip install "vocalinux[vad]"

Thread Safety: Fewer Crashes Under Load

The recognition pipeline got a serious threading pass:

  • Remote API HTTP sessions: lock-guarded session switching prevents use-after-close races during reconfiguration
  • IBus engine: _state_lock protection eliminates TOCTOU races in the socket handler
  • Text injection: concurrent reconfigure/reinitialize/resume cycles no longer crash

These fixes resolve crashes reported during rapid engine switching and system suspend/resume scenarios.

IBus: Dead Keys Work Again

IBus integration got two critical fixes for international keyboard layouts:

  • Dead key composition preserved: Vocalinux now temporarily activates its engine only while injecting text, preserving your previous engine for dead-key composition. This fixes ^+eê on French AZERTY and similar layouts
  • Scoped engine capture: the engine is now captured correctly during prepare_engine, fixing a regression where IBus injection silently fell back to ydotool on Wayland

Settings Dialog Polish

The settings UI received two quality-of-life improvements:

  • Remote Server section: now properly gated behind the Advanced toggle instead of appearing for all users
  • Dialog height reduced: from 50% to 40% of viewport height, preventing overflow on sub-1080p screens

Installer: CUDA Diagnostics with Auto-Remediation

The installer's CUDA diagnostics now:

  • Auto-remediate missing patchelf: installs it automatically when needed
  • Support ARM64 CUDA toolkit paths: works on ARM-based systems
  • Include behavioral tests: validates the diagnostics actually work

Model Metadata Corrections

Download sizes for whisper.cpp and VOSK models were corrected across the UI:

  • tiny: 39 MB → 74 MB
  • large: 1.5 GB → 3.0 GB
  • Other models updated to match actual download sizes

Bug Fixes by Category

Speech Recognition

  • Add Silero VAD with silence-only buffer dropping (#447)
  • Correct whisper.cpp and VOSK model download size metadata (#453)

Remote API

  • Add Remote API speech recognition engine with installation and configuration support (#335)
  • Harden thread safety for Remote API HTTP session switching (#452)

IBus

  • Preserve user engine for dead keys during scoped activation (#457)
  • Capture current engine during prepare_engine for scoped activation (#458)
  • Harden thread safety for IBus engine state protection (#452)

Text Injection

  • Harden thread safety for text injector concurrent operations (#452)

UI/Settings

  • Respect Advanced toggle for Remote Server section (#454)
  • Reduce settings dialog height to prevent overflow on lower-resolution screens (#456)

Installer

  • Harden CUDA diagnostics with auto-remediation and behavioral tests (#451)

Startup

  • Allow launch without pynput backend (#448)

Website

  • Clarify speech demo browser support (#449)

Under the Hood

  • Developer tooling: Remote API test server script (scripts/test_remote_server.py) for testing both API formats without a real server, documented in CONTRIBUTING.md
  • Website typecheck fix: added local react-syntax-highlighter declarations so the lazy-loaded syntax highlighter passes TypeScript checks
  • Community: GitHub Sponsors funding configuration added for those who want to support the project
  • Behavioral test coverage: CUDA diagnostics and release-facing reliability fixes include targeted tests

Contributors

This release represents 13 commits from our community. Thank you to everyone who contributed code, reported bugs, and suggested improvements.

Code Contributors

  • @jatinkrmalik (Jatin K Malik): Remote API thread safety, Silero VAD integration, IBus dead-key fixes, UI polish, installer hardening, model metadata corrections
  • @sevenjay: Remote API speech recognition engine (#335)
  • @Sbenazar: Original Silero VAD implementation (#351, carried forward in #447)
  • @juanfradb: Play stop sound immediately on release (#426)
  • @peplocanto (Pep Locanto): openSUSE Tumbleweed dependency fixes (#420)

Issue Reporters

Features and fixes in this release were inspired by reports from:


What is Vocalinux?

Vocalinux is a free, open-source voice dictation system for Linux. It lets you:

  • Dictate text into any application using your voice
  • Issue voice commands for common actions (copy, paste, select, delete, and more)
  • Run 100% offline: no cloud services, no data leaving your machine (unless you choose Remote API)
  • Choose your engine: whisper.cpp (fast, local, default), OpenAI Whisper, VOSK, or Remote API
  • Work on any display server: full support for both X11 and Wayland

The project is licensed under GPL-3.0 and welcomes contributions from the community.


Installation

The fastest way to install Vocalinux:

curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

For detailed installation instructions, including distro-specific notes, see the README.


Try It Out

If you run into issues, file a bug on GitHub. If you want to contribute, pull requests are always welcome.

Star the repo if you find it useful. It helps more people discover the project.

Thanks for being part of the Vocalinux community. Happy dictating.