Repository navigation
What's everyone using for Speech to Text? #3872
Replies: 8 comments 15 replies
|
I’m currently using hyprwhspr with push-to-talk (SUPER+F2) and a Waybar widget for state/tooltip feedback, plus audio cues. Because I’m on an older GTX 1070 Ti (Pascal), I moved transcription to a local REST backend (Parakeet) to avoid whisper.cpp/CUDA build issues — but that also means my GPU isn’t being used, which is exactly the kind of “same problem” you’re describing. I think? If Voxtype is the direction, what I’d love is: A stable Waybar integration contract: a command that prints Waybar JSON (text, class, tooltip) reflecting state (idle/recording/transcribing) + active model/backend/device. Haveing support for legacy GPU backends that works on older NVIDIA via Vulkan (so Pascal actually accelerates locally) + optional notifications/audio cues built-in. That would let me keep my current UX (SUPER + F2/hotkey + visual Waybar feedback) and make it so I can take full advantage of my older 1070 Ti without Python/venv/CUDA like you say. |
|
I've been mucking around with your code, comparing it to Hyperwhspr and trying to add a few Waybar/Hyperland features |
|
Here's an idea of what I've been playing with so far: Omarchy_251218_voxtype-showcase_v4.md The Transcribings: "Oh, there we go. And then it starts transcribing and you see a little bit of dynamic feedback. So you know that you're actually transcribing. Because before I was doing that and it was annoying because I would be talking to myself and then I realized I hit the key and it's not transcribing and it just started. So I lost a whole bunch of stuff that I had said. So I thought, I wonder if we can make this connect in there, which I guess with the flexibility of Arch and Amarchi and all that. It's fairly easy to do, especially with AI these days. And so I've been just messing around with this. The one thing I've noticed is that you, whenever, there must be some kind of bug with Signal Messenger because it seems to lose the pipe wire set up on, at least on my computer. And I don't know, some weird thing with Signal, which hopefully DHH will figure out or somebody will. What I wanted is both, you know, things to basically coexist at the same time and not interfere with each other. And I wanted it to work with my old NVIDIA 1070 TI card, which I guess is Vulcan, as it turns out. So it took me a while to figure out how to get it to compile so that it would because lately, most of the distros have pushed all of the GPU code for NVIDIA into the 10, 5, you know, what is it, 580 stuff. And so then you kind of have to write, you know, maintain it separately as a separate, then what do you call it, like in Python, you kind of create a separate environment with a separate set of drivers. So anyways, that's what I've done. I'll see if I can share that with you too. It would be nice if, if it could figure that out. I don't think without the help of AI, I wouldn't have figured out how to recompile it. But, but, you know, this stuff is getting a lot smarter these days. Yeah, so then I press F 14. You'll see it'll change to listening, has a similar little thing that's supposed to be moving, it's not as sensitive on this one. Um and uh it starts transcribing and eventually if I let go " |
|
/.config/voxtype/config.toml /.config/hyprwhspr/config.json |
|
Oh and one more thing I've been meaning to ask you. If I activate transcribe and say nothing and quickly press the button again. I often get the words "Thank you" I'm told that the thank you is something to do with open whisper or parakeet, but it's kind of weird. It doesn't happen on the um on the hyperwhspr, it only seems to happen on Voxtype. Yeah, anyway, kinda weird. Thought I'd ask. Maybe they've got some kind of code that um because it's kind of that the open whisper model is kind of thinking, so basically if you ask it to do something and you don't give it any input, it says thank you. I don't know, it's weird. |
|
Have you seen any of these other Local Open-Source Speech-to-Text Alternatives? Executive Overview of Local Speech-to-Text Alternatives (Updated)This overview summarizes key offline-capable speech-to-text (STT) tools suitable for high-accuracy local transcription, focusing on those that excel in handling large vocabularies, jargon, slang, and coding terms. All primary entries are open source and run fully locally on Linux setups like Omarchy, emphasizing compatibility with projects such as Voxtype (Rust-based push-to-talk STT for Wayland) or HyprWhspr (Hyprland-integrated STT). Based on the latest as of December 18, 2025, these tools are vetted for excellence (e.g., low WER on benchmarks like LibriSpeech/TED-LIUM, with fine-tuning support for domain-specific accuracy). To keep your Omarchy install fresh, align with @dhh's recent posts on X—no new updates since December 15, 2025, so v3.2.3 remains the current release, with ongoing polish and community fixes. Run
All licenses are non-viral (except MPL-2.0's file-level copyleft for Coqui), making them suitable for inclusion in open-source projects like Voxtype/HyprWhspr without forcing your code to adopt the same terms. They support free redistribution of models/weights, enabling offline deployment on Omarchy without cloud ties. For installation on Omarchy, prioritize AUR/yay methods (e.g., Stack Ranking for Inclusion in Projects Like Voxtype or HyprWhsprRanking prioritizes: (1) Ease of integration into Rust-based Linux tools (native bindings > Python bridges via PyO3/subprocess); (2) Active maintenance and community support (for long-term Omarchy freshness); (3) Resource efficiency on typical PC hardware (CPU-friendly > GPU-heavy); (4) Out-of-box accuracy/robustness to jargon/slang/coding terms (via fine-tuning); (5) Overall alignment with Omarchy's lightweight, community-driven ethos per @dhh's updates. Higher ranks are better for quick prototyping and production use in Wayland/Hyprland environments.
For your setup, start with Vosk in Voxtype (edit Addendum: Integration Potential of Microsoft Cognitive Services Speech SDK (Azure Speech SDK)The Azure Speech SDK, developed by Microsoft, is a proprietary software development kit that exposes Azure AI Speech service capabilities for building speech-enabled applications, including speech-to-text (STT), text-to-speech, translation, and speaker recognition. It supports both real-time and batch processing using local devices, files, Azure Blob Storage, or streams, making it versatile for scenarios like dictation in tools such as VS Code's Speech extension.
|
|
Have you seen this one being used in Goose AI called Faster Whisper? It looks cool too? Interesting: ### Faster Whisper Benchmarks (2025) Faster Whisper (CTranslate2-optimized Whisper) excels in local STT, with 2025 benchmarks showing strong accuracy/speed trade-offs: Accuracy (WER - Word Error Rate): Large-v3-turbo model achieves 7.7% WER on LibriSpeech (batched processing), outperforming base Whisper variants; distil-large-v3 at ~5-8% on English clean speech, robust to noise/accent/jargon (e.g., 95%+ on coding dictation per Reddit tests). Speed: 6x faster than standard Whisper Large-v3 (e.g., transcribes 1-hour audio in 30 seconds on M3-Max or equivalent NVIDIA GPU); real-time (1x) on CPU for medium model, up to 10x with quantization. Comparisons: Tops open-source charts (e.g., Northflank 2025 guide: 809M params for turbo, edges wav2vec2 on English tech audio); Modal benchmarks favor it for throughput vs. insanely-fast-whisper (hardware-intensive). Resource: ~1-5GB RAM, CPU-viable but GPU boosts (e.g., CUDA on Omarchy: yay -S cuda if NVIDIA). |
|
This is awesome, great job! Hyprwhspr has tons of issues with mise. Voxtype solved them all! |
Uh oh!
There was an error while loading. Please reload this page.
I have been following the interest in #2467 for speech-to-text integration. I like the work goodroot has done on hyprwhspr, but I am not personally a big fan of python implementations for this kind of utility. Having tried a number of other python STT tools, I get caught up in dependency hell pretty quickly.
After not finding exactly what I wanted in the community, I built Voxtype as a Rust and Whisper-based speech-to-text tool. I'm not trying to compete with hyprwhspr, but offering an alternative with some different trade-offs that might appeal to the Omarchy philosophy.
Voxtype is built as a single Rust binary because I prefer compiled tools rather than interpreted ones for stuff like this. I found Python venvs painful to manage, and often ran into dependency conflicts or incompatibility issues.
It's offline-only by design, so you'll have the best experience if you have a decent GPU. GPU support via the Vulkan backend works across GPU vendors. I ran in to some ROCm 7.x compatibility issues when trying out hyprwhspr and other similar python implementations.
I decided that I wanted desktop notifications in addition to a waybar widget to give me visual feedback when recording/transcribing. And also audio cues.
Trade-offs:
Happy to prepare a PR if there's interest, or help evaluate both options. The community clearly wants speech-to-text - just offering another option to consider.
https://voxtype.io | https://github.com/peteonrails/voxtype
All reactions