Whisper Desktop v3.0.0
Whisper Desktop v3.0.0 ✨
Whisper Desktop v3.0.0 marks our largest and most ambitious milestone to date. What began as a streamlined desktop transcription tool has evolved into an all-in-one, privacy-first local media studio — combining high-performance offline speech recognition, an enterprise-grade AI subtitle translation suite, and frame-accurate, hardware-accelerated video subtitle mastering.
Every layer of the application has been overhauled for native desktop speed, zero background overhead, and complete accessibility across Windows, macOS, and Linux.
⚡ v3.0.0 at a Glance (TL;DR)
- 🎬 Studio Hardsub 2.0: A complete canvas-powered video subtitle mastering suite with pixel-perfect preview, local streaming, dynamic hardware encoder probing, and granular typography controls.
- 🌐 AI Subtitle Translation Studio: Independent bilingual workspace supporting 122 languages, synchronized scrolling, cognitive thinking-tag stripping, and seamless connection to cloud & local LLMs.
- 🚀 Universal Dynamic CPU Dispatch: Zero-crash runtime architecture. Automatically detects your processor's vector capabilities (from legacy SSE4.1 up to AVX-512 & NEON) with zero illegal instruction errors.
- 🤖 5-Role Hardware Model Hub: Real-time hardware profiler that recommends the exact right Whisper model for your RAM and GPU (NVIDIA CUDA, AMD Vulkan, Intel Arc/OpenVINO, Apple Metal).
- 🌍 13 Fully Localized Languages & Native RTL: Complete UI translations with symmetrical RTL layouts, content-aware text direction, and Unicode RLM punctuation anchoring.
🌟 Major Highlights
🚀 Universal Hardware Runtime & Dynamic CPU Dispatch
- Dynamic Instruction Set Dispatch (Zero-Crash Guarantee): Architected the core execution pipeline with
GGML_CPU_ALL_VARIANTSandGGML_BACKEND_DL. The application packages all CPU vector extensions (AVX-512, AVX2, AVX, FMA, F16C, SSE4.1, baseline x86-64, and ARM64 NEON) and dynamically loads the optimal instructions for your specific host CPU at runtime. - Legacy & Modern CPU Parity: Completely eliminates
STATUS_ILLEGAL_INSTRUCTION(Windows0xc000001d) andSIGILL(Linux) crashes on older processors (e.g. Intel Core 2 Quad, 1st-gen Core i, legacy Xeons) while seamlessly unlocking maximum hardware vector throughput on cutting-edge architectures. - Automated Runtime Library Resolution: Engine runners automatically propagate companion shared libraries (
.so,.dll) and manage dynamic linker paths (LD_LIBRARY_PATH,DYLD_LIBRARY_PATH, andPATH) across AppImage, DEB, RPM, NSIS, MSI, and Portable ZIP distributions. - Silent Child Processes on Windows: Suppressed console window flicker across all child process invocations (
whisper-cli,ffmpeg,ffprobe,nvidia-smi,taskkill) using nativeCREATE_NO_WINDOWflags. - Auto-Provisioned VC++ Runtimes: The Windows setup wizard automatically verifies and installs the Microsoft Visual C++ 2015–2022 x64 Redistributable if absent.
🎬 Studio Hardsub 2.0 — Real-Time Video Subtitle Mastering
- Canvas-Powered 2D ASS/SSA Rendering Engine: Completely replaced legacy DOM/CSS overlays with a custom, high-fidelity 2D Canvas renderer. Directly mirrors
libassgeometry, margins, line wrapping, and font metrics for guaranteed 1:1 pixel parity between player preview and the final exported video. - Embedded Local HTTP Streaming Server: Built-in ultra-lightweight loopback server (
127.0.0.1with ephemeral security tokens) supporting RFC 7233 byte-range and suffix requests. Delivers instantaneous frame-accurate scrubbing, freeze-frame inspection, and zero playback lag even with high-bitrate 4K media. - Dynamic Hardware Codec Probing: Probes hardware accelerators (NVIDIA NVENC, Intel QSV, Apple VideoToolbox, Linux VA-API, and CPU) alongside individual video codecs (H.264, HEVC, AV1, ProRes) via fast parallel one-frame validation passes before listing, dynamically narrowing dropdown choices to verified codecs.
- Structured FFmpeg Diagnostics & Error Classification: Classifies FFmpeg encoder failures into 8 structured categories in the Rust backend, delivering actionable, localized recovery guidance rather than raw terminal stack traces.
- Safe In-App Fullscreen: Eliminated WebKitGTK full-viewport crashes in Linux AppImage environments by synchronizing in-app player controls with native window fullscreen states without DOM reparenting.
- Granular Media Reset & Concurrency Guards: Added dedicated Clear buttons on video/subtitle dropzones and a master "Reset All" action, backed by generation counters and serial guards to cleanly isolate and cancel in-flight preview decoding tasks.
- Companion Subtitle Auto-Detection: Automatically discovers matching subtitle files (
.srt,.vtt,.ass, including language-tagged variants) when loading videos, with intelligent fallback to preserve user selections. - Typography Polish & Style Toggles: Full support for Bold and Italic toggles in preview and ASS export, static Montserrat 400 Regular font bundling, dynamic background box padding based on typographic metrics (
textAscent/textDescent), and rounded-corner bounding boxes. - Native RTL Text Shaping & Punctuation Anchoring: Engine-level bidirectional run layout (
hardsubLayout.ts) and Unicode RLM (\u200F) punctuation anchoring — Persian, Arabic, and Hebrew subtitles never suffer from flipped punctuation marks, reversed parentheses, or broken glyph connections. - Linux AppImage Media Stabilization: Resolved GStreamer ALSA/PipeWire latency conflicts via
alsasinkranking andPULSE_LATENCY_MSEC=60, enabledTCP_NODELAYon streaming sockets to eradicate Nagle Delayed-ACK stalls, eliminated plugin scan stalls with background registry warmup, and sanitizedLD_LIBRARY_PATHagainst host library poisoning. - Live FFmpeg Command Inspection: Full transparency into the encoding pipeline with an instant one-click copyable FFmpeg command string.
🌐 Standalone AI Subtitle Translation Studio
- Independent Dual-Pane Workspace: Dedicated bilingual translation environment featuring side-by-side cue cards, real-time driver-locked synchronized scrolling (zero lag even on 10,000+ cue files), and inline caption editing.
- Content-Aware Bidirectional Card Layout: Independent, per-column text direction analysis (
textDirection.ts). Translating between RTL (Persian/Arabic) and LTR (English/French) renders both panes with native typographical alignment without interface clipping or reversed alignment. - Universal Parameter Negotiation: Seamlessly communicates with OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Groq, and local OpenAI-compatible endpoints (Ollama, LM Studio, vLLM). Automatically adapts parameter payloads (e.g. mapping
max_tokensvsmax_completion_tokens, stripping unsupported temperatures). - Cognitive Tag Stripping: Automatically parses and strips
<think>,<thought>, and<reasoning>thinking blocks from next-gen reasoning models (DeepSeek-R1, QwQ, Claude thinking) before passing text to line and timestamp alignment. - Crash-Resilient Atomic Checkpoints: Translates in SHA-256 fingerprinted transactional blocks. If interrupted, translations resume right where they stopped without corrupting files or consuming duplicate API tokens.
- OS-Level Keyring Security: Sensitive API keys are securely stored in your operating system's native keychain (Linux Secret Service / KWallet, Windows Credential Manager, macOS Keychain) — never written to plaintext config files.
- Pure Subtitle Text Export: One-click clipboard copying of clean translated subtitle text without timestamps or line index clutter.
- Direct Studio Handoff: Pass a finished translation directly into Studio Hardsub with a single click — companion video files are linked automatically.
- 122 Target Languages & Polish Pass: Instant-searchable language catalog with native-script names, dynamic reasoning effort parameter negotiation, and an optional second-pass rewriter tuned for natural subtitle cadence and rhythm.
🤖 Intelligent Hardware-Aware Model Hub
- 5-Role Hardware Recommendation Engine: Automatically inspects physical CPU core counts, system RAM, and GPU architecture (distinguishing discrete AMD RX/Pro from APUs, Intel Arc from integrated Iris Xe, NVIDIA CUDA, and Apple Metal) to map your hardware to 5 distinct operational profiles:
- Balanced Daily Driver: Optimal equilibrium between processing speed and linguistic accuracy for everyday workflows.
- Studio Quality: Maximum semantic precision and acoustic noise resilience for challenging audio recordings.
- Ultra-Fast: Near-instantaneous transcription with ultra-low memory footprints.
- English Specialist: Dedicated English-only architecture delivering 10–15% higher precision on English audio.
- Efficient Quantized: Compressed integer models slashing RAM and battery usage by 40–60%.
- Expanded GGML & VAD Catalog: 34 Whisper models (Tiny, Base, Small, Medium, Large-v1/v2/v3,
large-v3-turbo,.envariants, andq5_0/q5_1/q8_0quantizations) + 2 Silero VAD releases (v5.1.2&v6.2.0). - Speaker Diarization (
small.en-tdrz): Built-in catalog support for tinydiarize models enabling native speaker turn detection without external Python pipelines. - High-Throughput Async Rust Downloader: Completely replaced curl with a streaming Rust pipeline featuring live progress, speed metrics, ETA, pause/resume capability, and strict SHA-256 checksum verification.
- Pinned Segmented Filter Bar & Specifications: Responsive layout with quick category filters, detailed precision ratings, and automatic detection of existing
~/whisper.cpp/modelsfolders.
⚡ Transcription Studio & Core Refinements
- Direct Instant Landing View: Removed the redundant dashboard and background polling HUD to eliminate idle CPU/GPU draw and eradicated memory leaks. The application now boots directly into the Transcription Studio.
- Quick Configuration Deck: Prominent toolbar for fast adjustments to model selection, hardware backend, and spoken language without digging into deep menus.
- Auto-Expanding Transcript Cue Cards: Multi-line live transcripts dynamically expand to fit text content, preventing zero-height collapses and text clipping, with segment indices, duration badges, and LTR time isolation.
- Resilient Batch Queue: Sequential batch queue management with per-item status indicators, drag-and-drop reordering, duration sorting, and graceful individual item failure isolation.
- Auto-Translate After Transcription Pipeline: Option to automatically send finished transcripts directly to the Translation Studio.
- Configurable FFmpeg Engine: Seamless toggle between bundled static FFmpeg binaries and host system FFmpeg with automatic path escaping and version caching.
- 8 Export Formats & Output Folder Opener: Export transcripts to SRT, VTT, TXT, JSON, Full JSON, ASS, LRC, and CSV. Files land alongside source media by default with configurable directory fallback, middle-filename truncation for long paths, and one-click folder opening badges.
- Native Clipboard Integration: Integrated
tauri-plugin-clipboard-managerwith multi-tier fallback for reliable one-click copying of transcripts, logs, and FFmpeg commands.
🌍 Complete 13-Language Internationalization & Native RTL
- 13 Fully Localized Languages: Complete UI translations for English, Persian, Spanish, French, German, Simplified Chinese, Japanese, Russian, Arabic, Portuguese, Italian, Turkish, and Korean.
- True Symmetrical RTL Engine: Dedicated RTL stylesheet (
rtl.css) and directional layout logic ensuring navigation sidebars, modals, form controls, sliders, and video overlays mirror correctly for Persian and Arabic. - Content-Aware Direction Resolution: Subtitle cue cards dynamically evaluate text direction based on the actual subtitle content (
hardsubLayout.tsandtextDirection.ts), allowing seamless editing of mixed RTL and LTR files. - Tactile Toggle Switch Consistency: Enforced persistent LTR directional orientation on toggle switches within RTL modes to maintain an intuitive tactile mental model.
- Persian Typography Polish: Strict adherence to zero-width non-joiners (ZWNJ / half-spaces) and Hamza orthography across all application strings.
🎨 Design System, Theming & Desktop Ergonomics
- 4 Bespoke Design Themes: Choose between Royal Blue (default), Carbon, Emerald, and Fiery Orange, each accompanied by its own matching pre-rendered SVG brand icon in the sidebar.
- Titanium Zero-FOUC Splash Boot Sequence: Native dark background and color-scheme configuration eliminate white flashes on application startup.
- Native UI Scaling & Zoom: Full zoom control (Ctrl/Cmd + +, -, 0) with persistent scale factor across reboots.
- Deep Desktop Integration: Minimize to system tray on close, native taskbar progress bars during transcription/encoding, active job exit protection, and modernized About modal with official vector wave branding.
- Central Activity Log Ring Buffer: In-memory diagnostic logs with category filtering, smart auto-scroll pause during manual reading, jump-to-bottom pill, and quick clipboard export.
- Linux Packaging Enhancements: Atomic zstd compression for RPM and DEB packages, plus precompiled Arch Linux AUR package (
whisper-desktop-bin).
🧠 whisper.cpp Core Engine Sync
The embedded speech recognition engine has been synchronized with upstream whisper.cpp:
- VAD Timeline Synchronization: Fixed token timestamp drift when Silero VAD is active; preview, export, and burned-in subtitle timings remain perfectly locked to the original audio.
- Clean Plain Text Output: Eliminated the legacy leading whitespace quirk on exported
.txtfiles. - Sub-Second Audio Safety: Resolved an upstream heap out-of-bounds read error during spectrogram processing of ultra-short audio clips.
- Intel OpenVINO Optimization: Fixed cross-attention scaling on Intel Arc and Iris Xe GPUs for significantly cleaner transcription.
- Deterministic Fallback Passes: Systematic decoder re-seeding guarantees 100% reproducible results during temperature fallback iterations.
- Hardware Compute Kernels: Vectorized Flash-Attention (
fa-vec) tuned for Apple Silicon (Metal) and optimized shared memory swizzle tiles for NVIDIA CUDA & Vulkan backends.
📦 Downloads & Installation Hub
Tip
Which edition should you choose?
- Universal Edition: Ideal for all standard setups (AMD Radeon, Intel Arc / Iris Xe / Core Ultra, Apple Silicon, and CPU-only systems). Dynamically leverages Vulkan, OpenVINO, Metal, or optimized AVX2/NEON backends.
- NVIDIA CUDA Edition: Best for PCs with dedicated NVIDIA GeForce RTX / GTX graphics cards (ships with pre-configured CUDA 12 runtimes for maximum inference speed).
🪟 Windows (x64 & ARM64 via Prism)
Compatible with Windows 10, Windows 11, and Windows Server 2019+
Universal Edition (AMD Radeon, Intel Arc / Iris Xe, Standard CPU)
— Standard setup wizard (auto-installs Visual C++ 2015–2022 redistributable if missing)
— Zero-install standalone archive (run from any folder or USB drive)
— Windows Installer package for enterprise managed environments
NVIDIA CUDA Edition (Dedicated RTX / GTX Hardware Acceleration)
— Full installer with bundled CUDA 12 runtimes
— Standalone portable archive with CUDA acceleration
— Enterprise MSI with CUDA acceleration
🍏 macOS (Apple Silicon M1–M4 & Intel)
Compatible with macOS 12 Monterey, macOS 13 Ventura, macOS 14 Sonoma, and macOS 15 Sequoia
🐧 Linux (x86_64 & ARM64)
Compatible with Ubuntu 20.04+, Debian 11+, Fedora 38+, Arch Linux, and RHEL 9+
Arch Linux (AUR)
yay -S whisper-desktop-binx86_64 Universal Edition (Vulkan, Intel OpenVINO, CPU)
— Standalone binary; runs on any modern distribution
— Native package for Debian, Ubuntu, and Linux Mint (zstd compressed)
— Native package for Fedora, RHEL, and openSUSE (zstd compressed)
x86_64 NVIDIA CUDA Edition (Dedicated RTX / GTX Acceleration)
— Portable AppImage with bundled CUDA acceleration
— Native
.debpackage with CUDA support— Native
.rpmpackage with CUDA support
ARM64 (aarch64) Universal Edition (SBCs & ARM Linux Desktops)
— Standalone AppImage for Raspberry Pi 5, Rockchip, and ARM64 servers
— Native
.debfor 64-bit ARM Debian/Ubuntu— Native
.rpmfor 64-bit ARM Fedora/RHEL
Note
💬 Community Feedback & Bug Reports
Whisper Desktop is designed to run entirely offline across an enormous variety of operating systems and hardware configurations — from modern workstations with dedicated GPUs to legacy laptops, Apple Silicon Macs, and ARM SBCs.
Because it is impossible for a single developer to physically test every hardware combination, driver setup, and operating system environment, your feedback is invaluable!
If you encounter any bugs, performance issues, UI glitches, or have feature ideas on any platform (Windows, Linux, or macOS):
- Please open an issue on GitHub.
- Helpful context: your OS, CPU/GPU model, and diagnostic logs from the in-app Activity Logs view.
Every bug report, test result, and suggestion directly helps make Whisper Desktop more stable, robust, and accessible for everyone. Thank you for helping improve the project! 💙
Empowering your media locally — enjoy private, high-performance transcription, translation, and video subtitling! ✨
Full Changelog: v2.3.0...v3.0.0