Skip to content

Releases: AtomicError/whisper-desktop

Whisper Desktop v3.0.1

Choose a tag to compare

@github-actions github-actions released this 30 Sep 21:40

Whisper Desktop v3.0.1 is a focused maintenance and polish release delivering substantial usability enhancements to the Hardsub Studio, smoother video preview playback, refined RTL/UI ergonomics, and hardened cross-platform packaging.


🌟 What's New & Highlights

🎬 Hardsub Studio & Video Player Enhancements

  • Direct Video Playback Action: Added a dedicated "Play Exported Video" action button immediately upon completing a hardsub render, allowing instant verification in your default media player.
  • Persistent & Synchronized Volume Controls: The video player volume slider is now permanently visible with refined spacing and zero desynchronization between initial audio levels and player controls.
  • Playback Stall Auto-Recovery: Implemented an intelligent watchdog mechanism that detects and automatically recovers stalled or frozen video preview streams.
  • Ergonomic Subtitle Cue Editor: Cue textareas now dynamically auto-expand to fit multiline content with an ergonomic 48px minimum height floor, eliminating awkward nested scrollbars.
  • RTL Canvas Overlay Fix: Eliminated initial horizontal canvas offsets in Right-to-Left (RTL) mode, ensuring subtitles align with millimeter precision from the first frame.

🎨 UI Refinements & Internationalization

  • Ultra Reasoning Support: Integrated the Ultra Reasoning tier in the translation engine, complete with proper RTL layout alignment and robust settings null-safety.
  • Adaptive Themed Scrollbars: Implemented modern, unobtrusive scrollbars styled to match the dark theme and eliminated layout shifting (scrollbar gutter polish).
  • Stable Transcribe Layout: Locked grid dimensions in the transcription studio to prevent responsive jitter across different window sizes.

⚙️ CI/CD & Cross-Platform Packaging Hardening

  • Accelerated Windows Builds: Enabled parallel multi-core MSVC compilation (-j) across all Whisper backend targets in the CI workflow.
  • Linux AppImage Packaging Fixes: Resolved linuxdeploy failures concerning Level Zero loader symlinks and guaranteed explicit POSIX executable permissions across all bundled runtimes.
  • Bundled Runtime Assets: Improved distribution bundles with companion OpenVINO backend plugins and hardware cache files.

📝 Commit Changelog

  • 9920c7f ci(linux): resolve AppImage linuxdeploy failure with Level Zero loader and symlink hygiene
  • a636957 ci(windows): enable parallel multi-core compilation for whisper backends
  • 713f0ed feat(hardsub): add direct play exported video action button
  • bebb55d style(ui): harmonize reasoning dropdown width and implement adaptive themed scrollbar
  • d86eaa0 fix(ui): add ultra reasoning, fix rtl reasoning alignment and settings null-safety
  • d6e2431 fix(ui): polish provider cards i18n and hardsub scrollbar gutter
  • bddb66b fix(ui): enforce 48px ergonomic floor for hardsub cue textareas
  • de6fef1 fix(hardsub): eliminate cue textarea internal scroll and auto-fit heights
  • 80a0b64 fix(ui): auto-expand hardsub cue textareas and lock transcribe grid layout
  • 98372de chore(release): bump application version to 3.0.1
  • a206561 fix(hardsub): polish preview status banner and package OpenVINO cache on Windows
  • 413ba11 fix(hardsub): polish volume controls and auto-recover stalled video playback
  • bd3ab05 fix(hardsub): synchronize initial video audio and tighten volume slider spacing
  • a7f9595 feat(hardsub): keep video player volume slider permanently visible
  • ba5aff6 fix(hardsub): prevent initial canvas overlay offset in RTL and first-load
  • 90aca13 fix(ci): enforce POSIX execute permissions on AppImage repackaging
  • 6d55e39 docs(screenshots): update hardsub and models preview assets
  • fc026d7 ci(workflow): bundle OpenVINO device backend plugins and GPU cache

📦 Downloads & Installation Hub

Tip

Which edition should you choose?

  • Universal Edition: Ideal for all standard setups (AMD Radeon, Intel Arc / Iris Xe / Core Ultra, Apple Silicon, and CPU-only systems). Dynamically leverages Vulkan, OpenVINO, Metal, or optimized AVX2/NEON backends.
  • NVIDIA CUDA Edition: Best for PCs with dedicated NVIDIA GeForce RTX / GTX graphics cards (ships with pre-configured CUDA 12 runtimes for maximum inference speed).

🪟 Windows (x64 & ARM64 via Prism)

Compatible with Windows 10, Windows 11, and Windows Server 2019+

Universal Edition (AMD Radeon, Intel Arc / Iris Xe, Standard CPU)

  • Setup Installer — Standard setup wizard (auto-installs Visual C++ 2015–2022 redistributable if missing)
  • Portable ZIP — Zero-install standalone archive (run from any folder or USB drive)
  • MSI Package — Windows Installer package for enterprise managed environments

NVIDIA CUDA Edition (Dedicated RTX / GTX Hardware Acceleration)

  • CUDA Setup Installer — Full installer with bundled CUDA 12 runtimes
  • CUDA Portable ZIP — Standalone portable archive with CUDA acceleration
  • CUDA MSI Package — Enterprise MSI with CUDA acceleration

🍏 macOS (Apple Silicon M1–M4 & Intel)

Compatible with macOS 12 Monterey, macOS 13 Ventura, macOS 14 Sonoma, and macOS 15 Sequoia

  • Universal DMG — Native Mach-O Universal binary with Metal GPU acceleration for Apple Silicon and Intel Macs

🐧 Linux (x86_64 & ARM64)

Compatible with Ubuntu 20.04+, Debian 11+, Fedora 38+, Arch Linux, and RHEL 9+

Arch Linux (AUR)

yay -S whisper-desktop-bin

x86_64 Universal Edition (Vulkan, Intel OpenVINO, CPU)

  • AppImage — Standalone binary; runs on any modern distribution
  • Debian / Ubuntu Package — Native package for Debian, Ubuntu, and Linux Mint (zstd compressed)
  • Fedora / RHEL Package — Native package for Fedora, RHEL, and openSUSE (zstd compressed)

x86_64 NVIDIA CUDA Edition (Dedicated RTX / GTX Acceleration)

  • CUDA AppImage — Portable AppImage with bundled CUDA acceleration
  • CUDA Debian / Ubuntu Package — Native .deb package with CUDA support
  • CUDA Fedora / RHEL Package — Native .rpm package with CUDA support

ARM64 (aarch64) Universal Edition (SBCs & ARM Linux Desktops)

Read more

Whisper Desktop v3.0.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 20:31

Whisper Desktop v3.0.0 ✨

Whisper Desktop v3.0.0 marks our largest and most ambitious milestone to date. What began as a streamlined desktop transcription tool has evolved into an all-in-one, privacy-first local media studio — combining high-performance offline speech recognition, an enterprise-grade AI subtitle translation suite, and frame-accurate, hardware-accelerated video subtitle mastering.

Every layer of the application has been overhauled for native desktop speed, zero background overhead, and complete accessibility across Windows, macOS, and Linux.


⚡ v3.0.0 at a Glance (TL;DR)

  • 🎬 Studio Hardsub 2.0: A complete canvas-powered video subtitle mastering suite with pixel-perfect preview, local streaming, dynamic hardware encoder probing, and granular typography controls.
  • 🌐 AI Subtitle Translation Studio: Independent bilingual workspace supporting 122 languages, synchronized scrolling, cognitive thinking-tag stripping, and seamless connection to cloud & local LLMs.
  • 🚀 Universal Dynamic CPU Dispatch: Zero-crash runtime architecture. Automatically detects your processor's vector capabilities (from legacy SSE4.1 up to AVX-512 & NEON) with zero illegal instruction errors.
  • 🤖 5-Role Hardware Model Hub: Real-time hardware profiler that recommends the exact right Whisper model for your RAM and GPU (NVIDIA CUDA, AMD Vulkan, Intel Arc/OpenVINO, Apple Metal).
  • 🌍 13 Fully Localized Languages & Native RTL: Complete UI translations with symmetrical RTL layouts, content-aware text direction, and Unicode RLM punctuation anchoring.

🌟 Major Highlights

🚀 Universal Hardware Runtime & Dynamic CPU Dispatch

  • Dynamic Instruction Set Dispatch (Zero-Crash Guarantee): Architected the core execution pipeline with GGML_CPU_ALL_VARIANTS and GGML_BACKEND_DL. The application packages all CPU vector extensions (AVX-512, AVX2, AVX, FMA, F16C, SSE4.1, baseline x86-64, and ARM64 NEON) and dynamically loads the optimal instructions for your specific host CPU at runtime.
  • Legacy & Modern CPU Parity: Completely eliminates STATUS_ILLEGAL_INSTRUCTION (Windows 0xc000001d) and SIGILL (Linux) crashes on older processors (e.g. Intel Core 2 Quad, 1st-gen Core i, legacy Xeons) while seamlessly unlocking maximum hardware vector throughput on cutting-edge architectures.
  • Automated Runtime Library Resolution: Engine runners automatically propagate companion shared libraries (.so, .dll) and manage dynamic linker paths (LD_LIBRARY_PATH, DYLD_LIBRARY_PATH, and PATH) across AppImage, DEB, RPM, NSIS, MSI, and Portable ZIP distributions.
  • Silent Child Processes on Windows: Suppressed console window flicker across all child process invocations (whisper-cli, ffmpeg, ffprobe, nvidia-smi, taskkill) using native CREATE_NO_WINDOW flags.
  • Auto-Provisioned VC++ Runtimes: The Windows setup wizard automatically verifies and installs the Microsoft Visual C++ 2015–2022 x64 Redistributable if absent.

🎬 Studio Hardsub 2.0 — Real-Time Video Subtitle Mastering

  • Canvas-Powered 2D ASS/SSA Rendering Engine: Completely replaced legacy DOM/CSS overlays with a custom, high-fidelity 2D Canvas renderer. Directly mirrors libass geometry, margins, line wrapping, and font metrics for guaranteed 1:1 pixel parity between player preview and the final exported video.
  • Embedded Local HTTP Streaming Server: Built-in ultra-lightweight loopback server (127.0.0.1 with ephemeral security tokens) supporting RFC 7233 byte-range and suffix requests. Delivers instantaneous frame-accurate scrubbing, freeze-frame inspection, and zero playback lag even with high-bitrate 4K media.
  • Dynamic Hardware Codec Probing: Probes hardware accelerators (NVIDIA NVENC, Intel QSV, Apple VideoToolbox, Linux VA-API, and CPU) alongside individual video codecs (H.264, HEVC, AV1, ProRes) via fast parallel one-frame validation passes before listing, dynamically narrowing dropdown choices to verified codecs.
  • Structured FFmpeg Diagnostics & Error Classification: Classifies FFmpeg encoder failures into 8 structured categories in the Rust backend, delivering actionable, localized recovery guidance rather than raw terminal stack traces.
  • Safe In-App Fullscreen: Eliminated WebKitGTK full-viewport crashes in Linux AppImage environments by synchronizing in-app player controls with native window fullscreen states without DOM reparenting.
  • Granular Media Reset & Concurrency Guards: Added dedicated Clear buttons on video/subtitle dropzones and a master "Reset All" action, backed by generation counters and serial guards to cleanly isolate and cancel in-flight preview decoding tasks.
  • Companion Subtitle Auto-Detection: Automatically discovers matching subtitle files (.srt, .vtt, .ass, including language-tagged variants) when loading videos, with intelligent fallback to preserve user selections.
  • Typography Polish & Style Toggles: Full support for Bold and Italic toggles in preview and ASS export, static Montserrat 400 Regular font bundling, dynamic background box padding based on typographic metrics (textAscent/textDescent), and rounded-corner bounding boxes.
  • Native RTL Text Shaping & Punctuation Anchoring: Engine-level bidirectional run layout (hardsubLayout.ts) and Unicode RLM (\u200F) punctuation anchoring — Persian, Arabic, and Hebrew subtitles never suffer from flipped punctuation marks, reversed parentheses, or broken glyph connections.
  • Linux AppImage Media Stabilization: Resolved GStreamer ALSA/PipeWire latency conflicts via alsasink ranking and PULSE_LATENCY_MSEC=60, enabled TCP_NODELAY on streaming sockets to eradicate Nagle Delayed-ACK stalls, eliminated plugin scan stalls with background registry warmup, and sanitized LD_LIBRARY_PATH against host library poisoning.
  • Live FFmpeg Command Inspection: Full transparency into the encoding pipeline with an instant one-click copyable FFmpeg command string.

🌐 Standalone AI Subtitle Translation Studio

  • Independent Dual-Pane Workspace: Dedicated bilingual translation environment featuring side-by-side cue cards, real-time driver-locked synchronized scrolling (zero lag even on 10,000+ cue files), and inline caption editing.
  • Content-Aware Bidirectional Card Layout: Independent, per-column text direction analysis (textDirection.ts). Translating between RTL (Persian/Arabic) and LTR (English/French) renders both panes with native typographical alignment without interface clipping or reversed alignment.
  • Universal Parameter Negotiation: Seamlessly communicates with OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Groq, and local OpenAI-compatible endpoints (Ollama, LM Studio, vLLM). Automatically adapts parameter payloads (e.g. mapping max_tokens vs max_completion_tokens, stripping unsupported temperatures).
  • Cognitive Tag Stripping: Automatically parses and strips <think>, <thought>, and <reasoning> thinking blocks from next-gen reasoning models (DeepSeek-R1, QwQ, Claude thinking) before passing text to line and timestamp alignment.
  • Crash-Resilient Atomic Checkpoints: Translates in SHA-256 fingerprinted transactional blocks. If interrupted, translations resume right where they stopped without corrupting files or consuming duplicate API tokens.
  • OS-Level Keyring Security: Sensitive API keys are securely stored in your operating system's native keychain (Linux Secret Service / KWallet, Windows Credential Manager, macOS Keychain) — never written to plaintext config files.
  • Pure Subtitle Text Export: One-click clipboard copying of clean translated subtitle text without timestamps or line index clutter.
  • Direct Studio Handoff: Pass a finished translation directly into Studio Hardsub with a single click — companion video files are linked automatically.
  • 122 Target Languages & Polish Pass: Instant-searchable language catalog with native-script names, dynamic reasoning effort parameter negotiation, and an optional second-pass rewriter tuned for natural subtitle cadence and rhythm.

🤖 Intelligent Hardware-Aware Model Hub

  • 5-Role Hardware Recommendation Engine: Automatically inspects physical CPU core counts, system RAM, and GPU architecture (distinguishing discrete AMD RX/Pro from APUs, Intel Arc from integrated Iris Xe, NVIDIA CUDA, and Apple Metal) to map your hardware to 5 distinct operational profiles:
    1. Balanced Daily Driver: Optimal equilibrium between processing speed and linguistic accuracy for everyday workflows.
    2. Studio Quality: Maximum semantic precision and acoustic noise resilience for challenging audio recordings.
    3. Ultra-Fast: Near-instantaneous transcription with ultra-low memory footprints.
    4. English Specialist: Dedicated English-only architecture delivering 10–15% higher precision on English audio.
    5. Efficient Quantized: Compressed integer models slashing RAM and battery usage by 40–60%.
  • Expanded GGML & VAD Catalog: 34 Whisper models (Tiny, Base, Small, Medium, Large-v1/v2/v3, large-v3-turbo, .en variants, and q5_0/q5_1/q8_0 quantizations) + 2 Silero VAD releases (v5.1.2 & v6.2.0).
  • Speaker Diarization (small.en-tdrz): Built-in catalog support for tinydiarize models enabling native speaker turn detection without external Python pipelines.
  • High-Throughput Async Rust Downloader: Completely replaced curl with a streaming Rust pipeline featuring live progress, speed metrics, ETA, pause/resume capability, and strict SHA-256 checksum verification.
  • Pinned Segmented Filter Bar & Specifications: Responsive layout with quick category filters, detailed precision ratings, and automatic detection of existing ~/whisper.cpp/models folders.

⚡ Transcription Studio & Core Refinements

  • Direct Instant Landing View: Removed the redundan...
Read more

Whisper Desktop v2.3.0

Choose a tag to compare

@github-actions github-actions released this 26 Jul 16:37

🚀 Whisper Desktop v2.3.0

v2.3.0 is officially live! A major release featuring complete UI/UX modernizations, hybrid TypeScript architecture, real-time downloader telemetry, and an upgraded AI translation engine.


🌟 What's New in This Release?

🎛️ 1. UI Modernization & Top-Layer Portal Dropdowns

  • Top-Layer Dropdown Menus: Select menus now render seamlessly in the top layer, completely eliminating dropdown clipping across all overflow boundaries and scroll containers.
  • Unified Batch Processing Queue: The multi-file batch queue has been redesigned into a cohesive glassmorphic panel featuring an integrated header bar, live file count badges, and aligned action controls.
  • Refined Visual Affordance: Improved card hover behavior across all views—eliminating white flashes on static layout containers while preserving soft accent highlights on interactive controls.
  • Interactive Toast Notifications: Enhanced system notifications with animated progress countdowns, hover-to-pause interactions, and quick close controls.

📥 2. Real-Time Model Downloader & Telemetry

  • Event-Driven Progress Streaming: Download progress updates dynamically in real time without refreshing the page or interrupting navigation across other tabs.
  • Standardized Mbps Telemetry: Download speeds are now calculated and displayed in standard Megabits per second (Mbps).
  • Smart ETA Calculations: Estimated completion times exceeding 60 minutes automatically format into clean hours, minutes, and seconds.
  • Smoother Speed Accuracy: Baseline file tracking eliminates initial speed spikes when resuming partial downloads.
  • Clean Pause & Resume: Pausing an active download is tracked gracefully without generating false error notifications.

🌐 3. Multi-Language AI Subtitle Translation Suite

  • 40+ World Languages Support: Expanded translation support for over 40 global languages.
  • Sliding Context Window: Keeps conversational context intact across long media subtitle segments to ensure natural, narrative dialogue.
  • High-Fidelity Dialogue Prompting: Subtitle translations preserve speaker formatting, punctuation, line numbers, and exact pronoun fidelity.
  • Real-Time Chunk Progress: Stream line-by-line translation progress directly inside the interface as AI processing occurs.

⚡ 4. Modern TypeScript & Vite Toolchain

  • Hybrid TypeScript Support: The frontend build system now supports TypeScript, enabling modern, strongly typed module development alongside existing JavaScript components.
  • Blazing Fast Build Performance: Integrated Vite tooling for near-instant development updates and optimized production builds.

🛠️ 5. Engine Reliability & Diagnostics

  • Stabilized Output Monitoring: Improved stdout reader stability during active transcription sessions.
  • Automatic Silero VAD Synchronization: Automated state alignment between Voice Activity Detection settings and command-line execution parameters.

📜 Full Commit Log

  • be78b13 chore(release): bump version to v2.3.0
  • 4447bfb feat(infra): configure typescript and vite hybrid environment for future module development
  • 5d317e7 feat(dashboard & downloader): update dashboard feature highlights, refactor model downloader events, and fix pause/speed telemetry
  • e115982 fix(ui): refactor batch queue card layout and implement top-layer portal positioning for custom selects
  • eec30be style(ui): eliminate noisy white card hovers and standardize affordance across components
  • 6ed6eb4 style(ui): redesign models storage directory field into responsive read-only path control
  • 14b4b57 feat(translation): expand target languages list to support 40+ major languages
  • b092a70 fix(translation): balance system prompt for complete fidelity and natural spoken dialogue
  • d23433a feat(translation): emit per-chunk progress events and cap chunk size at 30 items
  • 301e34c fix(transcribe): prevent premature stdout reader loop termination and fix misleading abort notification
  • 277a77e feat(translation): implement sliding context window and conversational subtitle prompt optimization
  • c1b44dd feat(ui): upgrade toast notifications with longer durations, progress bar countdown, hover-pause, and close button

📥 Getting Started

Download the package corresponding to your Linux distribution from the Assets section below:

  • Ubuntu/Debian: .deb
  • Fedora/RedHat: .rpm
  • Portable Linux: .AppImage

Arch Linux (AUR):
Arch Linux users can install the application directly from the AUR:

paru -S whisper-desktop-bin

(or using your preferred AUR helper like yay)


Enjoy high-performance, private, and offline speech-to-text transcription! 🎙️✨

Full Changelog: v2.2.0...v2.3.0

Whisper Desktop v2.2.0

Choose a tag to compare

@github-actions github-actions released this 19 Jul 18:17

🚀 Whisper Desktop v2.2.0 - Standalone Linux Packaging, Wayland Support, and Smart Model Hub

Welcome to Whisper Desktop v2.2.0! This major release focuses on enhancing the local user experience on Linux (Wayland & X11), making installation completely dependency-free, and introducing smart, flexible model directory scanning.


📦 Key Features & Improvements

⚡ 1. Precompiled & Bundled Acceleration Backends (No Local Builds!)

  • Zero Compilation Required: All four high-performance computation backends (Standard CPU, Vulkan GPU, OpenVINO Intel, and NVIDIA CUDA) are now precompiled automatically on our GitHub Actions pipeline and bundled directly inside the installer resources.
  • Static Linking Fix: Resolved the runtime dynamic linker error (exit code 127: libwhisper.so.1 cannot open shared object file) by statically compiling the whisper.cpp core engine. Whisper Desktop is now a completely self-contained application on Linux!

🎨 2. Wayland & X11 Taskbar Integration

  • Wayland Taskbar Matching: Configured the native GTK application ID mapping, enabling flawless window grouping and app icon rendering on modern Wayland compositors (such as Hyprland, GNOME, and KDE).
  • X11 Launcher Mapping: Configured StartupWMClass matching in the desktop file template to fix dock grouping on X11 environments.

📂 3. Smart Recursive Model Scanning

  • No Subfolder Restrictions: Removed the hardcoded models/ subfolder requirement. You can now store your model files directly in the root of your custom Models Storage Directory.
  • Recursive Detection: The scanner now traverses the selected folder recursively to find model files, while ignoring hidden (.git), dependency (node_modules), and build (target, build) directories to maintain lightning-fast UI performance.
  • Legacy Backward Compatibility: Seamlessly checks, downloads, tracks, and deletes existing models located in old models/ subfolders so your existing setup won't be disrupted.

🛠️ 4. Cross-Platform Stability & UI Polish

  • Windows Path Parsing: Integrated a new cross-platform path parsing utility to avoid file matching bugs under Windows environments.
  • Cleaner Configuration UI: Removed misleading code badge labels (selectedBackend and modelsDir) and simplified descriptions to make options easier to read.

📥 Getting Started

Download the package corresponding to your Linux distribution from the Assets section below:

  • Ubuntu/Debian: .deb
  • Fedora/RedHat: .rpm
  • Portable Linux: .AppImage

Arch Linux (AUR):
Arch Linux users can install the application directly from the AUR:

paru -S whisper-desktop-bin

(or using your preferred AUR helper like yay)


Enjoy high-performance, private, and offline speech-to-text transcription! 🎙️✨

Whisper Desktop v2.0.0

Choose a tag to compare

@github-actions github-actions released this 11 Jul 23:59

Whisper Desktop v2.0.0

A ground-up redesign plus AI translation!

🤖 AI Translation — The headline feature

Translate your transcriptions through OpenAI, Anthropic, Gemini, or any OpenAI-compatible provider. Manage multiple providers with keyring-backed API keys, configure per-provider models with reasoning levels (None/Low/Medium/High), and let the system auto-detect context windows and recover from token-limit errors automatically. Original subtitle formatting (SRT, VTT, LRC, TXT) is preserved throughout. Preview lines before translating, and cancel mid-request with immediate HTTP abort.


🚀 Other Major Features

Premium UI Overhaul

  • Complete glassmorphic redesign with royal-blue gradient accents
  • Collapsible sidebar with staggered stagger animations
  • 3-step wizard for transcription flow with auto-convert
  • Custom-styled dropdowns, toggles, radio buttons, scrollbars throughout
  • Responsive layout stacking for small viewports
  • Disabled context menu on non-input elements

Voice Activity Detection (VAD)

  • Built-in Silero VAD model download and auto-usage
  • Configurable threshold, speech tolerance, minimum segment duration
  • Reduces processing time on silent sections

Model Hub Enhancements

  • Integrated model downloader with pause/resume/delete
  • Real-time speed, ETA, and progress tracking
  • Search and category filtering
  • Spec-guided recommendations based on GPU/hardware

Batch Processing

  • Batch queue with drag reordering and sort controls
  • Sequential per-file progress with detailed status
  • Error dialogs for per-file retry/skip/abort

🛠️ Engineering & Hardening

  • 35+ correctness fixes across Rust backend and JS frontend
  • Replaced all .unwrap() with .expect() on regex/path invariants
  • XSS hardening — DOM builder pattern replacing innerHTML with textContent
  • Tauri event listener lifecycle management with scoped cleanup
  • TDZ fixes for timer/handler/state declarations
  • IPC rate limiter removed (caused log line drops during FFmpeg bursts)
  • Sequential batch probing (fixed parallel race condition)
  • Error notifications for settings load/save failures
  • Strip ANSI codes from log output
  • Null-safety guard for non-string inputs

📦 Downloads

Assets below — .deb, .rpm, and .AppImage for Linux.
Enjoy local, private speech-to-text transcription!

Full Changelog: v1.0.1...v2.0.0

Whisper Desktop v1.0.1

Choose a tag to compare

@github-actions github-actions released this 03 Jun 01:03

Release Notes (v1.0.1)

Whisper Desktop v1.0.1 focuses on major Linux compatibility fixes, specifically addressing graphics, sandboxing, and library dependency conflicts when running the AppImage bundle on various distributions.

🚀 What's New & Fixed

🐛 Linux & AppImage Compatibility Fixes

  • Resolved Graphics Initialization Crash (EGL_BAD_PARAMETER): Switched to a post-build purging script that extracts the generated AppImage and purges conflicting system graphics/Wayland libraries (libEGL, libGL, libgbm, libglapi, libwayland-client). This forces the WebKit webview to use the host's native GPU drivers.
  • Fixed WebKit subprocess ICU Crashes (libicudata.so.74 not found): Enabled bundling of system ICU v74 libraries inside the AppImage target and dynamically configured LD_LIBRARY_PATH inside the backend entrypoint to resolve them at runtime.
  • Eliminated GLIBC Mismatches (GLIBC_PRIVATE symbol errors): Reverted from the experimental truly-portable Tauri CLI to the standard stable Tauri CLI compiler, ensuring system base libraries (libc.so.6) are loaded from the host OS.
  • Disabled WebKit Sandboxing inside AppImage: Programmatically set WEBKIT_DISABLE_SANDBOX=1 inside main.rs when executed as an AppImage to allow helper processes (WebKitNetworkProcess / WebKitWebProcess) to spawn successfully.
  • CI/CD Pipeline Robustness: Updated the release runner to resolve FUSE permissions on virtualized environments by extracting appimagetool and running it in a FUSE-free mode. Wrapped path variables in double quotes to prevent spacing errors during packaging.

📚 Documentation & Packaging

  • Updated README.md with instructions on installing Whisper Desktop directly on Arch Linux via the AUR package (paru -S whisper-desktop-bin).
  • Standardized release filenames and paths in the documentation.

Full Changelog: v1.0.0...v1.0.1

Whisper Desktop v1.0.0

Choose a tag to compare

@github-actions github-actions released this 02 Jun 01:08

🎙️ Whisper Desktop v1.0.0 - Initial Release

Welcome to the initial release of Whisper Desktop! This is a premium, state-of-the-art, and gorgeous native GUI for running high-performance local speech-to-text models using whisper.cpp.


✨ Key Features in this Release

  • ⚡ Hardware Acceleration Backends: Choose between CPU, Vulkan, OpenVINO, or CUDA directly from the UI to maximize transcription speed.
  • 📂 Sequential Batch Queue: Load multiple audio/video files, see total queue specs, and transcribe them one by one.
  • 🎤 Live Microphone Recording: Record audio in real-time with responsive custom wave visualizations.
  • 🎞️ Media Converter: Seamlessly extract high-quality audio tracks from video files using integrated utilities.
  • 🎨 Cyber-Neon Glassmorphic UI: A dark-mode design with fluid micro-animations and intuitive layout states.
  • 🔒 100% Local & Private: No external API requests, no cloud uploads. Your data stays completely on your machine.

📦 Installation

Choose the package matching your Linux distribution from the Assets section below:

  • Debian/Ubuntu: Download the .deb file and install with sudo dpkg -i.
  • Fedora/RedHat: Download the .rpm file and install with sudo dnf install.
  • Universal Linux: Download the .AppImage file, make it executable (chmod +x), and run.