Skip to content

Releases: KrishRVH/speakeasy

Speakeasy 0.3.4

Choose a tag to compare

@github-actions github-actions released this 05 Oct 02:12
3f10dac

Release notes

Speakeasy 0.3.4 reduces application work on the path from finishing a recording to recognition.

  • Sealed audio reaches the speech worker before native microphone teardown finishes. A replacement
    recording still waits for the previous microphone to finish cleanup.
  • Speech windows are classified while recording. Sealing uses the accumulated classification,
    preserving the recording and audible-speech gates, quiet-edge trimming, word padding, and interior
    pauses.
  • Native insertion rechecks released shortcut modifiers every 2 ms.

Existing settings, engines, models, filler removal, and dictation gestures are preserved. Recordings
retain the five-minute limit. Audio and recognition stay local, with no transcript history or
uploads.

Windows x64 and Apple silicon macOS zip packages, plus experimental Linux x86_64 AppImage and
bundled tar packages, are attached with SHA-256 checksums. CI verifies all three platforms and runs
owned-window rendering checks. Live microphone, keyboard-hook, clipboard/editor, and compositor
acceptance remain separate; see architecture and performance.
Native end-to-end latency gains remain unmeasured.

Speakeasy 0.3.3

Choose a tag to compare

@github-actions github-actions released this 02 Oct 15:34
f307912

Release notes

Speakeasy 0.3.3 removes standalone English “um” and “uh” locally before inserting dictation.

  • Remove um / uh is enabled by default. Turn it off in Settings for literal transcription or
    non-English Parakeet dictation. Saving this preference keeps the model loaded.
  • Both speech engines share a linear cleanup pass with at most one output allocation. Cleanup
    handles case and pause punctuation, preserves compounds and individually quoted tokens, and sends
    nothing for filler-only recognition.
  • Whisper removes fillers only with its language set to English. Automatic and other language
    requests preserve literal words. Parakeet detects languages automatically; its cleanup switch
    expresses your English dictation preference.

Existing settings, engines, models, and dictation gestures are preserved. Recordings retain the
five-minute limit. Audio and recognition stay local, with no transcript history or uploads.

Windows x64 and Apple silicon macOS zip packages, plus experimental Linux x86_64 AppImage and
bundled tar packages, are attached with SHA-256 checksums. CI verifies all three platforms and runs
owned-window rendering checks. Live microphone, keyboard-hook, clipboard/editor, and compositor
acceptance remain separate; see architecture and performance.

Speakeasy 0.3.2

Choose a tag to compare

@github-actions github-actions released this 02 Oct 00:10
cd302b4

Release notes

Speakeasy 0.3.2 brings the codebase to an idiomatic Rust standard, gives every thread, process, and
native event loop one explicit owner, and fixes small Settings, pill, Linux, and error-reporting
issues.

  • Keep keyboard focus on a Settings button when its label changes, such as a toggle switching
    between On and Off. A restarted setup shows its own progress rather than the paused attempt's, and
    Settings shows its status from the first frame.
  • Show the newest audio level at the center of the pill's grille for every bar count, including
    while the capsule resizes.
  • Report a failed capture thread as "Microphone stopped unexpectedly. Try recording again." instead
    of leaving the session finishing. Zero unencoded recording audio whenever its buffer is released.
  • On Linux, let specific desktop-access errors, such as a keyboard mapping change, reach Settings
    instead of a generic stop message. Wayland manual paste asks you to release the shortcut and
    paste, and --help lists --toggle and --cancel. Error dialogs include their cause.
  • Own background threads, native input monitors, and engine processes through shared helpers that
    join, wake, or kill and reap them before replacement. Split Settings, the pill, and the tray into
    focused modules; share pill geometry, shortcuts, and service state from the platform crate; and
    use typed AppKit calls on macOS. Windows and macOS insertion share one pre-commit check.

Windows x64 and Apple silicon macOS zip packages, plus Linux x86_64 AppImage and bundled tar
packages, are attached with SHA-256 checksums. Existing settings, engines, models, dictation
gestures, recognition policy, and the five-minute recording limit are preserved.

Linux remains experimental. This release changes native adapters on all three platforms; CI builds,
lints, and tests them natively and runs owned-window rendering checks, while live microphone,
keyboard-hook, clipboard/editor, and compositor acceptance remain separate. Forced native
termination retains synchronous cleanup. Current contracts and measurement limits are in
architecture and performance.

Speakeasy 0.3.1

Choose a tag to compare

@github-actions github-actions released this 01 Oct 15:19
23c3b71

Release notes

Speakeasy 0.3.1 tightens Rust ownership and failure handling, adds a reproducible tooling gate, and
makes the current implementation easier to navigate and verify.

  • Enforce documented public interfaces, checked integer arithmetic, explicit fallible-result
    handling, and locally justified synchronization. Keep unsafe FFI in the native adapters and
    preserve the bounded, allocation-free capture callback.
  • Validate WAV lengths and sample rates, reject invalid process-group identities, and check setup
    staging failures with actionable errors. Regression tests cover these boundaries and bounded
    relaunch commands; generated gesture sequences exercise timer invariants.
  • Keep Windows minimize handling owned by Settings, preserve file-dialog errors, and evaluate setup
    directory defaults only when needed.
  • Run locked Rust, Python, shell, and documentation checks through mise. Verify native Rust on
    Windows and macOS, and build packages with owned-window rendering checks on all three platforms.
  • Consolidate living documentation around ownership, current behavior, and relevant checks. Remove
    duplicate tray notification handling and update the dependency-patch instructions.

Windows x64 and Apple silicon macOS zip packages, plus Linux x86_64 AppImage and bundled tar
packages, are attached with SHA-256 checksums. Existing settings, engines, models, dictation
gestures, recognition policy, and the five-minute recording limit are preserved.

Linux remains experimental. Live microphone, keyboard-hook, clipboard/editor, and compositor
acceptance are separate from fake and owned-window checks. Forced native termination retains
synchronous cleanup; a stuck native driver can delay completion. Current contracts and measurement
limits are in architecture and performance.

Speakeasy 0.3.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 04:42

Speakeasy 0.3.0 makes session and native-resource ownership explicit, improves cancellation and shutdown handling, and updates the Rust toolchain and workspace dependencies.

  • Own each session's capture, inference, and insertion stages. A cancelled or replaced job cannot publish a stale result; microphone and speech replacements wait for native cleanup while input remains responsive.
  • Coordinate application-owned Quit across dictation, shortcut monitoring, setup, insertion, and requested settings saves. Repeated Quit is harmless, and delayed completions cannot re-enable dictation. Cleanup keeps the UI responsive without abandoning native owners.
  • Preserve microphone-worker startup causes and report actionable local-engine exit diagnostics without forwarding potentially private stderr. The speech meter retains its documented −60 to −6 dBFS range.
  • Prevent automatic Linux insertion into Speakeasy's own windows. Run portal focus checks on an owned worker, keep cancellation responsive during blocked X calls, and share insertion eligibility between native adapters.
  • Wake Windows shortcut monitoring with an owned event rather than relying on a successfully posted shutdown message. Portable Windows/macOS keyboard tests exercise modifier sides, missed releases, hands-free Space, and passive Escape.
  • Test the production session owner with a paused clock for desktop readiness, gestures, and recording deadlines. Strengthen cancellation, warmup failure, retirement, save, and lifecycle regression coverage, with Windows/macOS/Linux verification in CI.
  • Pin nightly-2026-10-01, update workspace dependencies to stable releases, and document module contracts and the relevant checks for future changes. A repository-wide finishing pass removes redundant cleanup and configuration and preserves independent checksum-test expectations.

Windows x64 and Apple silicon macOS zip packages, plus Linux x86_64 AppImage and bundled tar packages, are attached with SHA-256 checksums. Existing settings, engines, models, dictation gestures, recognition policy, and the five-minute recording limit are preserved.

Linux remains experimental. Live microphone, keyboard-hook, clipboard/editor, and compositor acceptance are separate from fake and owned-window checks. Forced native termination retains synchronous cleanup; a stuck native driver can delay completion. Engine diagnostics discard stderr, and upstream dependency constraints retain some older transitive versions. Validation details and measurement limits are in docs/refactor-validation.md.

Speakeasy 0.2.2

Choose a tag to compare

@github-actions github-actions released this 30 Sep 20:26

Speakeasy 0.2.2 reduces rendering memory, improves recording and setup responsiveness, and fixes native window resource and scheduling defects.

  • Settings uses embedded SVG diamonds, and all three renderers allocate path surfaces only when needed. Paired native Windows measurements showed 38.5–39.0% lower GPU committed memory and 23.4–25.3% lower process private commitment with Settings open. After 12 resizes, reductions were 73.2% and 69.2%. Windows RSS stayed essentially flat. Linux software rendering used about 31 MiB less resident memory at 2× scale. The pill retains its animated paths and antialiasing quality; complete Windows Settings images differed only at 66 diamond-edge pixels.

  • macOS window hide/show retains one owned frame source and subscribes it to a shared display link. This backports an upstream fix for native resources retained on each stop/start cycle. Portable lifecycle tests cover repeated visibility, display changes, failed starts, and window destruction; native macOS memory savings remain unmeasured.

  • Finishing or cancelling capture wakes its owned consumer immediately. Paired component measurements reduced the polling delay from about 2.9 ms to 0.02 ms median on native Windows, and from 2.6 ms to 0.06 ms on Linux. The p95 delay fell from about 5.1 ms to 0.03 ms on Windows and from 4.8 ms to 0.09 ms on Linux. This shared Windows/macOS/Linux change preserves callback work and recording cadence; native microphone teardown and full dictation latency are separate measurements.

  • Verifying resumed or completed model downloads uses bounded 1 MiB asynchronous reads. The same 714 MB public model verified in 340 ms versus 881 ms on Linux and 371 ms versus 505 ms in a native Windows component probe. Cancellation still yields between chunks and preserves partial data for retry. These are verification timings, not download or inference gains.

  • Linux GUI startup no longer reaches GPUI's unimplemented X11 native-handle methods.

  • The idle Linux pill starts unmapped, with its nonactivating and click-through configuration applied before it appears.

  • Buffered XCB startup events start frame scheduling without needing another desktop event. This fixes a reproduced black Settings window in an isolated desktop.

  • Buffered events are also drained after foreground tasks, so windows opened after launch render reliably. A disconnected X server terminates the event loop even with every window hidden, avoiding a reproduced one-core CPU loop.

  • Windows refresh timing preserves fractional clock precision. Valid reduced refresh-rate ratios no longer divide by zero in the fallback calculation; invalid intervals use the existing default. The arithmetic failure is reproduced; the native trigger has not been observed.

  • Release checks exercise SVG/path rendering, resize, and repeated show/hide on owned windows. Linux checks inspect actual rendered pixels and packaged GUIs on a private virtual display. Windows checks also exercise hidden paint acknowledgement; native demo and tray acceptance passed locally.

The native fixes use a documented patch to the pinned GPUI source. Recognition, recording limits, display pacing, and local-only behavior are preserved. Native microphone, clipboard, GNOME/KDE compositor, Windows display-power, and macOS interaction acceptance remain separate. Detailed evidence and measurement limits are in docs/performance.md.

Speakeasy 0.2.1

Choose a tag to compare

@github-actions github-actions released this 30 Sep 05:56

Speakeasy 0.2.1 keeps settings, cancellation, and Linux clipboard preparation responsive while tightening ownership of native resources.

  • Validate and save settings off the UI thread. Saves preserve later edits, serialize durable writes, and respect Pause; Quit finishes requested saves. Resume also validates configured files in the background.
  • Initialize Windows and macOS shortcut monitoring asynchronously. Readiness is reported after native initialization, and Pause owns listener cleanup without waiting on the UI thread.
  • Own microphone teardown and serialize rapid retries. Cancellation checks run between native startup stages; Pause waits for microphone and speech cleanup together.
  • Cancel model startup cooperatively and explicitly reap owned speech processes before replacement. Healthy GPU inference cancellation still keeps the warm model; input remains responsive during recovery.
  • Keep Wayland clipboard preparation and bounded selection transfers responsive to desktop events and cancellation. X11 fallback clipboard calls use an owned worker; automatic X11 paste uses one synchronization after its key releases, and manual copy skips modifier waiting.
  • Verify a completed partial download locally instead of downloading it again. Resumed hashing yields between bounded chunks so setup cancellation can proceed. Cancel keeps cleanup off the UI thread; Quit waits for owned setup children to exit.
  • Release large unused PCM allocations after heavy quiet-edge trimming. A synthetic five-minute 48 kHz recording retaining eleven seconds of padded speech released about 26.5 MiB of buffer capacity; ordinary recordings keep their allocation. This is a capacity measurement, not a whole-app RSS or GPU-memory claim.

Windows x64 and Apple silicon macOS zip packages, plus Linux x86_64 AppImage and bundled tar packages, are attached with SHA-256 checksums. Existing settings, engines, and models are preserved.

Linux remains experimental. This release uses safe fixtures and subprocess checks; live microphone, keyboard hooks, compositor permissions, clipboard/editor insertion, displayed frame pacing, and macOS runtime acceptance were not exercised. Native acceptance requires explicit opt-in. Recognition policy, the five-minute recording limit, warm model residency, and the display-aware animation budget are unchanged.

Speakeasy 0.2.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 03:50

Speakeasy 0.2.0 adds experimental Linux desktop builds and reduces audio processing work while retaining official GPUI and responsive interaction.

  • Add X11 and Wayland desktop adapters, Linux Settings and tray integration, XDG storage, and pinned local speech setup. The default Linux dictate chord is Ctrl+Super+Space; Ctrl+Super+Escape cancels. Desktops can use explicit command bindings and manual paste.
  • Own asynchronous insertion with session-bound cancellation. The controller remains responsive during modifier and clipboard preparation; stale work cannot paste for a later recording.
  • Batch audio ring publication and consumption without an added timer. Paired synthetic packet-processing measurements took 40–47% less time; this is a small component cost, not a whole-app CPU reduction.
  • Share immutable shortcut labels between snapshots, filter Linux tray events, and serialize native insertion work. Windows/Linux profiling scripts report CPU, memory, threads, and handles or descriptors.
  • Fix portal double-tap event ordering, check X11 clipboard payload ownership before paste, and show platform-specific gesture guidance in Settings.

Windows x64 and Apple silicon macOS zip packages, plus Linux x86_64 AppImage and bundled tar packages, are attached with SHA-256 checksums. Existing settings and models are preserved. Engines and models are downloaded separately on first launch.

Linux support is experimental. The app targets glibc 2.35+, X11/Xwayland, Vulkan, and desktop audio. Install the launcher as described in the README so portals can identify the app. Native GNOME/KDE/X11 acceptance is pending, and automatic paste depends on compositor permissions and modifier feedback. The pinned upstream Linux CPU engine can fail on some CPUs without AVX-512; a synthetic inference readiness check catches this and directs users to a compatible engine. Broad CPU fallback compatibility is not established.

X11 automatic paste requires the first keyboard layout with an unshifted V key; other layouts receive manual-paste guidance.

The existing display-aware 200-FPS ceiling and recording policy are unchanged. Native whole-app footprint and displayed frame pacing remain separate measurements.

Speakeasy 0.1.1

Choose a tag to compare

@github-actions github-actions released this 30 Sep 00:09

Speakeasy 0.1.1 improves capture reliability and reduces rendering work while keeping the official GPUI framework.

  • Handle a Windows audio discontinuity before the first microphone packet, fixing the observed startup failure with Valorant open. Interruptions during speech still fail visibly, and microphone errors retain actionable driver details.
  • Stop the Windows pill from spinning on hidden paint messages, removing the measured idle CPU spike.
  • Budget pill animation at 200 FPS, coalesce meter updates, and keep session changes immediate. Paired Windows preview measurements showed 36% fewer process cycles per second; displayed frame pacing and power consumption were not measured.
  • Serialize automatic setup across retries and app instances, and terminate and wait for its child process before canceled setup releases the directory.
  • Keep held Space repeats out of the editor after enabling hands-free, and preserve the saved motion preference when Settings validation fails.

Windows x64 and Apple silicon macOS packages are attached. On first launch, Speakeasy downloads the speech engine and model suited to the machine. Existing settings and models are preserved.

Speakeasy 0.1.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 17:21

Windows and macOS builds of 458f7f3. On first launch, Speakeasy downloads the speech engine and model that suit the machine.