Releases: ddv1982/echo
Releases · ddv1982/echo
Release list
v0.11.0
- Whisper runs now report separate WAV encoding, child-process, parsing, runtime, backend, decoding, and attempt detail while keeping the existing
inferMsboundary compatible with old history and CLI consumers. - One typed execution plan owns the selected Whisper runtime, model, VAD, protocol, and any explicit decoding overrides. Normal runs preserve the runtime's own tuning defaults, managed CPU keeps its existing precedence, and system or manually imported assets remain supported.
- The file CLI adds unsaved Whisper tuning overrides for reproducible experiments. The benchmark records outer wall time, artifact hashes, host identity, seeds, warmups, randomized candidate order, resolved tuning, and every VAD retry.
- New managed Whisper installs include the matching
whisper-serverfrom the already verified upstream archive. Existing one-shot installations remain valid without repair. A separate loopback-only probe measures model load, first request, warm requests, memory, and cleanup without enabling resident dictation before it passes the quality and latency gates. - Advanced diagnostics show the actual cold path, runtime source, backend, split timing, decoding values, and VAD retries. General Settings gains no performance knobs.
- Echo retries without VAD only when Whisper reports a VAD model or context failure, failed VAD computation, or an exact unsupported VAD flag. Decoder and model failures now preserve their original error instead of paying for an unrelated second inference.
v0.10.0
- Recommended setup now installs Large v3 Turbo Q5_0 on machines with at least 8 GiB RAM. The 547 MiB quantized model replaces Small as the normal high-quality multilingual choice, while Base Q5_1 remains the low-memory fallback and existing Small or manual models stay supported.
- Settings uses one Speech model row: Whisper exposes installed choices with honest quality guidance, while Parakeet shows its fixed TDT 0.6B v3 model and automatic 25-language capability instead of making the control disappear.
- Parakeet now parses the pinned sherpa-onnx JSON protocol, passes only transcript text to cleanup and insertion, reports its model path, and uses the required NeMo transducer model type.
- Linux clipboard fallback leaves the dictated text available after paste instead of racing the target by immediately restoring old clipboard contents. Wayland clipboard tools are preferred in Wayland sessions and direct typing remains clipboard-free.
- A manifest-driven benchmark runs installed speech candidates through the shipping CLI, fails on missing or broken candidates, and reports per-language WER, real-time factor, and silence hallucinations in JSON Lines and Markdown.
v0.9.1
- Managed Whisper setup now installs the pinned 1.9.2 runtime even when its shared-library symlinks appear before their targets in the archive.
- Extraction validates the complete selected symlink graph, including cycles, missing targets, escapes, and flattened destination mismatches, before any staged link can reach activation.
- CI downloads the exact pinned Whisper archive and drives it through digest verification, extraction, payload verification, the real Linux runtime probe, immutable activation, receipt checks, and post-install Verify.
v0.9.0
- Echo now discovers Linux microphones through native PipeWire or PulseAudio before falling back to ALSA, so Bluetooth, USB, and built-in sources use the same recognizable names exposed by the desktop sound server.
- The normal microphone picker shows the Linux system default and primary input sources. Playback sinks, ALSA plugins, aliases, resamplers, and raw endpoint IDs remain available under Advanced audio endpoints.
- Speech setup is now one compact readiness card. Installed component paths and maintenance actions, alternative models, and inactive engine plans start collapsed without removing repair, verification, removal, system-runtime, or manual-model support.
- Settings now adapts cleanly at the 760-pixel minimum and across the navigation breakpoint. Pinned Chromium tests check eight widths in both themes for horizontal overflow and closed disclosures.
- Debian and RPM packages declare the PipeWire and PulseAudio runtime libraries, and release CI inspects the generated dependency metadata before publication.
v0.8.0
- Recording length is now a visible General setting shared by timed, button, tray, CLI, and shortcut capture, with 30-second, 1-minute, 2-minute, 5-minute, and 10-minute choices. Ten minutes is the default and ceiling.
- Active sessions snapshot their limit, Home shows that value while recording, preview behavior matches the backend, and existing
record_secondsconfig plusECHO_RECORD_SECONDSoverrides remain supported. - Ten-minute capture avoids the previous native-sample clone and full mono intermediate buffer while preserving exact conversion output on tested mono, stereo, and multichannel inputs.
- Shortcut verification always cleans up its test recording, and token-scoped stop requests cannot cancel a replacement session. Fixture capture now obeys the same limit and cancellation contract as live capture.
v0.7.0
- Microphones now use CPAL stable device IDs, keep equal labels distinct, show available metadata, preserve disconnected choices with an explicit fallback, and test the exact selected input.
- Linux x86_64 users can install complete Whisper or Parakeet setups inside Echo. Managed components use resumable downloads, SHA-256, bounded archive extraction, immutable activation records, Verify, Repair, and managed-only removal.
- System runtimes and existing cache files remain external, visible, and untouched. Healthy managed components take precedence while corrupt managed components fall back to those external inputs.
- Recommended setup chooses a multilingual Whisper model from detected memory and installs its runtime and VAD. Downloads expose cumulative disk needs, progress, cancellation, resume, retry, repair, verification, and removal without activating partial or corrupt files.
v0.6.0
echo-desktop transcribe FILE.wavnow writes clean text, raw text, or schema-versioned JSON to stdout or an exact output path without starting recorder or desktop side effects.- One prepared transcription request now resolves the engine, Whisper model, language, cleanup mode, and bounded dictionary recognition hints for both microphone and file runs.
echo-desktop languagesreports model-aware Whisper languages and Parakeet's 25 automatic-only languages in text or JSON.- Engine, model, and language precedence is source-aware, failed inference processes cannot leak partial output, and microphone cleanup retains its dictionary-only fallback.
v0.5.0
- Echo now uses one fixed Super+Alt+Space toggle across the desktop portal, X11, GNOME setup, and manual compositor setup. Push-to-talk, raw-input fallback, shortcut customization, and the
rec --holdcommand have been removed. - Shortcut setup is reported through one typed status, remains available when unrelated Settings probes fail, supports explicit retry, and verifies activations against the effective binding.
v0.4.2
Publication-path hotfix for the fully verified v0.4.1 artifacts. The application changes and package contents remain the same.
- Release-candidate checks now download the staged Debian, RPM, and binary artifacts on every pull request and
mainbuild, then verify the exact directory layout consumed by the publisher. - The GitHub Release publisher follows the artifact service's preserved
deb/andrpm/subdirectories, preventing a valid tagged build from failing at its final attachment step. - The failed
v0.4.1tag run remains visible as an audit record;v0.4.2is the first release published entirely by the hardened workflow.
v0.4.0
Linux shortcuts are now configurable, source-aware, and resilient across modern Wayland, X11, and older GNOME sessions.
- Toggle and push-to-talk shortcuts support canonical multi-key chords, environment overrides, persisted settings, capture/reset controls, and effective-trigger reporting.
- Echo uses the GlobalShortcuts portal when available and native X11 grabs otherwise, with explicit conflict and registration errors instead of silent fallback claims.
- GNOME releases without the portal get an explicit, ownership-checked setup and repair action for the Echo custom toggle shortcut; startup and status polling never change desktop settings.
- Push-to-talk prefers native desktop shortcuts and falls back to a chord-aware evdev supervisor that handles multiple keyboards, hotplug, reconnect, cancellation, permission denial, and listener failure without granting privileges.
- Advanced Settings identifies toggle and push-to-talk sources independently, and Test shortcut accepts only a successful action from the configured shortcut command path.
- The recording HUD is smaller and its premultiplied ARGB edges render cleanly without a pale fringe.
- CI now includes a release build alongside frontend, test, and lint gates.