Skip to content

Releases: bglglzd/uxo

UXO v0.1.2 Preview

UXO v0.1.2 Preview Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 04 Oct 07:52

UXO v0.1.2 Preview

UXO v0.1.2 makes dictated text read the way people type it: brand names and technical terms are spelled correctly, dates and phone numbers become digits, and English mixed into another language keeps its own spelling. It also adds in-app updates and an opt-in mode for dictating into administrator terminals on every launch.

Warning

This is an unsigned Windows x64 Preview. Windows SmartScreen may show an “Unknown publisher” warning. Download only from this repository's official release and verify the package against the published SHA-256 checksums before installing.

Download

Asset Recommended use
UXO_0.1.2_x64-setup.exe Standard interactive installer; recommended for most people
UXO_0.1.2_x64_en-US.msi MSI package for managed or administrative installation
SHA256SUMS.txt SHA-256 checksums for the release assets

The speech model is downloaded separately during first-run setup and is not bundled in either installer. Installing over v0.1.1 keeps your settings, history and downloaded models.

What's new

  • Brand and tech dictionary. A built-in offline dictionary fixes the spelling of well-known brands and technical terms, including Russian transliterations: “гитхаб” → GitHub, “чат джипити” → ChatGPT, “вс код” → VS Code. Ordinary words that happen to sound like a brand are left alone. It can be turned off in Settings → Advanced → Transcription.
  • Replacement rules. A custom word can now be a rule such as спейс икс, спейсикс → SpaceX: whatever you say on the left is always written as the right-hand side.
  • Better custom words in Cyrillic. Fuzzy correction of custom words now works for Cyrillic as well as Latin terms and keeps Russian case endings intact.
  • Numbers as digits. Spelled-out dates, years, phone numbers, prices and other quantities are written with digits when the context implies it — “25 марта”, “2026 года”, “+7 916 123-45-67”, “125 рублей” — while small standalone numbers stay as words. Russian and English; on by default.
  • Second language. In the model settings, choose a language you often mix into your main one, such as English in Russian speech. Whisper models then keep those terms in their own spelling instead of transliterating or translating them.
  • In-app updates. UXO can check this repository's releases, show what changed in every newer version before anything is downloaded, verify the installer against SHA256SUMS.txt, install it and restart. Background checks can be turned off in Settings → About.
  • Always run as administrator. An opt-in setting in Settings → Advanced → Output starts UXO with administrator rights on every launch, so dictation reaches terminals and shells that run as administrator without a manual restart each session.
  • Interface polish. Menus, tooltips and onboarding cards have a proper background again; dark mode, contrast, screen-reader labels and keyboard support were fixed across the settings.

Updating from v0.1.1

v0.1.1 has no in-app updater, so install v0.1.2 manually once using the installer above. From v0.1.2 on, new releases are offered inside UXO.

In-app updates apply to the standard installer on Windows. Portable copies and other platforms show the release notes and a link to the release page instead.

Elevated Windows applications

Windows prevents a normally running application from sending simulated input into a higher-integrity administrator window. To dictate into an administrator terminal, either use Restart as administrator for the current session or enable Always run as administrator. With the latter, Windows asks for confirmation at every start, including autostart; declining keeps UXO running normally. Both options are unavailable in portable mode.

Because this Preview elevates the full unsigned application, use elevated mode only on a trusted personal installation and only while it is needed.

Privacy boundary

  • Core recording and speech recognition are local. A network connection is needed to download models.
  • The dictionary, number formatting and second-language handling run entirely on your device.
  • Update checks request only the public list of this repository's GitHub releases; nothing is downloaded or installed without confirmation, and background checks can be disabled.
  • Optional LLM post-processing is opt-in and separate from the local transcription guarantee. Invoking it sends transcript text to the configured provider under that provider's terms.

Verification

Download SHA256SUMS.txt beside the installers, then calculate their digests in PowerShell:

Get-FileHash .\UXO_0.1.2_x64-setup.exe -Algorithm SHA256
Get-FileHash .\UXO_0.1.2_x64_en-US.msi -Algorithm SHA256
Get-Content .\SHA256SUMS.txt

The calculated values must exactly match the entries in SHA256SUMS.txt.

Known limitations

  • Windows x64 is the only qualified binary target for v0.1.2.
  • The installers are not code-signed, so Windows may display a SmartScreen warning. In-app updates verify checksums from the same release; they are not a code-signed update channel.
  • Windows speech inference is CPU-only in this release.
  • Number formatting and the built-in dictionary cover Russian and English; other languages pass through unchanged.
  • The second-language hint applies to Whisper models only.
  • The Preview elevates the full UXO process; a future signed release should replace this compatibility mode with a minimal signed input helper.
  • Language quality and latency depend on the chosen model, hardware, accent, microphone, and acoustic environment.

Licensing

UXO source code is MIT-licensed. Bundled third-party components and separately downloaded models retain their own licenses and terms. See LICENSE and THIRD_PARTY_NOTICES.md.

UXO v0.1.1 Preview

UXO v0.1.1 Preview Pre-release
Pre-release

Choose a tag to compare

@bglglzd bglglzd released this 28 Aug 22:02

UXO v0.1.1 Preview

UXO v0.1.1 makes dictation more dependable across Windows applications and fixes selector menus throughout the interface.

Warning

This is an unsigned Windows x64 Preview. Windows SmartScreen may show an “Unknown publisher” warning. Download only from this repository's official release and verify the package against the published SHA-256 checksums before installing.

Download

Asset Recommended use
UXO_0.1.1_x64-setup.exe Standard interactive installer; recommended for most people
UXO_0.1.1_x64_en-US.msi MSI package for managed or administrative installation
SHA256SUMS.txt SHA-256 checksums for the release assets

The speech model is downloaded separately during first-run setup and is not bundled in either installer.

What's fixed

  • Dropdowns stay visible. Selector menus are now rendered above the glass-card layout instead of being hidden beneath the next settings row. The shared fix covers language, model, and other dropdown selectors across the application.
  • Dictation into elevated terminals. Windows Terminal, PowerShell, SSH, and tmux sessions opened with administrator rights can receive dictated text after an explicit, session-only Restart as administrator action in UXO's Advanced settings.
  • No silent loss on blocked paste. When Windows prevents input injection into a higher-integrity target, UXO preserves the complete transcription on the clipboard and reports the real paste failure instead of claiming success.
  • Localized recovery guidance. The elevated-target status, recovery action, and paste-error guidance are available in all 24 shipped interface locales.

Elevated Windows applications

Windows prevents a normally running application from sending simulated input into a higher-integrity administrator window. This is an operating-system security boundary, not a limitation of PowerShell, SSH, or tmux.

When UXO detects that boundary, it keeps the transcription on the clipboard so it can still be pasted manually. To dictate directly into an administrator terminal:

  1. Open Settings → Advanced in UXO.
  2. Choose Restart as administrator.
  3. Approve the Windows User Account Control prompt.

Elevation applies only to that UXO session. A later normal launch starts without administrator rights again. Because this Preview restarts the full unsigned application, use elevated mode only on a trusted personal installation and only while it is needed. It is unavailable in portable mode, and UXO disables external-script paste while elevated. Windows secure-desktop prompts, protected credential fields, and some protected or anti-cheat applications may still reject automated input.

Languages

UXO ships 24 interface languages: English, Simplified Chinese, Traditional Chinese, Spanish, French, German, Japanese, Korean, Vietnamese, Polish, Italian, Russian, Ukrainian, Portuguese, Czech, Turkish, Arabic, Hebrew, Swedish, Bulgarian, Dutch, Nepali, Hindi, and Danish.

Speech-language support is determined by the selected local model and is separate from the interface language.

Privacy boundary

  • Core recording and speech recognition are local. A network connection is needed to download models.
  • History and retained recordings may be stored locally when their corresponding settings are enabled.
  • Optional LLM post-processing is opt-in and separate from the local transcription guarantee. Invoking it sends transcript text to the configured provider under that provider's terms.
  • Remote-provider endpoints and API keys are stored in UXO's local settings in plaintext in this Preview, not in a dedicated OS credential vault. Use scoped, revocable credentials and avoid configuring this feature on an untrusted shared Windows account.

Verification

Download SHA256SUMS.txt beside the installers, then calculate their digests in PowerShell:

Get-FileHash .\UXO_0.1.1_x64-setup.exe -Algorithm SHA256
Get-FileHash .\UXO_0.1.1_x64_en-US.msi -Algorithm SHA256
Get-Content .\SHA256SUMS.txt

The calculated values must exactly match the entries in SHA256SUMS.txt.

Known limitations

  • Windows x64 is the only qualified binary target for v0.1.1.
  • The installers are not code-signed, so Windows may display a SmartScreen warning.
  • Windows speech inference is CPU-only in this release.
  • Model files require a separate download and retain their own licenses and usage terms.
  • There is no signed production update channel yet.
  • Direct input into elevated applications requires the explicit session-only restart described above.
  • The Preview elevates the full UXO process; a future signed release should replace this compatibility mode with a minimal signed input helper.
  • Language quality and latency depend on the chosen model, hardware, accent, microphone, and acoustic environment.
  • Optional remote post-processing is not offline and has the local API-key storage caveat described above.

Licensing

UXO source code is MIT-licensed. Bundled third-party components and separately downloaded models retain their own licenses and terms. See LICENSE and THIRD_PARTY_NOTICES.md.

UXO v0.1.0 Preview

UXO v0.1.0 Preview Pre-release
Pre-release

Choose a tag to compare

@bglglzd bglglzd released this 28 Aug 18:18

UXO v0.1.0 Preview

UXO's first public Preview turns speech into text locally on Windows: press a global shortcut, speak, and the result is pasted into the focused app without relying on the active keyboard layout.

Warning

This is an unsigned Windows x64 Preview. Windows SmartScreen may show an “Unknown publisher” warning. Download only from this repository's official release and verify the package against the published SHA-256 checksums before installing.

Download

Asset Recommended use
UXO_0.1.0_x64-setup.exe Standard interactive installer; recommended for most people
UXO_0.1.0_x64_en-US.msi MSI package for managed or administrative installation
SHA256SUMS.txt SHA-256 checksums for the release assets

The speech model is downloaded separately during first-run setup and is not bundled in either installer.

Highlights

  • Local, CPU-only core ASR on Windows. Ordinary microphone audio and speech recognition stay on the computer, while the GPU remains available to games and other demanding apps.
  • Fast tray workflow. UXO can launch at login, start hidden after onboarding, retain the selected model, and prewarm the model and voice-activity detector in the background.
  • Global dictation anywhere. Choose hold-to-talk or toggle recording, speak into the selected microphone, and paste into the currently focused text field.
  • Keyboard-layout-independent text. Unicode clipboard insertion prevents Russian or other non-Latin output from being corrupted by the active keyboard layout.
  • Cohesive UXO interface. A restrained navy-and-cyan visual system, unified line icons, accessible focus states, and subtle glass surfaces keep the tray app clear and compact.
  • Streaming Nemotron default. The recommended Q8 build of NVIDIA Nemotron 3.5 ASR Streaming 0.6B is approximately 751 MB and is downloaded separately.
  • Model choice. Whisper-family, Parakeet, and other compatible local models remain available for different language, accuracy, memory, and latency requirements.
  • Multilingual UI. The application interface ships in 24 languages.

Languages

Interface translations and speech-model coverage are separate:

  • UXO interface: 24 languages — English, Simplified Chinese, Traditional Chinese, Spanish, French, German, Japanese, Korean, Vietnamese, Polish, Italian, Russian, Ukrainian, Portuguese, Czech, Turkish, Arabic, Hebrew, Swedish, Bulgarian, Dutch, Nepali, Hindi, and Danish.
  • Recommended Nemotron model: 32 ready-to-use locales across 28 base languages. Coverage includes English, Spanish, French, Italian, Portuguese, Dutch, German, Turkish, Russian, Arabic, Hindi, Japanese, Korean, Vietnamese, Ukrainian, Polish, Swedish, Czech, Norwegian Bokmål, Danish, Bulgarian, Finnish, Croatian, Slovak, Mandarin Chinese, Hungarian, Romanian, and Estonian.

English, Spanish, French, and Portuguese each include multiple regional locales. Accuracy varies by language, accent, microphone, and environment; alternative models have different language coverage.

Privacy boundary

  • Core recording and speech recognition are local. A network connection is needed to download models.
  • History and retained recordings may be stored locally when their corresponding settings are enabled.
  • Optional LLM post-processing is opt-in and separate from the local transcription guarantee. Invoking it sends transcript text to the configured provider under that provider's terms.
  • Remote-provider endpoints and API keys are stored in UXO's local settings in plaintext in this Preview, not in a dedicated OS credential vault. Use scoped, revocable credentials and avoid configuring this feature on an untrusted shared Windows account.

Verification

Download SHA256SUMS.txt beside the installer, then calculate its digest in PowerShell:

Get-FileHash .\UXO_0.1.0_x64-setup.exe -Algorithm SHA256
Get-FileHash .\UXO_0.1.0_x64_en-US.msi -Algorithm SHA256
Get-Content .\SHA256SUMS.txt

The calculated values must exactly match the entries in SHA256SUMS.txt.

The release candidate is validated with frontend build, lint, translation, and formatting checks; 208 Rust library tests; two Playwright application-shell smoke tests; and Windows package launch and dependency audits.

Known limitations

  • Windows x64 is the only qualified binary target for v0.1.0.
  • The installers are not code-signed, so Windows may display a SmartScreen warning.
  • Windows speech inference is CPU-only in this release.
  • Model files require a separate download and retain their own licenses and usage terms.
  • There is no signed production update channel yet.
  • Language quality and latency depend on the chosen model, hardware, accent, microphone, and acoustic environment.
  • Optional remote post-processing is not offline and has the local API-key storage caveat described above.

Licensing

UXO source code is MIT-licensed. Bundled third-party components and separately downloaded models retain their own licenses and terms. See LICENSE and THIRD_PARTY_NOTICES.md.