Skip to content

Releases: alex37529/voxspica

VoxSpica 0.1.6

Choose a tag to compare

@alex37529 alex37529 released this 03 Oct 15:08

VoxSpica 0.1.6

Free speech recognition that runs offline: record from the microphoneor transcribe audio files, with no cloud and nothing sent to the
internet.

Fixes and improvements

  • A model's timings are stored, kept and shown. Every cue is saved with the entry instead of being thrown away with the buffer, so the timings are still there after a restart. The history pane shows one [mm:ss] line per cue, and a click on a line moves the player to that moment. A switch turns the stamps off for reading, and the column in the table stays plain in both modes so a scan of five hundred rows is not a column of numbers.
  • A recorded take is playable from the history. The audio of a microphone recording is kept next to its transcript, and the history has a sound column so you can tell at a glance which entries can be played at all - a table of texts cannot say that.
  • The player has a real transport. Play and pause are separate from stop, the bar shows how far in and how long the recording is, it can be dragged, and the end of a clip is reported rather than guessed. A build without an engine degrades to a bar that counts, and says so, instead of looking broken.
  • The recognition language you choose in the window now sticks. It was written nowhere: the window loaded the model for the language you picked, said so in the status line, and forgot the choice - so the next start went back to the previous language and loaded that model instead. The size next to it had the same failure one step earlier: written to the settings, never saved.
  • A click on a timestamped line used to land at the beginning. Loading a file is asynchronous, and asking the engine for a position before it is ready is not a jump but nothing at all - the request was lost. The position is now kept and applied as soon as the file is loaded, whether the sound is playing, paused or stopped, and whichever button you press next starts from there.
  • The sound no longer outlives the window. Closing the history window with the Close button, with Escape or with the title bar sends no close event, so the closeEvent that was supposed to stop the playback never ran and the recording played on with nothing on screen to stop it.
  • The upgrade no longer deletes your models. Two ways it did, both on the first start after installing: a directory copy interrupted partway through was taken for a finished one, which cost every downloaded model with no message; and in a portable build - the program unpacked in a folder you can write to - the directory the installer unpacks into is the same one models are downloaded to, so every model except the bundled one was judged a leftover of a previous installation and removed. A copy now marks itself finished, and only what the installer itself wrote is ever removed.
  • Long files are streamed, and the window stays alive while they are recognised. The file is decoded in pieces instead of being read into memory whole, the recognition runs off the GUI thread, and there is a progress bar and a working stop button. A temporary copy can no longer outlive the call or collide with the next one.
  • Recognising a file reuses the model the window already has. It used to load a second copy: a user who waited for the large model to open, watched the same few minutes again on the next file, with both copies in memory.
  • A disabled control says why. The history's sound column says whether there is no recording, the file is gone, or the microphone is open; the bar says the same, and explains that a build without QtMultimedia can play but cannot scrub.
  • The search finds Russian text. casefold on both sides rather than SQLite's own collation, which folds ASCII only and silently found nothing.
  • Going back to an older version works again. This release adds three settings - , and - and the previous version refused to open a settings file containing any key it did not know. Putting 0.1.6 on top of 0.1.5 and then going back therefore left a program that would not start, with a traceback and no way out but editing the file by hand. Unknown keys are now kept and written back untouched - this version does not apply them, and does not silently delete them either - and each one is named on stderr, so a typo in a key is still reported rather than quietly ignored.

Other

  • The window has one spacing scale and one type hierarchy, so the buttons, fields and headings line up instead of each being its own idea.
  • The About dialog's buttons sit on one row.
  • The portable build publishes a .sha256 next to the archive, so a package manager can verify it. Scoop reads that file, and with the table it had been updating silently to nothing.

Install

  1. Download VoxSpica-0.1.6-win64.zip and unpack it into any folder — the archive contains a single file, VoxSpica.exe.
  2. Run it. On the first launch, pick the interface language and download a recognition model (models are not included in the download).
  3. The folder can be moved anywhere: models and transcripts are written next to the executable.

Note: there is nothing else to unpack. The program extracts itsown libraries into a Windows temporary directory on every launch, so
startup takes about a second longer. If Windows shows a SmartScreenwarning, that is expected for an unsigned build — choose "More info"→ "Run anyway".

Requirements: Windows 10/11 x64, ~500 MB of free disk space(including room for a model). Python is not needed.

The SHA256 sum is in SHA256SUMS.txt in this release.

Per-language installers, each with a recognition model already bundled, are on voxspica.4crytobot.xyz: en, ru, de, fr, es, it, zh.

This repository holds the releases only. The source code is notpublic and is not mirrored here.

VoxSpica 0.1.5

Choose a tag to compare

@alex37529 alex37529 released this 30 Sep 14:20

VoxSpica 0.1.5

Free speech recognition that runs offline: record from the microphone
or transcribe audio files, with no cloud and nothing sent to the
internet.

Install

  1. Download VoxSpica-0.1.5-win64.zip and unpack it into any folder —
    the archive contains a single file, VoxSpica.exe.
  2. Run it. On the first launch, pick the interface language and
    download a recognition model (models are not included in the
    download).
  3. The folder can be moved anywhere: models and transcripts are
    written next to the executable.

Note: there is nothing else to unpack. The program extracts its
own libraries into a Windows temporary directory on every launch, so
startup takes about a second longer. If Windows shows a SmartScreen
warning, that is expected for an unsigned build — choose "More info"
→ "Run anyway".

Requirements: Windows 10/11 x64, ~500 MB of free disk space
(including room for a model). Python is not needed.

The SHA256 sum is in SHA256SUMS.txt in this release.


This repository holds the releases only. The source code is not
public and is not mirrored here.

VoxSpica 0.1.4

Choose a tag to compare

@alex37529 alex37529 released this 29 Sep 08:02

VoxSpica 0.1.4

Free speech recognition that runs offline: record from the microphone
or transcribe audio files, with no cloud and nothing sent to the
internet.

Install

  1. Download VoxSpica-0.1.4-win64.zip and unpack it into any folder —
    the archive contains a single file, VoxSpica.exe.
  2. Run it. On the first launch, pick the interface language and
    download a recognition model (models are not included in the
    download).
  3. The folder can be moved anywhere: models and transcripts are
    written next to the executable.

Note: there is nothing else to unpack. The program extracts its
own libraries into a Windows temporary directory on every launch, so
startup takes about a second longer. If Windows shows a SmartScreen
warning, that is expected for an unsigned build — choose "More info"
в†’ "Run anyway".

Requirements: Windows 10/11 x64, ~500 MB of free disk space
(including room for a model). Python is not needed.

The SHA256 sum is in SHA256SUMS.txt in this release.


This repository holds the releases only. The source code is not
public and is not mirrored here.

VoxVault 0.1.3

Choose a tag to compare

@alex37529 alex37529 released this 29 Sep 07:04

VoxVault 0.1.3

Free speech recognition that runs offline: record from the microphone
or transcribe audio files, with no cloud and nothing sent to the
internet.

This is a historical build. The program was called VoxVault until
0.1.4, which renamed it to VoxSpica. The file names below are the ones
this release was published under.

Install

  1. Download VoxVault-0.1.3-win64.zip and unpack it into any folder —
    the archive contains a single file, VoxVault.exe.
  2. Run it. On the first launch, pick the interface language and
    download a recognition model (models are not included in the
    download).
  3. The folder can be moved anywhere: models and transcripts are
    written next to the executable.

Note: there is nothing else to unpack. The program extracts its
own libraries into a Windows temporary directory on every launch, so
startup takes about a second longer. If Windows shows a SmartScreen
warning, that is expected for an unsigned build — choose "More info"
→ "Run anyway".

Requirements: Windows 10/11 x64, ~500 MB of free disk space
(including room for a model). Python is not needed.

The SHA256 sum is in SHA256SUMS.txt in this release.


This repository holds the releases only. The source code is not
public and is not mirrored here.