Skip to content

v4.4.0 — Parakeet Ultra Transcription

Latest

Choose a tag to compare

@github-actions github-actions released this 01 Oct 20:31

What's New

More Accurate Transcription with Parakeet Ultra

  • Transcription runs on Parakeet Ultra, a post-trained Parakeet v3, through FluidAudio 0.17.5
  • Same 25 languages, word timings and speed as before, with fewer errors: 2.13% vs 2.27% word error rate on LibriSpeech test-clean, 3.81% vs 4.12% on test-other
  • Lower error rate in all 24 FLEURS languages FluidAudio measures (mean 11.67% vs 14.81%)
  • The model (about 600 MB) downloads once, with the first transcription after updating

Share Before the Transcript Is Ready

  • If you share a recording while it's still transcribing, the share page gets the transcript, title, summary and chapters as soon as transcription finishes
  • Regenerating a transcript on a shared recording updates its page as well
  • Sharing while a transcription runs, or transcribing in the player while a share uploads, no longer throws away the other's result

Transcriptions Resume After a Quit or Crash

  • A transcription cut off by quitting, a crash or a force quit starts again the next time Voom opens
  • Before, the recording stayed on "Transcribing…" indefinitely and its Transcribe button stayed hidden
  • Meeting recordings resume through the speaker-labeling pipeline

Current AI Models

  • OpenAI: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna
  • Anthropic: Claude Opus 5.5, Fable 5.1, Sonnet 5.5, Haiku 4.5
  • Google: Gemini 3.1 Pro, Gemini 3.8 Flash, Gemini 3.5 Flash Lite
  • xAI: Grok 4.7, Grok 4.3
  • If your saved model was dropped from the list, Voom switches you to its successor at the same tier

Under the Hood

  • FluidAudio 0.17.1 → 0.17.5, which includes the upstream fixes for the Nemotron 3 Neural Engine issue (FluidInference/FluidAudio#951). Diarization stays on CPU + GPU so the Neural Engine is free for transcription running alongside it
  • The share server replaces a recording's transcript and chapters when the app re-sends them, instead of adding a second copy
  • Claude requests allow up to 16,000 output tokens and AI requests wait up to three minutes, since current models think before answering and that thinking counts against the limit

Known Limitations

  • The first transcription after updating waits for the Parakeet Ultra download
  • Self-hosted share servers get the replace-on-resend change when they are next redeployed. Until then, regenerating a transcript on a recording that already had one on its page adds a second copy
  • Viewers who loaded a share page before its transcript arrived may see cached captions and link previews for up to an hour

Full Changelog: v4.3.0...v4.4.0