Skip to content

Releases: aritropaul/voom

v4.4.0 — Parakeet Ultra Transcription

Choose a tag to compare

@github-actions github-actions released this 01 Oct 20:31

What's New

More Accurate Transcription with Parakeet Ultra

  • Transcription runs on Parakeet Ultra, a post-trained Parakeet v3, through FluidAudio 0.17.5
  • Same 25 languages, word timings and speed as before, with fewer errors: 2.13% vs 2.27% word error rate on LibriSpeech test-clean, 3.81% vs 4.12% on test-other
  • Lower error rate in all 24 FLEURS languages FluidAudio measures (mean 11.67% vs 14.81%)
  • The model (about 600 MB) downloads once, with the first transcription after updating

Share Before the Transcript Is Ready

  • If you share a recording while it's still transcribing, the share page gets the transcript, title, summary and chapters as soon as transcription finishes
  • Regenerating a transcript on a shared recording updates its page as well
  • Sharing while a transcription runs, or transcribing in the player while a share uploads, no longer throws away the other's result

Transcriptions Resume After a Quit or Crash

  • A transcription cut off by quitting, a crash or a force quit starts again the next time Voom opens
  • Before, the recording stayed on "Transcribing…" indefinitely and its Transcribe button stayed hidden
  • Meeting recordings resume through the speaker-labeling pipeline

Current AI Models

  • OpenAI: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna
  • Anthropic: Claude Opus 5.5, Fable 5.1, Sonnet 5.5, Haiku 4.5
  • Google: Gemini 3.1 Pro, Gemini 3.8 Flash, Gemini 3.5 Flash Lite
  • xAI: Grok 4.7, Grok 4.3
  • If your saved model was dropped from the list, Voom switches you to its successor at the same tier

Under the Hood

  • FluidAudio 0.17.1 → 0.17.5, which includes the upstream fixes for the Nemotron 3 Neural Engine issue (FluidInference/FluidAudio#951). Diarization stays on CPU + GPU so the Neural Engine is free for transcription running alongside it
  • The share server replaces a recording's transcript and chapters when the app re-sends them, instead of adding a second copy
  • Claude requests allow up to 16,000 output tokens and AI requests wait up to three minutes, since current models think before answering and that thinking counts against the limit

Known Limitations

  • The first transcription after updating waits for the Parakeet Ultra download
  • Self-hosted share servers get the replace-on-resend change when they are next redeployed. Until then, regenerating a transcript on a recording that already had one on its page adds a second copy
  • Viewers who loaded a share page before its transcript arrived may see cached captions and link previews for up to an hour

Full Changelog: v4.3.0...v4.4.0

v4.3.0 — Nemotron 3 Speaker Diarization

Choose a tag to compare

@github-actions github-actions released this 24 Sep 14:57

What's New

Speaker Diarization with NVIDIA Nemotron 3

  • Meetings are diarized with NVIDIA's Nemotron 3 Diarization, running on your Mac through FluidAudio 0.17.1
  • Tells up to eight remote speakers apart end-to-end, instead of by clustering voice embeddings
  • Detects speakers who talk over each other
  • Falls back to the previous pyannote pipeline automatically if Nemotron can't run
  • The model (193 MB) downloads with the first meeting transcription

Word-Level Speaker Labels

  • Speakers are assigned per word from the transcription's word timings, not per sentence
  • Lines split exactly where the speaker changes, including mid-sentence
  • Words the mic's voice detection misses are credited to you when no system audio is playing

Speaker Echo No Longer Labeled as You

  • Recording on laptop speakers, the mic re-records remote voices, and those lines were labeled "You"
  • Voom now estimates how much system audio leaks into the mic and ignores mic activity the echo explains
  • Nothing changes with headphones

Catching Up

  • The in-app What's New also covers 4.0 through 4.2, which shipped without one: window recording and camera selection, the new share page with timestamped comments, editing tools, crash-safe recordings, and one-click self-hosting

Under the Hood

  • FluidAudio 0.15.2 → 0.17.1; the pyannote fallback is retuned for its corrected clustering thresholds
  • Nemotron runs on CPU + GPU (about 500× real time) while an Apple Neural Engine issue in FluidAudio is open (FluidInference/FluidAudio#951)
  • Meeting audio is streamed in chunks, so long meetings are never decoded into memory whole
  • 17 new unit tests for segmentation, speaker attribution, and echo detection

Known Limitations

  • Transcription produces one stream of words, so words spoken over each other still land in a single line
  • Beyond eight speakers, extra voices can be merged into others

Full Changelog: v4.2.0...v4.3.0

v4.2.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 04:14

Full Changelog: v4.1.3...v4.2.0

v4.1.3

Choose a tag to compare

@github-actions github-actions released this 17 Sep 22:18

Full Changelog: v4.1.2...v4.1.3

v4.1.2

Choose a tag to compare

@github-actions github-actions released this 30 Jul 16:12

Full Changelog: v4.1.1...v4.1.2

v4.1.1

Choose a tag to compare

@github-actions github-actions released this 30 Jul 14:22

Full Changelog: v4.1.0...v4.1.1

v4.1.0

Choose a tag to compare

@github-actions github-actions released this 10 Jun 19:04

Full Changelog: v4.0.2...v4.1.0

v4.0.2

Choose a tag to compare

@github-actions github-actions released this 09 Jun 17:09

Full Changelog: v4.0.1...v4.0.2

v4.0.1 — Share Page Polish

Choose a tag to compare

@github-actions github-actions released this 09 Apr 22:32

Fixes

Share Page Legibility

  • Add backdrop blur layer over hero video (matching Linear's technique)
  • mix-blend-mode: plus-lighter on description text
  • Summary clamped to 3 lines

Content Section Redesign

  • Editorial layout: 680px max-width, 1px section dividers
  • Transcript and comment lists with bottom fade masks
  • Pill-shaped reaction buttons with hover lift
  • Comment cards with surface background and hover border
  • Refined typography hierarchy: 22px title, 10px section labels
  • Tabular nums on timestamps, tighter letter-spacing
  • Full-bleed hover states on transcript and chapter rows
  • Smooth micro-interactions throughout

v4.0.0 — Cinematic Share Page & Editing Tools

Choose a tag to compare

@github-actions github-actions released this 09 Apr 22:19

What's New

Cinematic Share Page

  • Full-viewport hero with blurred video background, vignette gradient overlay
  • Vertically centered content with play button that opens modal video player
  • Astro SPA frontend with Inter font, dark theme, responsive design
  • Backdrop blur layer + mix-blend-mode for text legibility over video
  • Transcript, chapters, reactions, and comments visible below hero

Editing Tools

  • AI Clips — intelligent clip extraction from recordings
  • Transcript Editor — edit transcript text inline
  • Privacy Blur — blur sensitive regions in recordings
  • Transcript Export — export transcripts in multiple formats

Recording Enhancements

  • Cursor and input tracking during recording with sidecar export
  • Recording presets for quick capture setup
  • Auto-zoom analysis service

Sharing Fixes

  • Upload raw HEVC video directly (fixes frame truncation on high-res recordings)
  • Apply missing D1 migrations (summary, chapters, speakers, is_meeting)
  • Add comment rate limiting (client_ip tracking)

Infrastructure

  • Share page rebuilt as Astro 5.0 static SPA
  • Custom video player with keyboard shortcuts, captions, chapters, seekbar
  • Password protection, expiration, reactions, comments on shared recordings