v4.3.0 — Nemotron 3 Speaker Diarization
What's New
Speaker Diarization with NVIDIA Nemotron 3
- Meetings are diarized with NVIDIA's Nemotron 3 Diarization, running on your Mac through FluidAudio 0.17.1
- Tells up to eight remote speakers apart end-to-end, instead of by clustering voice embeddings
- Detects speakers who talk over each other
- Falls back to the previous pyannote pipeline automatically if Nemotron can't run
- The model (193 MB) downloads with the first meeting transcription
Word-Level Speaker Labels
- Speakers are assigned per word from the transcription's word timings, not per sentence
- Lines split exactly where the speaker changes, including mid-sentence
- Words the mic's voice detection misses are credited to you when no system audio is playing
Speaker Echo No Longer Labeled as You
- Recording on laptop speakers, the mic re-records remote voices, and those lines were labeled "You"
- Voom now estimates how much system audio leaks into the mic and ignores mic activity the echo explains
- Nothing changes with headphones
Catching Up
- The in-app What's New also covers 4.0 through 4.2, which shipped without one: window recording and camera selection, the new share page with timestamped comments, editing tools, crash-safe recordings, and one-click self-hosting
Under the Hood
- FluidAudio 0.15.2 → 0.17.1; the pyannote fallback is retuned for its corrected clustering thresholds
- Nemotron runs on CPU + GPU (about 500× real time) while an Apple Neural Engine issue in FluidAudio is open (FluidInference/FluidAudio#951)
- Meeting audio is streamed in chunks, so long meetings are never decoded into memory whole
- 17 new unit tests for segmentation, speaker attribution, and echo detection
Known Limitations
- Transcription produces one stream of words, so words spoken over each other still land in a single line
- Beyond eight speakers, extra voices can be merged into others
Full Changelog: v4.2.0...v4.3.0