0.1.0 is the first versioned release of NeMo-Speech.cpp.
Highlights
- Speech recognition with Nemotron 3.5 ASR, Nemotron Speech Streaming, Parakeet TDT, and Parakeet CTC
- Speaker diarization with Streaming Sortformer v2, standalone or integrated with ASR
- Speech synthesis with MagpieTTS and NanoCodec
- Text and speech translation with Riva Translate 4B Instruct v2
- Subtitles, VAD, endpointing, punctuation, capitalization, and structured output
- Unified nemo-speech CLI with managed model discovery, downloads, and caching
- Local HTTP API with OpenAI-compatible transcription and speech subsets, realtime WebSocket transcription, and playground
- Native C API, shared libraries, and exported CMake package for application integration
- CPU, CUDA, Vulkan, and Metal backend support
Prebuilt archives cover Linux and Windows on x86-64 and ARM64, plus Apple Silicon macOS, using the platform-appropriate CPU, CUDA, Vulkan, or Metal backend.
Note: this is an early release and API interfaces may evolve.