Releases: nyosegawa/speech.cpp
Release list
v0.6.0
speech-worker, the other tools and the shared library libspeech with its C API, speech.h, for macOS arm64 (Metal), Windows x64 (Vulkan) and Linux x64 (Vulkan, and the CPU alone; glibc 2.34 or later). The README lists the model families the worker runs and where their converted weights are on Hugging Face.
What's Changed
- Make speech recognition a task of the C API, the worker, the server and a speech-asr tool by @nyosegawa in #28
- Decode parakeet-tdt_ctc-0.6b-ja with its TDT decoder and drop the CTC head by @nyosegawa in #29
- Raise the version to 0.6.0 by @nyosegawa in #30
Full Changelog: v0.5.0...v0.6.0
v0.5.0
speech-worker, the other tools and the shared library libspeech with its C API, speech.h, for macOS arm64 (Metal), Windows x64 (Vulkan) and Linux x64 (Vulkan, and the CPU alone; glibc 2.34 or later). The README lists the model families the worker runs and where their converted weights are on Hugging Face.
What's Changed
- Let a request set Irodori-TTS's length and speed, refuse what a model cannot do, and answer every line the worker cannot read by @nyosegawa in #21
- Replace qwen3-tts and irodori-tts with one command-line tool on the C API, speech-tts by @nyosegawa in #22
- Serve a model over HTTP with OpenAI's speech API, speech-server by @nyosegawa in #23
- Port FastConformer's frontend, encoder and CTC head for parakeet-tdt_ctc-0.6b-ja, checked stage by stage against NeMo by @nyosegawa in #24
- Share the graph helper between Irodori-TTS and FastConformer from src/common by @nyosegawa in #25
- Build and release Linux x64 in a Vulkan and a CPU build on glibc 2.34 by @nyosegawa in #26
- Raise the version to 0.5.0 by @nyosegawa in #27
Full Changelog: v0.4.0...v0.5.0
v0.4.0
speech-worker, the other tools and the shared library libspeech with its C API, speech.h, for macOS arm64 (Metal) and Windows x64 (Vulkan). The README lists the model families the worker runs and where their converted weights are on Hugging Face.
What's Changed
- Serve every program through one C API that the worker itself uses by @nyosegawa in #19
- Speak plain JSON Lines from the worker, and name the release in one VERSION file by @nyosegawa in #20
Full Changelog: v0.3.1...v0.4.0
v0.3.1
speech-worker and the other tools for macOS arm64 (Metal) and Windows x64 (Vulkan). The README lists the model families the worker runs and where their converted weights are on Hugging Face.
What's Changed
- Build ggml once in the Windows CI and restore it from the cache by @nyosegawa in #17
- Open model paths as UTF-8 on Windows and stop a cancelled Irodori-TTS sentence between its sampler's steps by @nyosegawa in #18
Full Changelog: v0.3.0...v0.3.1
v0.3.0
speech-worker and the other tools for macOS arm64 (Metal) and Windows x64 (Vulkan). The README lists the model families the worker runs and where their converted weights are on Hugging Face.
What's Changed
- Record the Irodori-TTS decisions and its measurements on Vulkan by @nyosegawa in #12
- Turn off Metal's tensor API before any tool starts a device by @nyosegawa in #14
- Name every family's languages with BCP 47 tags by @nyosegawa in #15
- Point to the Irodori-TTS weights on Hugging Face by @nyosegawa in #16
Full Changelog: v0.2.0...v0.3.0
v0.2.0
speech-worker and the other tools for macOS arm64 (Metal) and Windows x64 (Vulkan). The worker runs Qwen3-TTS, whose models are on https://huggingface.co/sakasegawa/qwen3-tts-ggml, and Irodori-TTS v4.1-Small-MF and v4.1-Small, whose models reference/irodori-tts converts (see the README).
v0.1.1
Binaries for macOS arm64 (Metal) and Windows x64 (Vulkan). The models are on https://huggingface.co/sakasegawa/qwen3-tts-ggml.
v0.1.0
Binaries for macOS arm64 (Metal) and Windows x64 (Vulkan). The models are on https://huggingface.co/sakasegawa/qwen3-tts-ggml.