First prebuilt release, so you no longer need a compiler or the Vulkan SDK to try it. Unpack and run.
Downloads
| File | For |
|---|---|
breeze-tts-2-v0.1.0-windows-x64-vulkan.zip |
Windows 10 and 11, 64 bit |
breeze-tts-2-v0.1.0-linux-x64-vulkan.tar.gz |
Linux x64 with glibc 2.35 or newer (Ubuntu 22.04, Debian 12, Fedora 36 and up) |
Both are Vulkan builds, so they run on NVIDIA, AMD and Intel GPUs, integrated ones included, and fall back to the CPU if no GPU is found. All you need is a current GPU driver. The CPU has to support AVX2, which anything from roughly 2013 onwards does.
The Linux build comes out of GitHub Actions and shows up here a few minutes after the release goes live. It needs the Vulkan loader, which any desktop with GPU drivers already has (libvulkan1 on Debian and Ubuntu).
Getting started
- Unpack the archive.
- Download a model from HoppouAI/Breeze-TTS-2.cpp on Hugging Face. Q8_0 is the recommended one, Q4_K if you are short on VRAM.
- Start the server with the web UI:
breeze-server breeze-tts-2-q8_0.gguf --webui - Open http://localhost:8080/
breeze-cli writes straight to a WAV file if you'd rather not run a server, run it with no arguments to see the options.
Windows might show a SmartScreen warning the first time because the exe isn't signed. Click More info, then Run anyway.
What's in the archive
- breeze-server streams audio over HTTP and WebSocket as it generates, with the web UI built in
- breeze-cli generates speech to a WAV file
- breeze-convert respeaks a recording in another voice
- breeze-quantize makes your own quantized models
- libbreeze and
include/for the C API, if you want to call it from another language - docs/ is the full documentation
New since the last commits you might have built
- OpenAI compatible TTS endpoint (#7).
/v1/audio/speechalso takes OpenAI style JSON now, so SillyTavern and the OpenAI SDKs can use it as a TTS provider. It streams MP3, WAV or PCM, and can send server sent events. SillyTavern setup is indocs/server.md. - CORS (#5). Start the server with
--corsto let pages on other origins call it. - Integrated GPUs (#6). When there's no discrete GPU it picks up the integrated one instead of dropping straight to the CPU.
- Guidance and the depth decoder got faster, and the WebSocket flush and end messages actually speak the whole message now instead of leaving the last bit stuck in the buffer.
Please report bugs in issues.