-
-
Notifications
You must be signed in to change notification settings - Fork 377
TTS
Over the years I've built up a small ecosystem of C/C++ libraries and examples that bring Text-to-Speech (TTS) to microcontrollers like the ESP32, and to AudioTools as the glue that gets the resulting audio out to I2S, Bluetooth, or anywhere else a Print/Stream can go.
There's no single "one size fits all" solution for speech synthesis on embedded devices — the right approach depends heavily on your available flash, RAM, CPU budget, network connectivity, and how natural you need the voice to sound. Broadly there are two families of solutions:
- On-device synthesis — a library running on the microcontroller itself turns text into audio, no network required.
- External/network TTS — the microcontroller sends text to a TTS service (cloud or local) over Wi-Fi and streams back the resulting audio.
This page gives an overview of both, plus the ready-to-run examples that ship in examples/examples-tts of this repo.
Roughly ordered from "lightest and most robotic" to "heaviest and most natural":
| Library | Approach | Best for |
|---|---|---|
| arduino-SAM | Formant synthesis ("Software Automatic Mouth") | Smallest footprint, classic retro voice |
| arduino-flite | Formant/diphone synthesis (Flite) | Clear speech, low resources |
| arduino-espeak-ng | Formant synthesis (eSpeak NG) | Multilingual, 100+ languages/voices |
| arduino-simple-tts | Pre-recorded audio snippets | Best voice quality, fixed vocabulary |
| TinyTTSTools | Modular G2P + vocoder pipeline (formant / phoneme / diphone / PSOLA, optional neural G2P fallback) | Header-only, pick-your-own quality/memory tradeoff |
| TinyTTS | Neural TTS (proof of concept) | Experimenting with neural synthesis |
arduino-SAM — Software Automatic Mouth
A port of the classic 1982 "Software Automatic Mouth" speech synthesizer — the same engine behind the iconic C64 speech demos. It's tiny, extremely fast, and produces that unmistakable retro robotic voice. If you want speech on the smallest possible microcontroller with minimal flash and no dependency on pre-recorded audio, this is about as lean as it gets.
arduino-flite — Flite (Festival-Lite)
A port of Flite, a small, fast, portable speech synthesis engine derived from the Festival and FestVox projects at Carnegie Mellon. Flite produces clearer, more intelligible speech than SAM while still being compact enough to run on an Arduino-class device.
arduino-espeak-ng — eSpeak NG
A port of eSpeak NG, a mature, formant-based open-source speech synthesizer supporting more than a hundred languages and accents. This is the most feature-complete "classic" TTS engine of the bunch, and my most popular TTS repo.
arduino-simple-tts — Pre-recorded Audio TTS
Sometimes you don't need a synthesis engine at all. arduino-simple-tts stitches together pre-recorded audio snippets (words, phrases, numbers) to construct spoken output. This trades flexibility for quality — since the audio is pre-recorded, the result sounds far more natural than any real-time formant synthesizer, at the cost of being limited to a known vocabulary.
TinyTTSTools — Text-to-Speech Tools for Microcontrollers
A flexible, header-only TTS toolkit built around a three-stage pipeline: Grapheme-to-Phoneme (G2P) → Phoneme-to-Audio (Vocoder) → Audio Output, with several interchangeable implementations for each stage so you can pick your own quality/memory tradeoff instead of being locked into one engine.
-
G2P:
G2PDictionaryModel(dictionary lookup),G2PRuleBasedModel(rule-based English),G2PNeuralModel(a dependency-free GRU neural fallback for out-of-dictionary words, ported from the sibling TinyTTS project, no TensorFlow Lite needed) — and hybrids combining several of these. -
Vocoders:
FormantVocoder(pure rule-based synthesis, no audio data required),PhonemeVocoder/DiphoneVocoder(playback of pre-recorded, compressed phoneme/diphone samples from PROGMEM or SD/LittleFS), andPSOLAVocoder(TD-PSOLA re-synthesis of those same samples for genuine pitch-shifting and time-stretching). -
Output: any
Printtarget — works directly with AudioTools'I2SStream,AnalogAudioStream,PWMAudioStream, or a raw callback.
Unlike TinyTTS, it doesn't run a full neural TTS model on-device (microcontrollers are too slow for that in real time); the "neural" piece here is only an optional G2P fallback model. What you get instead is a practical, working set of formant/sample-based/PSOLA synthesizers usable today, with a design that's easy to extend with custom G2P models or vocoders.
TinyTTS — Neural TTS Proof of Concept
A header-only C++ port of a ~1.6M-parameter, VITS-style, end-to-end neural TTS model, with no external inference-runtime dependency (no TFLite Micro) — a proof of concept evaluating whether ML-based TTS can run on current microcontrollers at all. The conclusion: memory is not the bottleneck (~4.2MB flash, ~743KB PSRAM), but current microcontrollers are still ~19-32x too slow for real-time generation; the same unmodified code runs comfortably at or faster than real-time on a desktop or a Raspberry Pi 4.
If your device has Wi-Fi, you don't need to run the synthesizer on the microcontroller at all: send the text to a TTS service and stream the returned audio straight into an AudioTools output chain. This gets you the most natural-sounding voices (real neural TTS), at the cost of needing connectivity and, for cloud services, an API key.
The pattern is always the same:
URLStream --copy--> [optional Decoder] --> I2S / AudioBoardStream / Bluetooth
URLStream (in AudioTools/Communication/AudioHttp.h) issues the HTTP(S) request — GET or POST, with custom headers and a body — and the response body is read as a Stream, so it can be piped directly into a StreamCopy, decoded on the fly if needed (WAV, MP3 via Helix, etc.), and written out to I2S, an audio board, or A2DP.
Azure Cognitive Services Speech (streams-azure_tts-i2s) — POSTs SSML to the Azure REST endpoint and streams back raw PCM:
URLStream AzureURLStream("ssid", "pwd");
I2SStream i2s;
StreamCopy copier(i2s, AzureURLStream);
String ssml = "<speak version='1.0' xml:lang='en-US'><voice xml:lang='en-US' "
"xml:gender='Female' name='en-US-JennyNeural'>Hello there</voice></speak>";
AzureURLStream.addRequestHeader("Ocp-Apim-Subscription-Key", speechKey.c_str());
AzureURLStream.addRequestHeader("X-Microsoft-OutputFormat", "raw-8khz-16bit-mono-pcm");
AzureURLStream.begin(url_str.c_str(), "audio/wav", POST, "application/ssml+xml", ssml.c_str());Since Azure returns raw PCM directly, no decoder is needed — the I2S output just needs to be configured to match the requested sample rate/format.
Google Translate TTS (streams-google-audiokit) — a fun, keyless option that (ab)uses Google Translate's free translate_tts endpoint. It returns MP3, so it's decoded with arduino-libhelix on the way to I2S:
URLStream url("ssid", "password");
AudioBoardStream i2s(AudioKitEs8388V1);
EncodedAudioStream dec(&i2s, new MP3DecoderHelix());
StreamCopy copier(dec, url);
const char* url_str = tts("this is an english text"); // builds the translate_tts URL
url.begin(url_str, "audio/mp3");Being a free, undocumented endpoint, it's great for demos but not something to rely on for a production device.
rhasspy (streams-url_wav-i2s) — if you'd rather not depend on a cloud provider at all, rhasspy can run locally (e.g. on a Raspberry Pi) and exposes a simple /api/text-to-speech REST endpoint that returns WAV, decoded here with the built-in WAVDecoder:
URLStream url("ssid", "password");
I2SStream i2s;
WAVDecoder decoder;
EncodedAudioStream out(&i2s, &decoder);
StreamCopy copier(out, url);
url.begin("http://192.168.1.37:12101/api/text-to-speech?play=false",
"audio/wav", POST, "text/plain", "Hallo, my name is Alice");This same pattern works for any self-hosted TTS server that speaks HTTP — Piper, Coqui TTS, OpenAI-compatible TTS endpoints, etc. — just point URLStream at the right URL, set the appropriate content type/headers, and pick the matching decoder (or none, if the service streams raw PCM).
More background on the network-TTS approach: Text to Speech in Arduino - Conclusions.
All examples live under examples/examples-tts in this repo. Check each folder's README for the extra library you need to install.
| Example | Library / Service | Output |
|---|---|---|
streams-sam-i2s, streams-sam-audiokit
|
arduino-SAM | I2S / AudioKit |
streams-flite-i2s, streams-flite-audiokit
|
arduino-flite | I2S / AudioKit |
streams-espeak-i2s, streams-espeak-audiokit
|
arduino-espeak-ng | I2S / AudioKit |
streams-simple_tts-i2s, streams-simple_tts-a2dp
|
arduino-simple-tts | I2S / Bluetooth (A2DP) |
streams-talkie-a2dp, streams-talkie-audiokit
|
TalkiePCM | Bluetooth (A2DP) / AudioKit |
streams-azure_tts-i2s |
Azure Cognitive Services (network) | I2S |
streams-google-audiokit |
Google Translate TTS (network) | AudioKit |
streams-url_wav-i2s |
rhasspy (network, self-hosted) | I2S |
Additional TTS examples using the on-board webserver are available in the webserver examples folder.
- Need a handful of fixed announcements ("Door open", "Battery low") with the best possible voice quality and no network? Use arduino-simple-tts.
- Need to speak arbitrary, dynamic text entirely offline? Use arduino-espeak-ng for multilingual support, or arduino-flite/arduino-SAM for a lighter, English-only footprint.
- Want to mix and match dictionary/rule-based/neural G2P with formant, sample-based, or PSOLA resynthesis, and tune the tradeoff yourself? Use TinyTTSTools.
- Have Wi-Fi and want the most natural voice with the least on-device effort? Stream from a cloud TTS service (Azure, or similar), or a self-hosted engine like rhasspy/Piper if you'd rather not depend on the cloud.
- Curious whether a full neural TTS model can run on-device at all? TinyTTS is the proof-of-concept — works great on a desktop or Raspberry Pi 4, but current microcontrollers are still too slow for real-time generation.
All of these projects are open source and available on GitHub — issues, stars, and contributions are always welcome.