๐ฃ๏ธ TTS Application v3.0
A powerful, cross-platform text-to-speech app featuring Google Cloud TTS and ElevenLabs integration, real-time playback, advanced voice controls, and flexible audio output options.
โจ Whatโs New in v3.0
๐ง Dual TTS Provider Support
- ๐ Switch seamlessly between:
- Google Cloud TTS โ Enterprise-grade performance with SSML and audio profiles
- ElevenLabs โ Realistic voices with emotional expression and custom sliders
๐๏ธ Enhanced Voice Customization
- Google Cloud
- ๐๏ธ Control speed, pitch, and audio profiles
- ๐ Full SSML support (
<break>,<prosody>, etc.)
- ElevenLabs
- ๐ง Adjust stability, similarity boost, style, speaker boost, and speed
๐ฅ๏ธ All-New Cross-Platform UI
- ๐ Available on Windows, macOS, and Linux
- ๐จ Dynamic UI that updates based on the selected provider
- ๐ Real-time quota usage tracking (for both providers)
- ๐ง Integrated audio playback, pause/resume, and stop controls
- ๐ฝ One-click download in selected format (
.mp3,.wav,.ogg, etc.)
๐ฅ Quick Start
- Download the latest release from the Releases page
- Add credentials:
- ๐ Google:
credentials/google.json - ๐ ElevenLabs:
credentials/elevenlabs.json
- ๐ Google:
- Run the app:
- Windows:
.\TTS_App_Windows.exe - macOS:
chmod +x TTS_App_macOS ./TTS_App_macOS
- Linux:
chmod +x TTS_App_Linux ./TTS_App_Linux
- Windows:
๐งฉ System Requirements
- OS:
- Windows 10/11 (64-bit)
- macOS 11.0+ (Big Sur and above)
- Ubuntu 20.04+ (glibc โฅ 2.31)
- Python: 3.10 or later
๐ Credentials Setup (For detailed setup instructions, please refer to UserGuide.md.)
๐ Google Cloud TTS
-
Enable APIs in the Google Cloud Console:
- Text-to-Speech API
- Vertex AI API (for advanced voices)
- Cloud Storage API (optional for caching)
-
Create a Service Account with roles:
Cloud Text-to-Speech API UserVertex AI Service AgentStorage Object Admin(if needed)
-
Generate JSON Key and save as:
credentials/google.json
๐ ElevenLabs TTS
-
Create an account at https://www.elevenlabs.io
-
Generate API Key from Profile > API Keys
-
Save as:
credentials/elevenlabs.json{ "api_key": "YOUR_API_KEY" }
๐ฎ How to Use
1. Select TTS Provider
Use the dropdown to switch between Google Cloud and ElevenLabs.
UI updates dynamically based on provider.
2. Enter Text
Use plain text or SSML (for Google Cloud):
<speak>
Hello <break time="0.5s"/> world!
</speak>3. Select Language & Voice
๐ฆ Google Cloud
- Choose from:
- Standard voices
- Wavenet voices
- Neural2 voices
- Controls available:
- Speed
- Pitch
- Audio Profile
- โ Full SSML support
๐จ ElevenLabs
- Choose from:
- Multilingual v2
- Turbo v2 models
- Controls available:
- Stability
- Similarity Boost
- Style
- Speaker Boost
- Speed
4. Playback and Download
Use the control buttons for audio interaction:
โถ๏ธ Play- โฏ๏ธ Pause / Resume
- โน๏ธ Stop
- โฌ๏ธ Download the generated audio in your selected format
๐ Supported Audio Formats
| Provider | Formats Supported | Default |
|---|---|---|
| Google Cloud | MP3, LINEAR16 (WAV), OGG_OPUS | MP3 |
| ElevenLabs | MP3, PCM, ULAW | MP3 |
โ ๏ธ Known Limitations
Google Cloud
- โ Some voices (e.g., Chirp) do not support SSML or pitch control
- ๐ API usage statistics may be unavailable
ElevenLabs
- โณ Some models take longer to generate audio
- ๐ Retry logic may be triggered if response is delayed
- ๐ Usage stats may lag slightly behind real-time
๐ ๏ธ Developer Setup
# Install Python dependencies
pip install -r requirements.txt
# Build standalone executable
pyinstaller --onefile --add-data "credentials;credentials" src/main.py

