Skip to content

v3.0

Latest

Choose a tag to compare

@github-actions github-actions released this 02 Aug 08:15

๐Ÿ—ฃ๏ธ TTS Application v3.0

A powerful, cross-platform text-to-speech app featuring Google Cloud TTS and ElevenLabs integration, real-time playback, advanced voice controls, and flexible audio output options.

Google Cloud TTS UI
ElevenLabs TTS UI

โœจ Whatโ€™s New in v3.0

๐Ÿง  Dual TTS Provider Support

  • ๐Ÿ”„ Switch seamlessly between:
    • Google Cloud TTS โ€“ Enterprise-grade performance with SSML and audio profiles
    • ElevenLabs โ€“ Realistic voices with emotional expression and custom sliders

๐ŸŽ›๏ธ Enhanced Voice Customization

  • Google Cloud
    • ๐ŸŽš๏ธ Control speed, pitch, and audio profiles
    • ๐Ÿ“ Full SSML support (<break>, <prosody>, etc.)
  • ElevenLabs
    • ๐Ÿ”ง Adjust stability, similarity boost, style, speaker boost, and speed

๐Ÿ–ฅ๏ธ All-New Cross-Platform UI

  • ๐ŸŒ Available on Windows, macOS, and Linux
  • ๐ŸŽจ Dynamic UI that updates based on the selected provider
  • ๐Ÿ“Š Real-time quota usage tracking (for both providers)
  • ๐ŸŽง Integrated audio playback, pause/resume, and stop controls
  • ๐Ÿ”ฝ One-click download in selected format (.mp3, .wav, .ogg, etc.)

๐Ÿ“ฅ Quick Start

  1. Download the latest release from the Releases page
  2. Add credentials:
    • ๐Ÿ” Google: credentials/google.json
    • ๐Ÿ” ElevenLabs: credentials/elevenlabs.json
  3. Run the app:
    • Windows:
      .\TTS_App_Windows.exe
    • macOS:
      chmod +x TTS_App_macOS
      ./TTS_App_macOS
    • Linux:
      chmod +x TTS_App_Linux
      ./TTS_App_Linux

๐Ÿงฉ System Requirements

  • OS:
    • Windows 10/11 (64-bit)
    • macOS 11.0+ (Big Sur and above)
    • Ubuntu 20.04+ (glibc โ‰ฅ 2.31)
  • Python: 3.10 or later

๐Ÿ” Credentials Setup (For detailed setup instructions, please refer to UserGuide.md.)

๐Ÿ”‘ Google Cloud TTS

  1. Enable APIs in the Google Cloud Console:

    • Text-to-Speech API
    • Vertex AI API (for advanced voices)
    • Cloud Storage API (optional for caching)
  2. Create a Service Account with roles:

    • Cloud Text-to-Speech API User
    • Vertex AI Service Agent
    • Storage Object Admin (if needed)
  3. Generate JSON Key and save as:
    credentials/google.json

๐Ÿ”‘ ElevenLabs TTS

  1. Create an account at https://www.elevenlabs.io

  2. Generate API Key from Profile > API Keys

  3. Save as: credentials/elevenlabs.json

    {
      "api_key": "YOUR_API_KEY"
    }
    

๐ŸŽฎ How to Use

1. Select TTS Provider

Use the dropdown to switch between Google Cloud and ElevenLabs.
UI updates dynamically based on provider.


2. Enter Text

Use plain text or SSML (for Google Cloud):

<speak>
  Hello <break time="0.5s"/> world!
</speak>

3. Select Language & Voice

๐ŸŸฆ Google Cloud

  • Choose from:
    • Standard voices
    • Wavenet voices
    • Neural2 voices
  • Controls available:
    • Speed
    • Pitch
    • Audio Profile
  • โœ… Full SSML support

๐ŸŸจ ElevenLabs

  • Choose from:
    • Multilingual v2
    • Turbo v2 models
  • Controls available:
    • Stability
    • Similarity Boost
    • Style
    • Speaker Boost
    • Speed

4. Playback and Download

Use the control buttons for audio interaction:

  • โ–ถ๏ธ Play
  • โฏ๏ธ Pause / Resume
  • โน๏ธ Stop
  • โฌ‡๏ธ Download the generated audio in your selected format

๐Ÿ“ Supported Audio Formats

Provider Formats Supported Default
Google Cloud MP3, LINEAR16 (WAV), OGG_OPUS MP3
ElevenLabs MP3, PCM, ULAW MP3

โš ๏ธ Known Limitations

Google Cloud

  • โŒ Some voices (e.g., Chirp) do not support SSML or pitch control
  • ๐Ÿ“‰ API usage statistics may be unavailable

ElevenLabs

  • โณ Some models take longer to generate audio
  • ๐Ÿ” Retry logic may be triggered if response is delayed
  • ๐Ÿ“Š Usage stats may lag slightly behind real-time

๐Ÿ› ๏ธ Developer Setup

# Install Python dependencies
pip install -r requirements.txt

# Build standalone executable
pyinstaller --onefile --add-data "credentials;credentials" src/main.py