3.1.0
Release Notes - v3.1.0
Major New Features
JIT (Just-In-Time) Generation
Introducing real-time audio streaming! You no longer have to wait for a full document to be processed before you start listening.
- Real-time Playback: Audio begins playing as soon as the first segments are ready.
- Smart Buffering: The engine maintains a 2-minute buffer ahead of playback to ensure a smooth, uninterrupted experience.
- Interruptible: Stop at any time; the system will automatically save what was played and export the remaining text to a
_remaining.txtfile so you can resume later.
Integrated Audio FX Pipeline
Powered by Spotify's pedalboard library, you can now apply professional-grade audio effects directly within the GUI.
- Effects included: Reverb (Room/Hall simulation), Compressor (for consistent volume), and EQ (High/Low Shelf filters).
- FX Presets: Save your favorite effect chains and apply them globally or per-segment via script.
Pronunciation Lexicon
Take control over how the AI speaks.
- Custom Overrides: Define a "Find & Replace" dictionary for specialized terms, acronyms, or names that the base model might mispronounce.
- Persistent: Your lexicon is saved to your configuration and applied automatically to all generations.
Advanced Voice Mixing
Unlock a near-infinite variety of voices.
- Custom Voice Tab: Mix any two standard voices with precise ratio and operation control to create unique voices.
- Save & Reuse: Export your custom mixes as
.ptfiles to use them just like standard voices.
Enhancements & Fixes
- Scripted Multi-Speaker & FX: Added support for the advanced syntax
[Speaker_Preset:FX_Preset]: Your text here. - Intelligent Caching: Implemented a SHA-256 hashing system. If you generate the same text with the same settings twice, the GUI will instantly pull from the cache instead of re-generating.
- Windows
run.bat: A new one-click startup script that automatically handles virtual environment creation and dependency updates. - Performance: Optimized the dual-loop architecture for JIT mode to minimize CPU overhead during simultaneous generation and playback.
- Expanded Language Support: Full integration for Spanish, French, Italian, Portuguese, Japanese, and Chinese.
Requirements Update
This version introduces new dependencies:
pedalboard(for Audio FX)winsound(standard on Windows, used for JIT playback)
To update, simply run run.bat or execute:
pip install -r requirements.txt