Skip to content

LocalText2Voice 2.2.0

Latest

Choose a tag to compare

@estebanstifli estebanstifli released this 04 Oct 04:09
· 6 commits to main since this release

LocalText2Voice 2.2.0

This release adds IndexTTS-2.5 voice cloning and per-passage emotions, plus
short-text fixes and clearer generation estimates.

Highlights

  • An optional new local engine. Install IndexTTS-2.5 from Settings > TTS
    Engines, reuse a reference recording from the voice library, and keep the
    worker in memory through EngineHost. No reference transcript is required.
  • Direct emotions in your script. The editor's Emotion menu offers happy,
    sad, angry, disgust, fear, surprise, calm and melancholic presets, with an
    optional intensity. Off restores the reference recording's expression.
  • Custom emotional descriptions. QwenEmotion converts descriptions into
    vectors in a preparation pass, then exits before narration loads. Saved
    vectors are reused when regenerating segments; repeated descriptions are
    calculated once. Presets do not need to load QwenEmotion.
  • Short-text and progress improvements. Chatterbox handles the affected
    short-input alignment case; Qwen/Chatterbox respect the selected chunk size.
    Synthesis estimates account for text length after the first three blocks.
{{lang en}}
{{emotion happy 75%}}
We finally made it! There is so much to celebrate today.
{{emotion custom "Quiet sadness, with a little hope" 60%}}
I still miss those days, but tomorrow could be different.
{{emotion off}}
This passage uses the expression from the selected reference recording.

Select or import a voice first. These emotion commands target IndexTTS;
other engines ignore them with a warning. Presets are application settings,
not calibrated guarantees, and free-form emotion interpretation can be wrong.

Installation And Upgrade Notes

  • Optional IndexTTS models and dependencies are not bundled in the Windows
    installer. Reserve about 30 GB for its installation, downloads and caches.
  • CUDA/BF16 is the initial selection and requires a compatible NVIDIA GPU.
    CPU inference requires choosing FP32 explicitly; performance varies by PC.
  • IndexTTS uses the bilibili Model Use License Agreement, with conditions
    on commercial use, redistribution and AI training. Read the displayed license
    before installing. LocalText2Voice's own source remains MIT licensed.
  • Existing IndexTTS global emotion settings from development builds reset to
    reference mode. Put emotional direction in the script; explicit API settings
    remain available.
  • Let current jobs finish, then restart the app and idle EngineHost after
    updating so all clients load the new code.

In Development

H3 workflow support is not ready for public use. The private-endpoint prototype
is labeled In development and retained for testing only. General-purpose
local/remote H3 JSON workflows will be documented when ready.

Documentation