VideoLingo v3.1.1 ✨
More flexible subtitle workflows, safer long-video dubbing, and official Fish Audio voice cloning.
🚀 What’s new
- Official Fish Audio dubbing is here. The new
fish_audioprovider uses Fish Audio’s official API and defaults to thes2.1-promodel. It can clone the voice from 15–30 seconds of the source video without storing a voice in your Fish Audio account, or use a fixed Fish Audio voice or custom voice ID. - OpenAI-compatible TTS can now use OpenLux. The OpenAI TTS backend supports the OpenLux speech endpoint, with configurable base URL and model.
- Bring your own subtitles. Import an SRT file for translation, with or without a video, and use transcribe-only mode when you only need source subtitles.
- Review the pipeline before it continues. Optional checkpoints let you edit terminology before translation and review the translation spreadsheet before subtitle generation.
- More control over subtitle output. Set a translation style, customize font/size/color, and independently choose bilingual, translated, or source subtitles for the subtitle and dubbed videos.
- MAI-Transcribe-2 is available as an optional ASR backend through Azure Speech or OpenRouter, while local Qwen3-ASR remains the default.
- Long videos are more robust. Demucs processes long audio in chunks, Japanese NLP is batched, subtitle splitting has safer fallbacks, and translation/punctuation edge cases no longer stop the pipeline.
- Dubbing failures are visible and recoverable. Failed TTS requests stop after retries instead of becoming silent lines. Speaking speed is capped, lines that still cannot fit are listed in
output/log/dub_truncated.json, and dubbed videos are produced correctly even when burned subtitles are disabled.
⚠️ Before upgrading
- The old
fish_ttsprovider for 302.ai andazure_ttshave been removed. Choose Fish Audio, Edge TTS, or another available provider; Edge TTS is now the default. - Fish Audio requires
fish_audio.api_keyand an account with credit.msgpackis a new dependency and is installed with the normal setup flow. - The Qwen3-ASR model size is now an advanced
config.yamloption:whisper.qwen_model.
✅ Release validation
- Automated suite: 663 passed, 19 skipped, 12 subtests passed.
- Five real English videos (about 7.5–11.3 minutes each) were processed end to end into Simplified Chinese, Japanese, French, Spanish, and Traditional Chinese using OpenLux GPT-5.6 Luna and Fish Audio
s2.1-provoice cloning. - All 1,025 translated lines produced TTS audio, with 0 truncated lines. All five dubbed videos passed full video/audio decode checks and frame inspection.
VideoLingo v3.1.1 ✨
字幕流程更灵活,长视频配音更稳定,并正式接入 Fish Audio 声音克隆。
🚀 这次更新
- 新增 Fish Audio 官方配音。 新的
fish_audioProvider 直接调用 Fish Audio 官方 API,默认模型为s2.1-pro。可以从原视频中取 15–30 秒语音克隆音色,不会在 Fish Audio 账户中保存声音;也可以选择固定音色或自定义 Voice ID。 - OpenAI 兼容 TTS 现在可以使用 OpenLux。 OpenAI TTS 后端支持 OpenLux 的语音端点,并可自定义 Base URL 和模型。
- 直接使用现有字幕。 可导入 SRT 进行翻译,无论是否有视频;也可只跑转录阶段,仅生成原文字幕。
- 流程可以暂停检查。 可选择在翻译前修改术语表,也可在生成字幕前检查并修改翻译表格。
- 字幕输出更可控。 新增翻译风格,可调整字体、字号和颜色,并可分别为字幕视频和配音视频选择双语、译文或原文字幕。
- 新增可选的 MAI-Transcribe-2 识别后端,支持 Azure Speech 和 OpenRouter;默认仍然是本地 Qwen3-ASR。
- 长视频处理更稳定。 Demucs 会分块处理长音频,日语 NLP 支持分批,字幕切分有更稳妥的降级方案,翻译和标点边界情况不再让整条流程中断。
- 配音失败不再被隐藏。 TTS 重试失败后会明确中止,不会生成静音冒充成功。语速有上限,仍然放不下的句子会写入
output/log/dub_truncated.json;关闭烧录字幕时也能正常生成配音视频。
⚠️ 升级前留意
- 302.ai 的旧
fish_tts和azure_tts已移除。请改用 Fish Audio、Edge TTS 或其他可用 Provider;默认已改为 Edge TTS。 - Fish Audio 需要配置
fish_audio.api_key,账户也需要有可用额度。新依赖msgpack会随正常安装流程自动安装。 - Qwen3-ASR 模型大小现在是
config.yaml中的高级选项:whisper.qwen_model。
✅ 发布验证
- 自动化测试:663 passed,19 skipped,12 subtests passed。
- 使用 5 个真实英文视频(每个约 7.5–11.3 分钟)完整跑通简体中文、日语、法语、西班牙语和繁体中文,翻译使用 OpenLux GPT-5.6 Luna,配音使用 Fish Audio
s2.1-pro声音克隆。 - 全部 1,025 条译文都成功生成 TTS,0 条被截断。5 个配音成片全部通过完整音视频解码和抽帧检查。
Full Changelog / 完整改动:v3.1.0...v3.1.1