-
Notifications
You must be signed in to change notification settings - Fork 10
AI Features
AI features are lazy-loaded and run locally in supported browser workflows. The interface reports the backend actually used rather than silently substituting another model.
AI 功能按需加载,并在受支持的浏览器流程中本地运行。界面会显示实际使用的后端,不会静默换用其他模型。
-
Chinese: Piper/VITS browser ONNX, with Xiao Ya selected by default and Chaowen available.
-
English: Kokoro 82M v1.0 ONNX, preferring FP32 WebGPU and explicitly falling back to q8 WASM.
-
Browser Piper voices are available for supported German, Spanish, French, Italian, and Brazilian Portuguese workflows.
-
Voice cards play their matching sample when selected. Generated timeline audio remains a separate preview.
-
中文:Piper/VITS 浏览器 ONNX,默认选择小雅,并提供超文。
-
英文:Kokoro 82M v1.0 ONNX,优先 FP32 WebGPU,必要时明确回退至 q8 WASM。
-
德语、西班牙语、法语、意大利语和巴西葡萄牙语提供经过工作流验证的浏览器 Piper 声音。
-
点击声音卡片会播放对应样音;生成后的时间线语音使用独立播放器。
Automatic captions use Whisper small q8 ONNX. Chinese ASR runs on WASM for stability, applies conservative high-confidence cleanup, and adjusts coarse timestamps toward nearby audio energy.
自动字幕使用 Whisper small q8 ONNX。中文识别使用 WASM 以提高稳定性,只进行克制的高置信纠错,并根据附近音频能量修正粗略时间戳。
YOLOS tiny and MODNet support subject detection, smart framing, caption avoidance, and background removal. Video analysis produces a timestamped temporal track so preview and export resolve the result at the current source time instead of freezing the first frame.
YOLOS tiny 与 MODNet 用于主体检测、智能构图、字幕避让和背景移除。视频分析生成带时间戳的时序结果,预览和导出会按当前源时间解析,而不是固定使用第一帧。
AI vocal separation is available from the Audio workflow and supported clip context menus. It keeps the vocal result on the relevant audio lane and places the instrumental stem on Music.
AI 人声分离可从 Audio 工作区和受支持的片段菜单启动。人声结果保留在对应音频轨,伴奏放入 Music 轨。
The Digital Human workflow uses JoyVASA audio-to-motion and LivePortrait neural rendering on WebGPU. It offers a 256px preview path and a slower 512px quality path. Progress reporting reflects the real model stages and rendering cost.
数字人使用 JoyVASA 音频驱动与 LivePortrait 神经渲染,并通过 WebGPU 运行。提供 256px 快速预览和较慢的 512px 高质量路径,进度提示对应真实模型阶段与计算耗时。