build: qualify uv profiles with patched cu124 and experimental ROCm - #46
Conversation
- 依赖声明收敛到 pyproject.toml + uv.lock 唯一锁:删除 requirements.txt / requirements-dev.txt / requirements-cu124.txt / requirements/locks/ 与旧锁 生成器 tools/generate_python_locks.py - 单一 .venv 逐层 L1→L4:uv sync 前缀递增(voice/vad/local-cu124 extras); torch 经 [tool.uv.sources] 仅在 Windows + local-cu124 时路由到 PyTorch cu124 index;vad 钉 torch==2.5.1,L3→L4 为同版本 wheel 换源而非版本跳变 - Windows CI 重构为 uv 流,新增 single-venv-ladder job(L1→L4 单环境 + torch +cu124 断言 + 负向第二 venv 检查);新增 macOS CI(L1+L2 voice: brew portaudio + PyAudio 源码编译 + verify + ruff + 平台安全契约测试 + electron build) - 工具迁移:verify_python_environment 拆门(cpu/ci/voice 跨平台, vad/cu124 Windows-only fail-closed;pip check → uv pip check); audit_cu124_dependencies 声明侧改读 pyproject 展开; verify_clean_python_install.ps1 改 uv - 测试契约迁移:profile_contract / cu124-audit 改测 pyproject extras、 tool.uv.sources 路由、uv.lock 双 torch 平台条目 - 运行时与文档收敛:Electron venv 探测收敛 [.venv];删除 run_electron_cu124.bat(run_electron_utf8.bat 唯一入口);asr/llm 运行时 第二 venv 引用清零;README/CONTRIBUTING/release policy/README_EN 全部 uv 化 本地已验证(macOS M4):干净 venv L1→L3 逐层 sync+import、uv lock --check、 契约测试 26+50、ruff、electron tsc+build、macOS CI 本地等价复刻全绿; Windows L1→L4 ladder 以 CI job 绿为准。
## What and why The embedded GPT-SoVITS path assumes CUDA in several runtime-only code paths: device selection defaults to CUDA, synchronization and device contexts call `torch.cuda` unconditionally, and BigVGAN probes/loads CUDA kernels even when inference runs on another backend. On Apple Silicon this prevents the bundled local TTS backend from running through PyTorch MPS. This PR adds the smallest runtime compatibility layer needed for Apple Silicon: - resolve `TTS_DEVICE=auto` to `mps` on Apple Silicon and `cpu` on Intel macOS; - validate requested CUDA/MPS devices with actionable errors; - use device-neutral context and synchronization helpers; - keep BigVGAN CUDA kernels on CUDA while using its PyTorch implementation on MPS/CPU; - load supported GPT-SoVITS checkpoints with PyTorch 2.6 safe loading; - enable PyTorch MPS CPU fallback for the Electron-launched backend. Performance-policy changes from the local macOS prototype—lower sample steps, frame-rate caps, lazy animation loading, static secondary displays, and backend auto-restart—are intentionally excluded. Linked Issue for product-semantic or public-contract changes: #47 Related dependency-profile work: #46. This PR intentionally does not add Darwin packages to the CUDA-named `local-cu124` extra; the install-profile contract can be integrated with the ongoing dependency migration. ## Change class - [ ] Routine fix, documentation, test, maintenance, or presentation-only UI - [x] Product-semantic or public-contract change discussed in the linked Issue - [ ] Isolated, default-off experiment Owning layer: embedded GPT-SoVITS runtime and backend process environment User-visible effect, or `none`: Apple Silicon users can select `TTS_DEVICE=auto` and run compatible GPT-SoVITS v3 models through MPS. Compatibility or migration impact, or `none`: Explicit indexed CUDA, `mps`, and `cpu` values remain supported. Apple Silicon now resolves `auto`/unindexed `cuda` to `mps`; Intel macOS resolves them to `cpu`. A compatible PyTorch/MPS environment is still required; dependency installation is intentionally left to #46. ## Evidence Commands and manual journeys run: - New device/checkpoint tests — 11 passed - Existing relevant TTS/config/profile tests — 84 passed - Focused Ruff checks — passed - Electron `npm run build` — passed - Electron existing tests — 4 passed - Apple Silicon verifier: Python 3.12.13, torch/torchaudio 2.6.0, `mps_built=True`, `mps_available=True` - Manual GPT-SoVITS v3 synthesis on an M1 Max using the local model package — passed - `git diff --check` — passed The full macOS collection on current `main` is independently blocked by an eager import of the Windows-only pointer hook; #48 contains the isolated fix. The remaining 23 macOS failures (temporary-path alias expectations and Windows desktop-layer assertions) reproduce unchanged on `upstream/main`. - [x] Relevant Python tests pass - [x] CPU/model-less baseline remains supported - [x] Electron `npm run build` passes when Electron code changed - [ ] Dependency audit passes when dependencies changed - [ ] Before/after screenshots are attached for visible UI changes - [x] Documentation/examples are updated for changed settings or contracts ## Final check - [x] This PR addresses one coherent problem without unrelated cleanup - [x] It does not add a speculative API, fallback, or compatibility path - [x] No secrets, local state, model weights, voice material, or restricted assets are included - [x] Third-party notices and provenance are preserved --------- Co-authored-by: Lucas1479 <sli776@aucklanduni.ac.nz>
|
已完成本轮集成并将 #46 转为 Ready for review:
仓库保护规则仍要求另一位维护者批准;未使用管理员覆盖。 This integration is complete and #46 is now ready for review:
Repository protection still requires approval from another maintainer; no administrator override was used. |
|
@Mieluoxxx 你好,#46 已基于你在 #45 的原始提交完成维护集成,现请求你帮忙审阅。重点请看:单 Hello @Mieluoxxx, #46 has completed the maintainer integration based on your original commits in #45, and we would appreciate your review. The main areas are the single- |
The English README still presented CUDA local voice as the starting installation path after #46 moved the Chinese guide to capability tiers. This brings the English quick start, hardware requirements, macOS launch commands, profile/rollback guidance, first-run B2 behavior, and release boundaries into alignment. It also corrects two outdated Chinese statements about macOS CI and the distinction between installing assets and installing model dependencies. The shared Star History SVG is manually regenerated from GitHub's current stargazer timestamps using the existing renderer: **154 → 191 current stars**, latest star date **2026-09-05**. The current chart is referenced by both READMEs. Validation: checked local Markdown links, code fences and profile commands; parsed the SVG and visually inspected its rendered output; `git diff --check` passed. Documentation and chart changes only. --- #46 合并后,英文 README 仍把 CUDA 本地语音作为默认安装入口,没有完整同步中文的能力分层。本次补齐英文快速开始、硬件要求、macOS 启动命令、安装配置与回退说明、首次启动的 B2 行为和发布边界;同时修正中文关于 macOS CI 和“安装资产不等于补齐模型依赖”的两处过时表述。 使用现有生成器,根据 GitHub 最新 star 时间戳手动刷新两份 README 共用的 Star History:**154 → 191 星**,最近一颗 star 的日期为 **2026-09-05**。 已检查本地链接、代码围栏、安装配置命令、SVG XML 和实际图表渲染,`git diff --check` 通过。本次仅修改文档与图表。
This PR builds on #45 and preserves the original contributor commits. It qualifies one exact uv-managed project environment before changing the maintained installation path.
The default capability ladder is L1 core → L2 remote voice → L3 CPU VAD → L4 Windows cu124 local models. Every tier uses the project
.venv, and CPU, NVIDIA, and ROCm Torch builds are mutually exclusive. CPU VAD and cu124 now use PyTorch/Torchaudio 2.6.0 as the maintained baseline. This is the smallest official cu124 upgrade that fixes critical advisory GHSA-53q9-r3pm-6pq6; the previous 2.5.1 lock was rejected by dependency review. The Windows local-model profile also uses thepyopenjtalk-plusCPython 3.12 wheel so a clean install does not depend on compiling pyopenjtalk inside a long checkout path.This branch includes the merged portability, Apple Silicon MPS runtime, native macOS menu, and dependency-review changes from #48, #49, #51, and #52.
An opt-in
local-rocmcandidate is also included for Windows. It locks AMD's official ROCm 7.2.1 / Torch 2.9.1 packages as a third build selection in the same.venv. Qwen ASR and GPT-SoVITS may run in persistent sidecar processes while using that interpreter; sidecar mode is disabled unless explicitly selected. NVIDIA CUDA Graph, NVIDIA BigVGAN kernels, and unrelated performance tuning are not enabled by this candidate.The latest integration includes public main
f7e57e2and fixes the offline-loading boundary: inherited online flags, already-imported Hub/Transformers state, and cached/custom Hub sessions are now reset before local voice-model loading. Explicit local-file loading is applied to Qwen ASR, BERT, and BigVGAN. The new regression checks block remote requests while still loading a tiny locally generated BERT model. They ran successfully with the qualified cu124 and ROCm libraries. Current headf1f6397has all six remote checks green; the current Windows model-less suite reports 1769 passed / 11 skipped, and the Electron model-less smoke passed.Recorded validation (previous L4 qualification and latest integration checks):
uv.lockresolves CPU 2.6.0, cu124 2.6.0, and ROCm 2.9.1/7.2.1 branches; invalid combinations fail closed.A clean
local-cu124sync passed the 234-package environment contract anduv pip check.On an RTX 4070 Ti SUPER, PyTorch 2.6.0+cu124 completed real CUDA matrix compute, safely loaded the existing GPT and SoVITS v3 checkpoints, and generated a finite 1.612-second / 24 kHz v3 TTS sample through BERT, CNHubert, GPT, SoVITS LoRA, and BigVGAN.
The full L4 Python suite passed: 1783 passed, 2 skipped.
Exact sync back to L1+dev removed 127 voice/model packages and verified that Torch, Qwen ASR, ONNX Runtime, and pyopenjtalk were absent.
A clean
local-rocmsync installed the fixed candidate and passed its version/import contract plusuv pip check.Earlier revisions passed the remote Windows model-less, single-venv ladder, macOS voice, ROCm clean-install, and Electron build jobs. The Torch 2.5.1 critical advisory is removed. Dependency review has narrowly documented four temporary exceptions: NLTK GHSA-8mgp-746c-j5xp has no patched release, while qwen-asr 0.0.6 requires Transformers 4.57.6 exactly and therefore cannot consume the fixes for GHSA-29pf-2h5f-8g72, GHSA-fgcw-684q-jj6r, and GHSA-xrqw-3rrv-vx5w. Amadeus does not call the affected NLTK persistence, Transformers LightGlue, or
save_pretrainedpaths; ASR resolves a local directory and both ASR and TTS force Transformers/Hugging Face offline. Each exception must be removed when a compatible fix is published. All six remote checks have now passed again forf1f6397.Ruff, workflow YAML validation, lock consistency, focused profile/sidecar tests, and diff checks passed.
ROCm evidence remains explicitly experimental. Community history records successful RX 9070 XT ASR/TTS sidecars on another ROCm/PyTorch build. On the maintainer's Radeon 780M, the fixed 7.2.1 build installed and enumerated gfx1103, but its first FP32 tensor operation crashed in
amdhip64_7.dll; Radeon 780M is absent from AMD's Windows support matrix and is not treated as a supported result. A supported AMD GPU still needs to complete the fixed-combination ASR/TTS, microphone/playback, interruption, lifecycle, and long-running journeys.The PR is ready for maintainer review. ROCm remains experimental until supported AMD hardware completes the remaining real-device acceptance. No model weights, recordings, transcripts, generated audio, credentials, virtual environments, or validation caches are committed. Refs #44 and #45.
中文摘要
本 PR 基于 #45,并保留原贡献者提交,用于在切换维护基线前验证由 uv 管理的单一项目环境。
默认能力阶梯是 L1 core → L2 远程语音 → L3 CPU VAD → L4 Windows cu124 本地模型。所有梯级共用项目
.venv,CPU、NVIDIA 和 ROCm Torch 构建两两互斥。CPU VAD 与 cu124 现以 PyTorch/Torchaudio 2.6.0 作为正式维护基线;这是仍提供官方 cu124 wheel、同时修复 critical 漏洞 GHSA-53q9-r3pm-6pq6 的最小升级。旧 2.5.1 锁正是 dependency review 失败的原因。Windows 本地模型档也统一采用带 CPython 3.12 wheel 的pyopenjtalk-plus,避免在较长仓库路径中现场编译 pyopenjtalk。本分支现已合入 #48、#49、#51、#52 的跨平台导入、Apple Silicon MPS runtime、macOS 原生菜单和 dependency review 改动。
Windows
local-rocm仍是默认关闭的实验候选。它在同一个.venv中锁定 AMD 官方 ROCm 7.2.1 / Torch 2.9.1,并与 CPU/cu124 构建互斥。Qwen ASR 与 GPT-SoVITS 可使用同一解释器运行在常驻 sidecar 子进程;sidecar 只表示进程隔离,不额外要求虚拟环境,也不会默认启用 NVIDIA CUDA Graph、NVIDIA BigVGAN kernel 或社区补丁中的其他性能调优。本轮已接上公开主线
f7e57e2,并补齐模型离线加载边界:继承的在线环境变量、已导入的 Hub/Transformers 状态和缓存/自定义 Hub HTTP 会话都会在本地语音模型加载前恢复为离线;Qwen ASR、BERT、BigVGAN 的加载也明确只使用本地文件。新增测试验证远程请求被阻止,同时本地生成的小型 BERT 模型仍能正常加载,已在 cu124 与 ROCm 的实际依赖环境中通过。最新提交f1f6397的六项远程检查全绿,其中 Windows 无模型完整回归为 1769 passed / 11 skipped,Electron 无模型冒烟测试通过。验证记录(此前 L4 资格验证与本轮集成检查):
uv.lock可解析 CPU 2.6.0、cu124 2.6.0 和 ROCm 2.9.1/7.2.1,冲突组合会明确失败;全新
local-cu124同步通过 234 包环境合同和uv pip check;RTX 4070 Ti SUPER 上,2.6.0+cu124 完成真实 CUDA 矩阵计算,安全加载现有 GPT/SoVITS v3 权重,并经 BERT、CNHubert、GPT、SoVITS LoRA、BigVGAN 生成 1.612 秒、24 kHz 的有限值音频;
完整 L4 Python 回归 1783 通过、2 跳过;
同一
.venv精确返回 L1+dev 时移除 127 个语音/模型包,并确认 Torch、Qwen ASR、ONNX Runtime 与 pyopenjtalk 均不存在;全新
local-rocm同步通过固定候选版本/导入合同和uv pip check;此前提交的远程 Windows 无模型、单
.venv阶梯、macOS voice、ROCm clean-install 与 Electron build 均通过;Torch 2.5.1 critical 漏洞已移除。dependency review 现精确记录四条临时豁免:NLTK GHSA-8mgp-746c-j5xp 尚无修复版;qwen-asr 0.0.6 又严格要求 Transformers 4.57.6,暂时无法采用 GHSA-29pf-2h5f-8g72、GHSA-fgcw-684q-jj6r、GHSA-xrqw-3rrv-vx5w 的 5.x 修复。Amadeus 不调用相关 NLTK 持久化、LightGlue 或save_pretrained路径;ASR 只解析本地目录,ASR/TTS 均强制 Transformers/Hugging Face 离线。兼容修复发布后必须逐条移除。最新提交f1f6397的六项远程检查已全部通过;Ruff、工作流 YAML、锁一致性、profile/sidecar 聚焦测试和 diff check 均通过。
ROCm 继续明确标记为实验候选。社区资料记录了 RX 9070 XT sidecar ASR/TTS 历史成功;维护机 Radeon 780M 上,固定 7.2.1 环境可安装并枚举 gfx1103,但首次 FP32 计算在
amdhip64_7.dll中崩溃。780M 不在 AMD Windows 支持矩阵内,这个负结果只限定本机硬件,不否定社区候选。固定组合仍需由受支持 AMD GPU 完成真实 ASR/TTS、麦克风/播放、打断、生命周期和长时间运行验收。本 PR 已进入维护者审阅;ROCm 在受支持 AMD 硬件完成剩余实机验收前继续保持实验候选。仓库未提交模型、录音、转写、生成音频、凭证、虚拟环境或验证缓存。关联 #44、#45。