Skip to content

build: qualify uv profiles with patched cu124 and experimental ROCm - #46

Merged
Lucas1479 merged 10 commits into
mainfrom
codex/uv-install-integration
Sep 6, 2026
Merged

build: qualify uv profiles with patched cu124 and experimental ROCm#46
Lucas1479 merged 10 commits into
mainfrom
codex/uv-install-integration

Conversation

@Lucas1479

@Lucas1479 Lucas1479 commented Sep 3, 2026

Copy link
Copy Markdown
Member

This PR builds on #45 and preserves the original contributor commits. It qualifies one exact uv-managed project environment before changing the maintained installation path.

The default capability ladder is L1 core → L2 remote voice → L3 CPU VAD → L4 Windows cu124 local models. Every tier uses the project .venv, and CPU, NVIDIA, and ROCm Torch builds are mutually exclusive. CPU VAD and cu124 now use PyTorch/Torchaudio 2.6.0 as the maintained baseline. This is the smallest official cu124 upgrade that fixes critical advisory GHSA-53q9-r3pm-6pq6; the previous 2.5.1 lock was rejected by dependency review. The Windows local-model profile also uses the pyopenjtalk-plus CPython 3.12 wheel so a clean install does not depend on compiling pyopenjtalk inside a long checkout path.

This branch includes the merged portability, Apple Silicon MPS runtime, native macOS menu, and dependency-review changes from #48, #49, #51, and #52.

An opt-in local-rocm candidate is also included for Windows. It locks AMD's official ROCm 7.2.1 / Torch 2.9.1 packages as a third build selection in the same .venv. Qwen ASR and GPT-SoVITS may run in persistent sidecar processes while using that interpreter; sidecar mode is disabled unless explicitly selected. NVIDIA CUDA Graph, NVIDIA BigVGAN kernels, and unrelated performance tuning are not enabled by this candidate.

The latest integration includes public main f7e57e2 and fixes the offline-loading boundary: inherited online flags, already-imported Hub/Transformers state, and cached/custom Hub sessions are now reset before local voice-model loading. Explicit local-file loading is applied to Qwen ASR, BERT, and BigVGAN. The new regression checks block remote requests while still loading a tiny locally generated BERT model. They ran successfully with the qualified cu124 and ROCm libraries. Current head f1f6397 has all six remote checks green; the current Windows model-less suite reports 1769 passed / 11 skipped, and the Electron model-less smoke passed.

Recorded validation (previous L4 qualification and latest integration checks):

  • uv.lock resolves CPU 2.6.0, cu124 2.6.0, and ROCm 2.9.1/7.2.1 branches; invalid combinations fail closed.

  • A clean local-cu124 sync passed the 234-package environment contract and uv pip check.

  • On an RTX 4070 Ti SUPER, PyTorch 2.6.0+cu124 completed real CUDA matrix compute, safely loaded the existing GPT and SoVITS v3 checkpoints, and generated a finite 1.612-second / 24 kHz v3 TTS sample through BERT, CNHubert, GPT, SoVITS LoRA, and BigVGAN.

  • The full L4 Python suite passed: 1783 passed, 2 skipped.

  • Exact sync back to L1+dev removed 127 voice/model packages and verified that Torch, Qwen ASR, ONNX Runtime, and pyopenjtalk were absent.

  • A clean local-rocm sync installed the fixed candidate and passed its version/import contract plus uv pip check.

  • Earlier revisions passed the remote Windows model-less, single-venv ladder, macOS voice, ROCm clean-install, and Electron build jobs. The Torch 2.5.1 critical advisory is removed. Dependency review has narrowly documented four temporary exceptions: NLTK GHSA-8mgp-746c-j5xp has no patched release, while qwen-asr 0.0.6 requires Transformers 4.57.6 exactly and therefore cannot consume the fixes for GHSA-29pf-2h5f-8g72, GHSA-fgcw-684q-jj6r, and GHSA-xrqw-3rrv-vx5w. Amadeus does not call the affected NLTK persistence, Transformers LightGlue, or save_pretrained paths; ASR resolves a local directory and both ASR and TTS force Transformers/Hugging Face offline. Each exception must be removed when a compatible fix is published. All six remote checks have now passed again for f1f6397.

  • Ruff, workflow YAML validation, lock consistency, focused profile/sidecar tests, and diff checks passed.

ROCm evidence remains explicitly experimental. Community history records successful RX 9070 XT ASR/TTS sidecars on another ROCm/PyTorch build. On the maintainer's Radeon 780M, the fixed 7.2.1 build installed and enumerated gfx1103, but its first FP32 tensor operation crashed in amdhip64_7.dll; Radeon 780M is absent from AMD's Windows support matrix and is not treated as a supported result. A supported AMD GPU still needs to complete the fixed-combination ASR/TTS, microphone/playback, interruption, lifecycle, and long-running journeys.

The PR is ready for maintainer review. ROCm remains experimental until supported AMD hardware completes the remaining real-device acceptance. No model weights, recordings, transcripts, generated audio, credentials, virtual environments, or validation caches are committed. Refs #44 and #45.

中文摘要

本 PR 基于 #45,并保留原贡献者提交,用于在切换维护基线前验证由 uv 管理的单一项目环境。

默认能力阶梯是 L1 core → L2 远程语音 → L3 CPU VAD → L4 Windows cu124 本地模型。所有梯级共用项目 .venv,CPU、NVIDIA 和 ROCm Torch 构建两两互斥。CPU VAD 与 cu124 现以 PyTorch/Torchaudio 2.6.0 作为正式维护基线;这是仍提供官方 cu124 wheel、同时修复 critical 漏洞 GHSA-53q9-r3pm-6pq6 的最小升级。旧 2.5.1 锁正是 dependency review 失败的原因。Windows 本地模型档也统一采用带 CPython 3.12 wheel 的 pyopenjtalk-plus,避免在较长仓库路径中现场编译 pyopenjtalk。

本分支现已合入 #48#49#51#52 的跨平台导入、Apple Silicon MPS runtime、macOS 原生菜单和 dependency review 改动。

Windows local-rocm 仍是默认关闭的实验候选。它在同一个 .venv 中锁定 AMD 官方 ROCm 7.2.1 / Torch 2.9.1,并与 CPU/cu124 构建互斥。Qwen ASR 与 GPT-SoVITS 可使用同一解释器运行在常驻 sidecar 子进程;sidecar 只表示进程隔离,不额外要求虚拟环境,也不会默认启用 NVIDIA CUDA Graph、NVIDIA BigVGAN kernel 或社区补丁中的其他性能调优。

本轮已接上公开主线 f7e57e2,并补齐模型离线加载边界:继承的在线环境变量、已导入的 Hub/Transformers 状态和缓存/自定义 Hub HTTP 会话都会在本地语音模型加载前恢复为离线;Qwen ASR、BERT、BigVGAN 的加载也明确只使用本地文件。新增测试验证远程请求被阻止,同时本地生成的小型 BERT 模型仍能正常加载,已在 cu124 与 ROCm 的实际依赖环境中通过。最新提交 f1f6397 的六项远程检查全绿,其中 Windows 无模型完整回归为 1769 passed / 11 skipped,Electron 无模型冒烟测试通过。

验证记录(此前 L4 资格验证与本轮集成检查):

  • uv.lock 可解析 CPU 2.6.0、cu124 2.6.0 和 ROCm 2.9.1/7.2.1,冲突组合会明确失败;

  • 全新 local-cu124 同步通过 234 包环境合同和 uv pip check

  • RTX 4070 Ti SUPER 上,2.6.0+cu124 完成真实 CUDA 矩阵计算,安全加载现有 GPT/SoVITS v3 权重,并经 BERT、CNHubert、GPT、SoVITS LoRA、BigVGAN 生成 1.612 秒、24 kHz 的有限值音频;

  • 完整 L4 Python 回归 1783 通过、2 跳过;

  • 同一 .venv 精确返回 L1+dev 时移除 127 个语音/模型包,并确认 Torch、Qwen ASR、ONNX Runtime 与 pyopenjtalk 均不存在;

  • 全新 local-rocm 同步通过固定候选版本/导入合同和 uv pip check

  • 此前提交的远程 Windows 无模型、单 .venv 阶梯、macOS voice、ROCm clean-install 与 Electron build 均通过;Torch 2.5.1 critical 漏洞已移除。dependency review 现精确记录四条临时豁免:NLTK GHSA-8mgp-746c-j5xp 尚无修复版;qwen-asr 0.0.6 又严格要求 Transformers 4.57.6,暂时无法采用 GHSA-29pf-2h5f-8g72GHSA-fgcw-684q-jj6rGHSA-xrqw-3rrv-vx5w 的 5.x 修复。Amadeus 不调用相关 NLTK 持久化、LightGlue 或 save_pretrained 路径;ASR 只解析本地目录,ASR/TTS 均强制 Transformers/Hugging Face 离线。兼容修复发布后必须逐条移除。最新提交 f1f6397 的六项远程检查已全部通过;

  • Ruff、工作流 YAML、锁一致性、profile/sidecar 聚焦测试和 diff check 均通过。

ROCm 继续明确标记为实验候选。社区资料记录了 RX 9070 XT sidecar ASR/TTS 历史成功;维护机 Radeon 780M 上,固定 7.2.1 环境可安装并枚举 gfx1103,但首次 FP32 计算在 amdhip64_7.dll 中崩溃。780M 不在 AMD Windows 支持矩阵内,这个负结果只限定本机硬件,不否定社区候选。固定组合仍需由受支持 AMD GPU 完成真实 ASR/TTS、麦克风/播放、打断、生命周期和长时间运行验收。

本 PR 已进入维护者审阅;ROCm 在受支持 AMD 硬件完成剩余实机验收前继续保持实验候选。仓库未提交模型、录音、转写、生成音频、凭证、虚拟环境或验证缓存。关联 #44#45

Mieluoxxx and others added 3 commits September 3, 2026 19:03
- 依赖声明收敛到 pyproject.toml + uv.lock 唯一锁:删除 requirements.txt /
  requirements-dev.txt / requirements-cu124.txt / requirements/locks/ 与旧锁
  生成器 tools/generate_python_locks.py
- 单一 .venv 逐层 L1→L4:uv sync 前缀递增(voice/vad/local-cu124 extras);
  torch 经 [tool.uv.sources] 仅在 Windows + local-cu124 时路由到 PyTorch
  cu124 index;vad 钉 torch==2.5.1,L3→L4 为同版本 wheel 换源而非版本跳变
- Windows CI 重构为 uv 流,新增 single-venv-ladder job(L1→L4 单环境 +
  torch +cu124 断言 + 负向第二 venv 检查);新增 macOS CI(L1+L2 voice:
  brew portaudio + PyAudio 源码编译 + verify + ruff + 平台安全契约测试 +
  electron build)
- 工具迁移:verify_python_environment 拆门(cpu/ci/voice 跨平台,
  vad/cu124 Windows-only fail-closed;pip check → uv pip check);
  audit_cu124_dependencies 声明侧改读 pyproject 展开;
  verify_clean_python_install.ps1 改 uv
- 测试契约迁移:profile_contract / cu124-audit 改测 pyproject extras、
  tool.uv.sources 路由、uv.lock 双 torch 平台条目
- 运行时与文档收敛:Electron venv 探测收敛 [.venv];删除
  run_electron_cu124.bat(run_electron_utf8.bat 唯一入口);asr/llm 运行时
  第二 venv 引用清零;README/CONTRIBUTING/release policy/README_EN 全部
  uv 化

本地已验证(macOS M4):干净 venv L1→L3 逐层 sync+import、uv lock --check、
契约测试 26+50、ruff、electron tsc+build、macOS CI 本地等价复刻全绿;
Windows L1→L4 ladder 以 CI job 绿为准。
@tocekuma

tocekuma commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Apple Silicon runtime support is now split into #49. It deliberately avoids changing the CUDA-named local-cu124 extra, so the dependency-profile contract can stay with this migration PR. The remaining macOS desktop-host work is tracked in #47.

Lucas1479 added a commit that referenced this pull request Sep 4, 2026
## What and why

The embedded GPT-SoVITS path assumes CUDA in several runtime-only code
paths: device selection defaults to CUDA, synchronization and device
contexts call `torch.cuda` unconditionally, and BigVGAN probes/loads
CUDA kernels even when inference runs on another backend. On Apple
Silicon this prevents the bundled local TTS backend from running through
PyTorch MPS.

This PR adds the smallest runtime compatibility layer needed for Apple
Silicon:

- resolve `TTS_DEVICE=auto` to `mps` on Apple Silicon and `cpu` on Intel
macOS;
- validate requested CUDA/MPS devices with actionable errors;
- use device-neutral context and synchronization helpers;
- keep BigVGAN CUDA kernels on CUDA while using its PyTorch
implementation on MPS/CPU;
- load supported GPT-SoVITS checkpoints with PyTorch 2.6 safe loading;
- enable PyTorch MPS CPU fallback for the Electron-launched backend.

Performance-policy changes from the local macOS prototype—lower sample
steps, frame-rate caps, lazy animation loading, static secondary
displays, and backend auto-restart—are intentionally excluded.

Linked Issue for product-semantic or public-contract changes: #47

Related dependency-profile work: #46. This PR intentionally does not add
Darwin packages to the CUDA-named `local-cu124` extra; the
install-profile contract can be integrated with the ongoing dependency
migration.

## Change class

- [ ] Routine fix, documentation, test, maintenance, or
presentation-only UI
- [x] Product-semantic or public-contract change discussed in the linked
Issue
- [ ] Isolated, default-off experiment

Owning layer: embedded GPT-SoVITS runtime and backend process
environment

User-visible effect, or `none`: Apple Silicon users can select
`TTS_DEVICE=auto` and run compatible GPT-SoVITS v3 models through MPS.

Compatibility or migration impact, or `none`: Explicit indexed CUDA,
`mps`, and `cpu` values remain supported. Apple Silicon now resolves
`auto`/unindexed `cuda` to `mps`; Intel macOS resolves them to `cpu`. A
compatible PyTorch/MPS environment is still required; dependency
installation is intentionally left to #46.

## Evidence

Commands and manual journeys run:

- New device/checkpoint tests — 11 passed
- Existing relevant TTS/config/profile tests — 84 passed
- Focused Ruff checks — passed
- Electron `npm run build` — passed
- Electron existing tests — 4 passed
- Apple Silicon verifier: Python 3.12.13, torch/torchaudio 2.6.0,
`mps_built=True`, `mps_available=True`
- Manual GPT-SoVITS v3 synthesis on an M1 Max using the local model
package — passed
- `git diff --check` — passed

The full macOS collection on current `main` is independently blocked by
an eager import of the Windows-only pointer hook; #48 contains the
isolated fix. The remaining 23 macOS failures (temporary-path alias
expectations and Windows desktop-layer assertions) reproduce unchanged
on `upstream/main`.

- [x] Relevant Python tests pass
- [x] CPU/model-less baseline remains supported
- [x] Electron `npm run build` passes when Electron code changed
- [ ] Dependency audit passes when dependencies changed
- [ ] Before/after screenshots are attached for visible UI changes
- [x] Documentation/examples are updated for changed settings or
contracts

## Final check

- [x] This PR addresses one coherent problem without unrelated cleanup
- [x] It does not add a speculative API, fallback, or compatibility path
- [x] No secrets, local state, model weights, voice material, or
restricted assets are included
- [x] Third-party notices and provenance are preserved

---------

Co-authored-by: Lucas1479 <sli776@aucklanduni.ac.nz>
@Lucas1479 Lucas1479 changed the title build: qualify uv install profiles and migration gates build: qualify uv profiles and experimental ROCm sidecars Sep 4, 2026
@Lucas1479 Lucas1479 changed the title build: qualify uv profiles and experimental ROCm sidecars build: qualify uv profiles with patched cu124 and experimental ROCm Sep 4, 2026
@Lucas1479
Lucas1479 marked this pull request as ready for review September 4, 2026 13:52
@Lucas1479

Copy link
Copy Markdown
Member Author

已完成本轮集成并将 #46 转为 Ready for review:

  • CPU VAD 与 CUDA 12.4 的产品基线升级为 PyTorch/Torchaudio 2.6.0,移除了 2.5.1 的 critical RCE 告警;
  • cu124 clean install、RTX 4070 Ti SUPER CUDA 计算、现有 v3 权重加载与真实 TTS 均通过;完整 L4 回归为 1783 passed / 2 skipped;
  • pyopenjtalk-plus 解决 Windows 长路径下的源码编译失败;
  • ROCm 7.2.1/Torch 2.9.1 作为同一 .venv 中的互斥实验候选,远程 clean-install 通过;Radeon 780M 不在官方矩阵内,其失败不作为候选阻塞;
  • 六项远程检查全部通过。NLTK/Transformers 的四条临时 GHSA 豁免已逐条记录调用边界;本地模型加载强制离线,兼容修复发布后应移除。

仓库保护规则仍要求另一位维护者批准;未使用管理员覆盖。


This integration is complete and #46 is now ready for review:

  • The maintained CPU VAD and CUDA 12.4 baseline is PyTorch/Torchaudio 2.6.0, removing the critical RCE affecting 2.5.1.
  • The cu124 clean install, real CUDA compute on an RTX 4070 Ti SUPER, existing v3 checkpoint loading, and end-to-end TTS passed; the full L4 suite reports 1783 passed / 2 skipped.
  • pyopenjtalk-plus removes the Windows long-path source-build failure.
  • ROCm 7.2.1/Torch 2.9.1 remains a mutually exclusive experimental candidate in the same .venv; its remote clean install passed. The unsupported Radeon 780M result does not block the candidate.
  • All six remote checks passed. Four temporary NLTK/Transformers GHSA exceptions are documented by exact advisory and reachable code boundary; local model loading is forced offline, and each exception should be removed when a compatible fix is available.

Repository protection still requires approval from another maintainer; no administrator override was used.

@Lucas1479
Lucas1479 requested a review from Mieluoxxx September 4, 2026 13:55
@Lucas1479

Copy link
Copy Markdown
Member Author

@Mieluoxxx 你好,#46 已基于你在 #45 的原始提交完成维护集成,现请求你帮忙审阅。重点请看:单 .venv 的分级/互斥规则、Torch 2.6.0 cu124 正式基线、pyopenjtalk-plus Windows wheel,以及默认关闭的 ROCm 7.2.1 实验候选。当前本地完整 L4 回归为 1783 passed / 2 skipped,远程六项检查全部通过。若你发现问题,请直接留下 review 意见,我们会继续修正。

Hello @Mieluoxxx, #46 has completed the maintainer integration based on your original commits in #45, and we would appreciate your review. The main areas are the single-.venv tier/conflict rules, the maintained Torch 2.6.0 cu124 baseline, the pyopenjtalk-plus Windows wheel, and the opt-in ROCm 7.2.1 experimental candidate. The full local L4 suite reports 1783 passed / 2 skipped, and all six remote checks are green. Please leave review feedback directly if you find anything that should be corrected.

@Lucas1479
Lucas1479 merged commit f10a43d into main Sep 6, 2026
6 checks passed
@Lucas1479
Lucas1479 deleted the codex/uv-install-integration branch September 6, 2026 03:49
Lucas1479 added a commit that referenced this pull request Sep 6, 2026
The English README still presented CUDA local voice as the starting
installation path after #46 moved the Chinese guide to capability tiers.
This brings the English quick start, hardware requirements, macOS launch
commands, profile/rollback guidance, first-run B2 behavior, and release
boundaries into alignment. It also corrects two outdated Chinese
statements about macOS CI and the distinction between installing assets
and installing model dependencies.

The shared Star History SVG is manually regenerated from GitHub's
current stargazer timestamps using the existing renderer: **154 → 191
current stars**, latest star date **2026-09-05**. The current chart is
referenced by both READMEs.

Validation: checked local Markdown links, code fences and profile
commands; parsed the SVG and visually inspected its rendered output;
`git diff --check` passed. Documentation and chart changes only.

---

#46 合并后,英文 README 仍把 CUDA
本地语音作为默认安装入口,没有完整同步中文的能力分层。本次补齐英文快速开始、硬件要求、macOS 启动命令、安装配置与回退说明、首次启动的 B2
行为和发布边界;同时修正中文关于 macOS CI 和“安装资产不等于补齐模型依赖”的两处过时表述。

使用现有生成器,根据 GitHub 最新 star 时间戳手动刷新两份 README 共用的 Star History:**154 → 191
星**,最近一颗 star 的日期为 **2026-09-05**。

已检查本地链接、代码围栏、安装配置命令、SVG XML 和实际图表渲染,`git diff --check` 通过。本次仅修改文档与图表。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants