Skip to content

v0.3.1

Choose a tag to compare

@andimarafioti andimarafioti released this 15 Jul 14:53
· 10 commits to main since this release
8c3c791

Highlights

  • Add a public model.warmup(prefill_len=100) API shared by the Torch and GGML backends.
  • Keep the former private Torch _warmup() entry point as a compatibility alias.
  • Capture CUDA graphs once on Torch while treating GGML warmup as a safe no-op until qwentts.cpp exposes a native preparation hook.
  • Update internal, demo, example, and test callers to use the public API.
  • Install the published PyPI package in the demo Space deployment.

This resolves the backend interface mismatch reported in huggingface/speech-to-speech#354.

Installation

pip install faster-qwen3-tts==0.3.1

For the optional GGML backend:

pip install "faster-qwen3-tts[ggml]==0.3.1"

Changes

  • #128 — Expose public backend warmup API
  • #122 — Use the PyPI package in the demo Space

Full changelog: v0.3.0...v0.3.1