v0.105.1
Bug fixes
- Release unused CUDA cache after each transcription in the transformers ASR backend, keeping model weights loaded for the next request. Long recordings no longer leave their peak temporary GPU allocation reserved throughout the idle timeout. In a Qwen3-ASR test, idle process VRAM dropped from 10.89 GiB to 4.13 GiB. (#642)
- Clear temporary tensor references held by failed-request tracebacks, including chained and grouped exceptions, before releasing CUDA cache. CPU and MPS behavior is unchanged. (#642)
Full changelog: v0.105.0...v0.105.1