Releases: Quantatirsk/AsrServe
Releases · Quantatirsk/AsrServe
Release list
AsrServe v1.0.4
v1.0.4 replaces the whole model stack on both the backend and the browser client. In our own tests against v1.0.3, recognition accuracy improved by about 20% and end-to-end efficiency by about 40%.
Breaking changes:
- Project rename:
qwen3-asris now AsrServe. The GitHub repository, Python package and Docker images (quantatrisk/asrserve) use the new name; old GitHub URLs redirect automatically. - ASR model: Qwen3-ASR 1.7B/0.6B is replaced by Confucius4-R2T2. Offline and realtime share one R2T2 engine (vLLM on CUDA, vendored Rust on CPU); the
modelrequest field no longer switches models. - Speaker diarization: CAM++ is replaced by Nemotron-3-Diarization (up to 8 speakers). Realtime streams now also carry per-utterance speaker labels.
- Removed models: FSMN VAD, the three CAM++ models and the Qwen3-ASR checkpoints are gone. Nemotron speech activity drives offline segmentation. ModelScope and FunASR are no longer dependencies; all models come from Hugging Face at pinned revisions.
- Removed API: the Alibaba Cloud compatible REST API (
/stream/v1/asr*) is removed. Use the OpenAI-compatible/v1/audio/transcriptionsand the native/v1/streamWebSocket. - Deployment: images are published on Docker Hub as
quantatrisk/asrserve:gpu(amd64),:cpu(amd64/arm64) and:ascend(arm64, Ascend 910B branch), plus1.0.4-*version tags;build.shbuilds from source. Compose files arecompose.yml(GPU) andcompose.cpu.yml(CPU) and no longer build. The only mount is./models.docker-compose*.yml,deploy/prepare.shand the model export option are removed. The oldquantatrisk/qwen3-asrDocker Hub repository is no longer updated. - Runtime: the CUDA image moves to CUDA 13.0 and vLLM 0.30; the service listens on port
17003.
Quick start:
docker compose up -d # GPU: quantatrisk/asrserve:gpu
docker compose -f compose.cpu.yml up -d # CPU: quantatrisk/asrserve:cpumacOS Apple Silicon runs natively without Docker; see the README.
🤖 Generated with Claude Code
v1.0.3
Changes
- Remove the voiceprint database, APIs, and sqlite-vec dependency.
- Make HF_HUB_OFFLINE the single offline deployment switch and use local model snapshots.
- Reduce default deployment configuration and align Compose, environment examples, and documentation.
- Increase internal cold-start readiness timeout to 600 seconds.
- Deprecate numeric weights in hotword context.
Qwen3-ASR v1.0.1
Changes
- Fix speaker-diarization long segment handling by splitting overlong CAM++ segments at low-energy boundaries before the ASR pipeline.
- Change the default offline ASR segment limit to 60 seconds and sync env examples/docs.
- Bump package and runtime version metadata to 1.0.1.
Validation
- uv run python -m compileall app/utils/speaker_diarizer.py app/core/config.py
- uvx pyright app/utils/speaker_diarizer.py app/core/config.py
- Speaker split smoke test on the 474s sample produced segments under 60 seconds.
Qwen3-ASR v1.0.0
Qwen3-ASR v1.0.0
English
Qwen3-ASR v1.0.0 is the first release after the project rename from funasr-api to qwen3-asr. The service is now centered on Qwen3-ASR, while keeping the existing FunASR/Paraformer realtime websocket compatibility path where it is still part of the runtime.
Highlights
- Renamed the project metadata, Docker Compose services, Docker image names, default logs, startup UI, benchmark reports, and core documentation to
qwen3-asr. - Updated GitHub Actions Docker publishing to push
docker.io/quantatrisk/qwen3-asr:cpu-latest,docker.io/quantatrisk/qwen3-asr:gpu-latest, anddocker.io/quantatrisk/qwen3-asr:latest. - Set explicit Docker Compose project name
qwen3-asr, so generated networks no longer depend on the local checkout directory name. - Kept
funasr==1.3.1,/ws/v1/asr/funasr, andFUNASR_*settings where they still represent real dependency, protocol, or runtime compatibility semantics. - Preserved the current runtime split: CUDA uses official vLLM, CPU/macOS uses the vendored QwenASR Rust backend, and Paraformer remains the realtime-only websocket capability.
Docker Images
docker.io/quantatrisk/qwen3-asr:latestdocker.io/quantatrisk/qwen3-asr:gpu-latestdocker.io/quantatrisk/qwen3-asr:cpu-latest
Upgrade Notes
- Update any deployment scripts, compose files, CI variables, and pull commands from
funasr-apitoqwen3-asr. - If you rely on the legacy websocket compatibility endpoint,
/ws/v1/asr/funasris still intentionally available. - If you configured
FUNASR_*environment variables, keep them unchanged for this release.
中文
Qwen3-ASR v1.0.0 是项目从 funasr-api 改名为 qwen3-asr 后的首个正式发布版本。当前服务已经以 Qwen3-ASR 为核心,同时保留仍具备真实运行语义的 FunASR/Paraformer 实时 WebSocket 兼容路径。
主要变化
- 已将项目元数据、Docker Compose 服务、Docker 镜像名、默认日志、启动界面、benchmark 报告和核心文档统一改名为
qwen3-asr。 - 已更新 GitHub Actions 镜像发布配置,后续会推送到
docker.io/quantatrisk/qwen3-asr:cpu-latest、docker.io/quantatrisk/qwen3-asr:gpu-latest和docker.io/quantatrisk/qwen3-asr:latest。 - 已显式设置 Docker Compose project name 为
qwen3-asr,生成的网络名不再依赖本地目录名。 - 保留
funasr==1.3.1、/ws/v1/asr/funasr和FUNASR_*配置名,因为它们仍然分别对应真实依赖、兼容协议和运行时配置语义。 - 当前运行时拆分保持不变:CUDA 使用官方 vLLM,CPU/macOS 使用 vendored QwenASR Rust backend,Paraformer 继续作为仅实时流式能力存在。
Docker 镜像
docker.io/quantatrisk/qwen3-asr:latestdocker.io/quantatrisk/qwen3-asr:gpu-latestdocker.io/quantatrisk/qwen3-asr:cpu-latest
升级提示
- 请将部署脚本、compose 文件、CI 变量和镜像拉取命令中的
funasr-api更新为qwen3-asr。 - 如果仍依赖旧的 WebSocket 兼容端点,
/ws/v1/asr/funasr本版本仍会保留。 - 如果已经配置了
FUNASR_*环境变量,本版本不需要改名。