Prebuilt self-contained binaries for the FunASR llama.cpp / GGUF runtime, with SenseVoice, Paraformer, and Fun-ASR-Nano support plus built-in FSMN-VAD.
This release moves the SenseVoice Latest Release source snapshot onto current main, including the yaml.SafeLoader security fix from #312. Users downloading GitHub's generated source archives no longer receive the older unsafe YAML loader from runtime-llamacpp-v0.1.4.
Choose an asset
- Default
linux-x64/windows-x64: maximum CPU compatibility. linux-x64-avx2/windows-x64-avx2: higher throughput on CPUs with AVX2, FMA, F16C, and BMI2.linux-x64-vulkan/windows-x64-vulkan: Vulkan backend; requires a working Vulkan driver and ICD.windows-x64-cuda: CUDA backend targeting architecture 86; build from source for other GPU architectures.- Native
linux-arm64andmacos-arm64packages are included.
Download a default quantized model with:
pip install -U huggingface_hub
bash download-funasr-model.sh sensevoiceThen run llama-funasr-sensevoice; no Python ASR runtime or local compilation is required.
All nine files are byte-for-byte identical to the digest-verified assets from modelscope/FunASR runtime v0.1.9.
Documentation: https://github.com/QwenAudio/SenseVoice/blob/runtime-llamacpp-v0.1.9/runtime/llama.cpp/README.md