Voice Actions v0.1.0
Voice-to-voice processing pipeline: Audio → ASR (Qwen3) → WASM chain → TTS (Qwen3) → MP3 Output
Features
- Speech-to-text using Qwen3-ASR (supports 30+ languages)
- Text-to-speech using Qwen3-TTS (named speakers: Vivian, Ryan, etc.)
- Chainable WASM processing modules (echo, LLM, custom)
- MP3 output (192 kbps via libmp3lame)
- Multi-platform: Linux x86_64/ARM64 (libtorch), macOS Apple Silicon (MLX Metal GPU)
- Self-contained: all shared libraries bundled (WasmEdge, libopus, libmp3lame)
Release Archives
| Archive | Platform | Backend |
|---|---|---|
voice-actions-linux-x86_64.tar.gz |
Linux x86_64 | libtorch CPU |
voice-actions-linux-aarch64.tar.gz |
Linux ARM64 | libtorch CPU |
voice-actions-macos-aarch64.zip |
macOS Apple Silicon | MLX Metal GPU |
Quick Start
# Extract
tar xzf voice-actions-linux-x86_64.tar.gz
cd voice-actions-linux-x86_64
# Download models
pip install huggingface_hub transformers
huggingface-cli download Qwen/Qwen3-ASR-0.6B --local-dir models/Qwen3-ASR-0.6B
huggingface-cli download Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice --local-dir models/Qwen3-TTS-12Hz-0.6B-CustomVoice
# Generate tokenizer.json
python3 -c "
from transformers import AutoTokenizer
for model in ['Qwen3-ASR-0.6B', 'Qwen3-TTS-12Hz-0.6B-CustomVoice']:
path = f'models/{model}'
tok = AutoTokenizer.from_pretrained(path, trust_remote_code=True)
tok.backend_tokenizer.save(f'{path}/tokenizer.json')
"
# Run with bundled echo module
./voice-actions \
--input recording.mp3 \
--output response.mp3 \
--asr-model ./models/Qwen3-ASR-0.6B \
--tts-model ./models/Qwen3-TTS-12Hz-0.6B-CustomVoice \
--wasm echo.wasm