Skip to content

Releases: second-state/VoiceActions

Release list

v0.1.0

Choose a tag to compare

@juntao juntao released this 02 Mar 04:20

Voice Actions v0.1.0

Voice-to-voice processing pipeline: Audio → ASR (Qwen3) → WASM chain → TTS (Qwen3) → MP3 Output

Features

  • Speech-to-text using Qwen3-ASR (supports 30+ languages)
  • Text-to-speech using Qwen3-TTS (named speakers: Vivian, Ryan, etc.)
  • Chainable WASM processing modules (echo, LLM, custom)
  • MP3 output (192 kbps via libmp3lame)
  • Multi-platform: Linux x86_64/ARM64 (libtorch), macOS Apple Silicon (MLX Metal GPU)
  • Self-contained: all shared libraries bundled (WasmEdge, libopus, libmp3lame)

Release Archives

Archive Platform Backend
voice-actions-linux-x86_64.tar.gz Linux x86_64 libtorch CPU
voice-actions-linux-aarch64.tar.gz Linux ARM64 libtorch CPU
voice-actions-macos-aarch64.zip macOS Apple Silicon MLX Metal GPU

Quick Start

# Extract
tar xzf voice-actions-linux-x86_64.tar.gz
cd voice-actions-linux-x86_64

# Download models
pip install huggingface_hub transformers
huggingface-cli download Qwen/Qwen3-ASR-0.6B --local-dir models/Qwen3-ASR-0.6B
huggingface-cli download Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice --local-dir models/Qwen3-TTS-12Hz-0.6B-CustomVoice

# Generate tokenizer.json
python3 -c "
from transformers import AutoTokenizer
for model in ['Qwen3-ASR-0.6B', 'Qwen3-TTS-12Hz-0.6B-CustomVoice']:
    path = f'models/{model}'
    tok = AutoTokenizer.from_pretrained(path, trust_remote_code=True)
    tok.backend_tokenizer.save(f'{path}/tokenizer.json')
"

# Run with bundled echo module
./voice-actions \
  --input recording.mp3 \
  --output response.mp3 \
  --asr-model ./models/Qwen3-ASR-0.6B \
  --tts-model ./models/Qwen3-TTS-12Hz-0.6B-CustomVoice \
  --wasm echo.wasm