speech-to-speech v1.0.0
The first major release brings a single Realtime engine, explicit server and
microphone-client commands, more speech backends, and improved conversation
lifecycle handling. These notes cover changes since v0.2.12.
Highlights
- Three commands:
serve,talk, andlocal. Run the Realtime server,
connect the packaged microphone/speaker client, or start both in one terminal.
The client also supports local tool execution with--tool-module. - More speech backends. Add Qwen3-ASR speech recognition, OpenAI-compatible
STT and TTS endpoints, stateful OpenAI Realtime transcription, experimental
vLLM Realtime transcription, and optional OmniVoice and Supertonic TTS. - More reliable realtime conversations. Improve ordered text, audio, and
tool-call output; transcript and content-part events; interrupted responses;
remote-request cancellation; and session cleanup. The browser demo uses the
OpenAI Agents SDK, with integration tests for its WebSocket and WebRTC
transports against the implemented core Realtime event set. - Updated Apple Silicon setup. Refresh the MLX dependencies and default to
the 4-bitmlx-community/Qwen3-4B-Instruct-2507-4bitLLM with 6-bit Qwen3-TTS.
The Mac preset supplies defaults while respecting explicit overrides. - Clearer setup and diagnostics. The README provides local Mac, local
NVIDIA, and hosted-LLM starting configurations, a shared llama.cpp recipe,
offline setup guidance, and a speaker-feedback workaround. Full transcript
logging is now opt-in with--log_transcripts; operational logs omit
conversation text by default.
Upgrading from v0.2.12
In your Python environment:
pip install --upgrade "speech-to-speech==1.0.0"Update scripts and service definitions to use the explicit commands:
| Previous invocation | v1.0.0 invocation |
|---|---|
speech-to-speech |
speech-to-speech serve |
speech-to-speech --mode realtime |
speech-to-speech serve |
speech-to-speech --mode local |
speech-to-speech local |
speech-to-speech --local_mac_optimal_settings |
speech-to-speech local --mac-optimal-settings |
python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765 |
speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtime |
The --mode realtime and --mode local forms still work temporarily with a
deprecation warning. Other mode values, including socket and raw-websocket,
have been removed. Migrate raw PCM clients to the Realtime WebSocket or WebRTC
API; the old raw transport flags are no longer accepted.
The --mac-optimal-settings preset no longer selects a command: use local
for microphone/speaker interaction or serve for a server. serve binds to
127.0.0.1 by default; set --host explicitly for network access. local
uses loopback only.
The default hosted LLM is now gpt-5.6-terra, with reasoning effort none.
Pin --model_name if your deployment depends on a particular model. Local
Mac users can likewise override the preset's model with --model_name.
The redundant facebook-mms, language-detection, and websocket installation
extras were removed; their dependencies are included in the standard package.
Install speech-to-speech[webrtc] for WebRTC, speech-to-speech[omnivoice] for
OmniVoice, or speech-to-speech[supertonic] for Supertonic. Known options for
inactive backends are accepted but ignored with a warning.
Compatibility and setup
- Python 3.10+ is supported; the installation smoke tests use Python 3.11 on
Linux and Apple Silicon. Check the README's CUDA wheel guidance for Linux. - Realtime compatibility covers the documented core event set, not every
OpenAI Realtime API feature. See the
protocol reference. --local_audio_block_mic_during_playbackpauses microphone capture during
assistant playback to reduce speaker feedback. It disables spoken
interruptions during playback; it does not perform acoustic echo cancellation.
See the README for starting configurations and the
full changelog
for all changes and contributors.