speech-to-speech 0.2.12 is the final planned release in the 0.2.x line before
the next round of larger changes. It brings smarter turn-taking, WebRTC support,
direct audio input for audio-capable LLMs, more complete OpenAI Realtime protocol
behavior, and a significantly improved browser demo.
Highlights
- Smarter endpointing with Smart Turn v3.2. Realtime mode now enables the
quantized CPU Smart Turn model by default to distinguish completed turns from
mid-thought pauses while speculative STT and LLM work continues. Use
--no_smart_turnto retain Silero-only endpointing. (#192) - WebRTC transport for the OpenAI Realtime API. Install the new
webrtc
extra to use SDP negotiation atPOST /v1/realtime/calls, RTP audio, and the
oai-eventsdata channel alongside the existing WebSocket transport. Both
transports share the same pipeline pool, event dispatch, cancellation, and
interruption behavior. (#352) - Direct audio input for audio-capable LLMs. Run with
--stt none, the
Chat Completions backend, and an explicitly selected audio-capable model to
send completed VAD audio directly to the model. Audio history, cancellation,
tool calls, usage accounting, and provider failures are handled transactionally.
(#298) - Optional remote-LLM proxy.
--enable_llm_proxyexposes the configured
remote backend through the matching/v1/responsesor
/v1/chat/completionspath, with streaming passthrough and usage accounting.
The engine deliberately does not authenticate or throttle these routes, so
standalone deployments must keep them on a trusted network or behind an
authenticated gateway. (#368)
Realtime, demo, and backend improvements
- The browser demo can select microphone and speaker devices, start with a
configurable assistant greeting, forward signed-in Hugging Face identity for
session allocation, show user-turn lifecycle feedback, and replay the exact
captured user audio locally. (#376,
#389,
#390) - Qwen3-TTS GGML users can select quantization, load local talker and codec GGUF
files, configure voice-reference caching, and reuse precomputed.spkand
.rvqreferences. (#403) - Successful
session.updaterequests now receive the expected
session.updatedevent with the effective session configuration.
(#413,
#417,
#418) - Language prompting now covers every language reported by the bundled STT
handlers, including the full default Parakeet TDT language set.
(#395) - A new guide documents a local Gemma 4 12B Realtime setup on Apple Silicon,
including native-audio bypass, memory tuning, cancellation, barge-in, and
troubleshooting. (#422)
Reliability and security
- Session teardown now preserves drain sentinels, releases capacity after setup
failures, quarantines stuck pipeline units, reports them through/v1/pool,
and prevents stale teardown signals from releasing the wrong session.
(#358) - Speculative turn tracking no longer resurrects untracked revisions after a
reset or LRU eviction, preventing reused turn IDs from suppressing responses.
(#391) - Both MLX Audio Whisper and Lightning Whisper MLX now use the global MLX lock,
preventing Metal command-buffer crashes when STT overlaps other MLX work on
Apple Silicon. (#379,
#387) - Optional handler imports no longer reconfigure the host application's root
logger, and setup/cleanup output now uses module loggers.
(#409) - NLTK is upgraded to 3.10.0 to address CVE-2026-54293 / GHSA-p4gq-832x-fm9v.
(#400)
Breaking change
The raw PCM WebSocket mode is now named raw-websocket. Commands using
--mode websocket must change to --mode raw-websocket; the old value is no
longer accepted. Realtime mode is unchanged and still supports both WebSocket
and WebRTC transports. (#401)
Packaging notes
- Smart Turn adds
huggingface-hubandonnxruntimeas standard dependencies. - WebRTC support is available through
pip install "speech-to-speech[webrtc]". - The Smart Turn model is downloaded from the Hugging Face Hub on first use
unless a local model path is supplied.
All merged changes
- Fix stuck pipeline units in session teardown by @A-Mahla in #358
- Add WebRTC transport for the OpenAI Realtime API by @A-Mahla in #352
- Bump the actions group with 2 updates by @dependabot in #364
- Add mic/speaker device selectors to the demo Settings UI by @dcolley in #376
- Acquire the global MLX lock in MLXAudioWhisperSTTHandler by @Hotragn in #379
- Add demo startup greeting and HF user attribution by @andimarafioti in #389
- Update star history chart by @andimarafioti in #393
- Add an LLM proxy for remote LLM backends by @A-Mahla in #368
- Add GitHub Trending achievement badge by @andimarafioti in #398
- Acquire the MLX lock in LightningWhisperSTTHandler by @varunsahni18 in #387
- Stop SpeculativeTurnTracker from committing untracked turns by @ztcools in #391
- Give every STT language an LLM language name by @Hotragn in #395
- Refine the GitHub Trending badge by @andimarafioti in #399
- Upgrade NLTK to 3.10.0 by @andimarafioti in #400
- Rename raw PCM mode to
raw-websocketby @andimarafioti in #401 - Add user audio replay to the Realtime demo by @andimarafioti in #390
- Add missing Qwen3-TTS GGML features by @andimarafioti in #403
- Avoid configuring the root logger when importing optional handlers by @salignatmoandal in #409
- Add audio-input LLM STT bypass by @andimarafioti in #298
- Send
session.updatedaftersession.updateby @hrqiang in #413 - Fix the
session.updatedevent contract by @andimarafioti in #417 - Integrate the
session.updatedwork by @andimarafioti in #418 - Refresh the star history chart by @andimarafioti in #419
- Document the Gemma 4 12B Realtime setup by @andimarafioti in #422
- Bump the actions group with 3 updates by @dependabot in #406
- Integrate Smart Turn v3.2 endpointing by @andimarafioti in #192
New contributors
Thank you to @dcolley, @Hotragn, @varunsahni18, @ztcools,
@salignatmoandal, and @hrqiang for their first contributions to the project.
Full changelog: v0.2.11...v0.2.12