Skip to content

v0.2.12

Latest

Choose a tag to compare

@andimarafioti andimarafioti released this 05 Aug 13:03
56dc28f

speech-to-speech 0.2.12 is the final planned release in the 0.2.x line before
the next round of larger changes. It brings smarter turn-taking, WebRTC support,
direct audio input for audio-capable LLMs, more complete OpenAI Realtime protocol
behavior, and a significantly improved browser demo.

Highlights

  • Smarter endpointing with Smart Turn v3.2. Realtime mode now enables the
    quantized CPU Smart Turn model by default to distinguish completed turns from
    mid-thought pauses while speculative STT and LLM work continues. Use
    --no_smart_turn to retain Silero-only endpointing. (#192)
  • WebRTC transport for the OpenAI Realtime API. Install the new webrtc
    extra to use SDP negotiation at POST /v1/realtime/calls, RTP audio, and the
    oai-events data channel alongside the existing WebSocket transport. Both
    transports share the same pipeline pool, event dispatch, cancellation, and
    interruption behavior. (#352)
  • Direct audio input for audio-capable LLMs. Run with --stt none, the
    Chat Completions backend, and an explicitly selected audio-capable model to
    send completed VAD audio directly to the model. Audio history, cancellation,
    tool calls, usage accounting, and provider failures are handled transactionally.
    (#298)
  • Optional remote-LLM proxy. --enable_llm_proxy exposes the configured
    remote backend through the matching /v1/responses or
    /v1/chat/completions path, with streaming passthrough and usage accounting.
    The engine deliberately does not authenticate or throttle these routes, so
    standalone deployments must keep them on a trusted network or behind an
    authenticated gateway. (#368)

Realtime, demo, and backend improvements

  • The browser demo can select microphone and speaker devices, start with a
    configurable assistant greeting, forward signed-in Hugging Face identity for
    session allocation, show user-turn lifecycle feedback, and replay the exact
    captured user audio locally. (#376,
    #389,
    #390)
  • Qwen3-TTS GGML users can select quantization, load local talker and codec GGUF
    files, configure voice-reference caching, and reuse precomputed .spk and
    .rvq references. (#403)
  • Successful session.update requests now receive the expected
    session.updated event with the effective session configuration.
    (#413,
    #417,
    #418)
  • Language prompting now covers every language reported by the bundled STT
    handlers, including the full default Parakeet TDT language set.
    (#395)
  • A new guide documents a local Gemma 4 12B Realtime setup on Apple Silicon,
    including native-audio bypass, memory tuning, cancellation, barge-in, and
    troubleshooting. (#422)

Reliability and security

  • Session teardown now preserves drain sentinels, releases capacity after setup
    failures, quarantines stuck pipeline units, reports them through /v1/pool,
    and prevents stale teardown signals from releasing the wrong session.
    (#358)
  • Speculative turn tracking no longer resurrects untracked revisions after a
    reset or LRU eviction, preventing reused turn IDs from suppressing responses.
    (#391)
  • Both MLX Audio Whisper and Lightning Whisper MLX now use the global MLX lock,
    preventing Metal command-buffer crashes when STT overlaps other MLX work on
    Apple Silicon. (#379,
    #387)
  • Optional handler imports no longer reconfigure the host application's root
    logger, and setup/cleanup output now uses module loggers.
    (#409)
  • NLTK is upgraded to 3.10.0 to address CVE-2026-54293 / GHSA-p4gq-832x-fm9v.
    (#400)

Breaking change

The raw PCM WebSocket mode is now named raw-websocket. Commands using
--mode websocket must change to --mode raw-websocket; the old value is no
longer accepted. Realtime mode is unchanged and still supports both WebSocket
and WebRTC transports. (#401)

Packaging notes

  • Smart Turn adds huggingface-hub and onnxruntime as standard dependencies.
  • WebRTC support is available through pip install "speech-to-speech[webrtc]".
  • The Smart Turn model is downloaded from the Hugging Face Hub on first use
    unless a local model path is supplied.

All merged changes

New contributors

Thank you to @dcolley, @Hotragn, @varunsahni18, @ztcools,
@salignatmoandal, and @hrqiang for their first contributions to the project.

Full changelog: v0.2.11...v0.2.12