Skip to content

Releases: rootint/nexara-python-sdk

v0.7.0

Choose a tag to compare

@rootint rootint released this 15 Sep 19:16

Smart turn detection for realtime

client.realtime.connect(turn_detection="smart") tells a voice agent when the user starts speaking, stops, and has finished their turn. The session then yields three new event types alongside RealtimeTranscript:

  • RealtimeSpeechStarted — speech.started, ~100 ms after the user starts talking (interrupt the agent).
  • RealtimeSpeechEnded — speech.ended, ~280 ms after silence; not necessarily the end of the turn.
  • RealtimeTurnEnd — turn.end, the turn's complete text and words, with confidence (>= 0.5 model decided, < 0.5 3 s silence limit, None closed by finish()). Arrives after every transcript of its turn and before any of the next.

Every RealtimeTranscript gains turn: int | None. Works with any delay_ms, diarize and multichannel (turns are counted per channel).

Typing

RealtimeSession is now generic in its event type. connect() is overloaded: without turn_detection iteration still yields RealtimeTranscript, so existing typed code is unaffected; with turn_detection="smart" it yields the RealtimeEvent union — narrow on event.type.

New exports: RealtimeEvent, RealtimeSpeechStarted, RealtimeSpeechEnded, RealtimeTurnEnd, RealtimeTurnDetection.

v0.6.0

Choose a tag to compare

@rootint rootint released this 14 Sep 01:35

Realtime streaming

client.realtime.connect() now speaks the public streaming.nexara.ru WebSocket protocol (v1): GET /v1/transcribe with the session parameters in the query string and the API key as a Bearer token; raw PCM in, final words out.

async with client.realtime.connect(sample_rate=16000, diarize=True) as session:
    async for event in session.stream(microphone()):
        print(event.speaker, event.words[0].start, event.text)
    print(session.ended.text)
  • RealtimeSession: async iteration of RealtimeTranscript events, stream(), finish() (sends input_audio.end, resolves on session.ended), keepalive(), close(). Sending and receiving are concurrent.
  • Parameters: encoding (pcm_s16le default, pcm_f32le, pcm_mulaw, pcm_alaw), sample_rate (8–48 kHz), channels, multichannel, diarize, delay_ms, client_id; invalid values raise NexaraValidationError before connecting.
  • Every word is final, with integer-millisecond timestamps on the audio clock and, with diarize=True, a speaker label (0–3 or None).
  • New RealtimeError(code, reason, message) for server-side closes (session_full, idle_timeout, audio_backlog, service_restart, …); wallet exhaustion mid-session raises InsufficientBalanceError; a rejected handshake raises the usual status-mapped errors.
  • Requires the realtime extra: pip install nexara[realtime].

Breaking: the previous placeholder realtime API (RealtimeEvent, RealtimeToken, is_final) is gone — it never matched a real server.

v0.5.0

Choose a tag to compare

@rootint rootint released this 01 Aug 17:41

Mirrors apigateway 14.0: per-segment emotion recognition.

New

  • emotions=True on transcriptions.create() and create_job() (sync and async). Asks the ASR model to score each diarized segment.
  • Emotion — label, confidence, and probs when the server sends it. Exported from nexara, along with EMOTION_LABELS (angry, sad, neutral, positive).
  • DiarizedSegment.emotion — Emotion | None. Present only on segments the model could actually score, so check it per segment rather than per response. Never on words.
  • UsageItem.emotions — whether the call was charged the emotion surcharge.
call = client.transcriptions.create(
    file="call.mp3",
    task="diarize",
    model="nexara-ru",
    emotions=True,
)
for segment in call.segments:
    if segment.emotion:
        print(segment.speaker, segment.emotion.label, segment.emotion.confidence)

Server behavior this mirrors

Emotion is produced by the ASR model itself, so the server accepts emotions only with task="diarize", model="nexara-ru", and a JSON response format (json or verbose_json — the emotion object hangs off a segment, and text/srt/vtt have nowhere to put it). Every other combination is a 400. All three are now checked client-side and raise NexaraValidationError before the audio is uploaded, because the flag carries a per-second surcharge and a silently emotion-less paid response is the worst outcome available.

The format check runs after the prompt block that rewrites the format to verbose_json, matching the server's own order, so prompt= and emotions= still combine.

The server drops the surcharge when the backend scored nothing, so UsageItem.emotions means "emotion was delivered", not "emotion was requested".

Note on typing

Emotion.label is a str, not a Literal. Pydantic validates at runtime, so pinning the four labels would turn a label the server adds later into a hard parse failure for an entire diarization — losing the transcript over an advisory field. Compare against EMOTION_LABELS if you need exhaustiveness. (The TypeScript SDK uses a union for the same field, which is free there because TS types are erased.)

No breaking changes.

v0.4.0

Choose a tag to compare

@rootint rootint released this 31 Jul 13:33

...

v0.3.0

Choose a tag to compare

@rootint rootint released this 26 Jul 16:29

Realtime WebSocket streaming (nexara.realtime); requires the 'realtime' extra (websockets).

v0.2.0

Choose a tag to compare

@rootint rootint released this 23 Jul 15:27

Add typed SyncLLMTimeoutError (413) and BadGatewayError (502) for sync LLM enrichment on long audio; 413 is never retried — resubmit via create_job().

v0.1.0

Choose a tag to compare

@rootint rootint released this 20 Jul 16:07