Skip to content

v0.5.0

Choose a tag to compare

@rootint rootint released this 01 Aug 17:41
· 2 commits to main since this release

Mirrors apigateway 14.0: per-segment emotion recognition.

New

  • emotions=True on transcriptions.create() and create_job() (sync and async). Asks the ASR model to score each diarized segment.
  • Emotion — label, confidence, and probs when the server sends it. Exported from nexara, along with EMOTION_LABELS (angry, sad, neutral, positive).
  • DiarizedSegment.emotion — Emotion | None. Present only on segments the model could actually score, so check it per segment rather than per response. Never on words.
  • UsageItem.emotions — whether the call was charged the emotion surcharge.
call = client.transcriptions.create(
    file="call.mp3",
    task="diarize",
    model="nexara-ru",
    emotions=True,
)
for segment in call.segments:
    if segment.emotion:
        print(segment.speaker, segment.emotion.label, segment.emotion.confidence)

Server behavior this mirrors

Emotion is produced by the ASR model itself, so the server accepts emotions only with task="diarize", model="nexara-ru", and a JSON response format (json or verbose_json — the emotion object hangs off a segment, and text/srt/vtt have nowhere to put it). Every other combination is a 400. All three are now checked client-side and raise NexaraValidationError before the audio is uploaded, because the flag carries a per-second surcharge and a silently emotion-less paid response is the worst outcome available.

The format check runs after the prompt block that rewrites the format to verbose_json, matching the server's own order, so prompt= and emotions= still combine.

The server drops the surcharge when the backend scored nothing, so UsageItem.emotions means "emotion was delivered", not "emotion was requested".

Note on typing

Emotion.label is a str, not a Literal. Pydantic validates at runtime, so pinning the four labels would turn a label the server adds later into a hard parse failure for an entire diarization — losing the transcript over an advisory field. Compare against EMOTION_LABELS if you need exhaustiveness. (The TypeScript SDK uses a union for the same field, which is free there because TS types are erased.)

No breaking changes.