Skip to content

Releases: Gjusev/voice-evals

voice-evals v0.2.1

Choose a tag to compare

@Gjusev Gjusev released this 04 Oct 20:24

Patch release

Package changes since v0.2.0:

  • The protocol map's declared audio.input.frame_ms is no longer dead configuration: it becomes the caller pacing default unless --frame-ms overrides it.
  • default-v1.json now declares outcome_events: false: transport-source outcome scenarios fail fast at validation instead of silently scoring task completion 0 against agents that never emit an outcome event. The inbound outcome mapping is kept for maps whose servers do emit it.
  • Tests resolve bundled resources from the installed package (importlib.resources), so the offline Kaggle suite validates the wheel itself.

Also in this release (repository-side): multilingual real-speech fixtures (ElevenLabs, en/de/es) and open-model fixtures (Kokoro-82M, en/es at native 24 kHz) with per-language measured TTS->STT round-trip WER, integrity tests, the local reference agent, Kaggle kernels published and verified (offline Kernel A full pass with internet disabled; live Kernel B mock path), and honest documentation of the Qwen3-TTS bf16 hardware blocker for German clips.

Verified before tagging: 125 offline tests, ruff clean, wheel-install smoke, live adapter/E2E/multilingual suites green.

voice-evals v0.2.0

Choose a tag to compare

@Gjusev Gjusev released this 04 Oct 17:26

What's new

The harness becomes a synthetic caller: TTS-generated caller lines from a scenario script (v2), a real WebSocket transport driven by declarative protocol maps, deliberate barge-ins, full session recording, and scoring with the unchanged v0.1 replay evaluator.

Highlights

  • Scenario v2: bounded deterministic scripts with conditional branches, interrupt cues, fact placeholders/reveals; selection is seeded and order-independent. JSON Schema published, unknown fields rejected.
  • Live probe: full-duplex paced sending, response correlation (provider IDs or serial), E2E/STT timing with explicit measurement bases, barge-in observation with right-censored "not stopped" results, honest NOT SCORED reporting with exclusion reasons.
  • Transports & callers: default-v1 protocol map + real WebSocket client (httpx/websockets behind the optional probe extra); ElevenLabs native PCM caller, generic HTTP open-model caller, checksummed fixture caller, and offline mocks.
  • Reproduction: internet-disabled Kaggle Kernel A (pinned wheelhouse, hashed locks, frozen 123-call regression corpus with hand-checked WER cases) and live-diagnostic Kernel B with a full mock fallback.
  • Compatibility: replay API, demo metrics, and exit codes unchanged; base install still depends only on jiwer.

Verified

116 offline tests (no sockets, no providers); live validation 2026-10-04: real ElevenLabs PCM synthesis (16/24 kHz), scored WebSocket E2E, and a full real probe (ElevenLabs → Scribe STT) with measured WER 0.1379, turn E2E p50 1086 ms, barge-in stop 232.6 ms, replay-identical scores.

v0.1.1

Choose a tag to compare

@Gjusev Gjusev released this 04 Oct 13:48