Repository navigation
Releases: Gjusev/voice-evals
Release list
voice-evals v0.2.1
Patch release
Package changes since v0.2.0:
- The protocol map's declared
audio.input.frame_msis no longer dead configuration: it becomes the caller pacing default unless--frame-msoverrides it. default-v1.jsonnow declaresoutcome_events: false: transport-source outcome scenarios fail fast at validation instead of silently scoring task completion 0 against agents that never emit an outcome event. The inbound outcome mapping is kept for maps whose servers do emit it.- Tests resolve bundled resources from the installed package (
importlib.resources), so the offline Kaggle suite validates the wheel itself.
Also in this release (repository-side): multilingual real-speech fixtures (ElevenLabs, en/de/es) and open-model fixtures (Kokoro-82M, en/es at native 24 kHz) with per-language measured TTS->STT round-trip WER, integrity tests, the local reference agent, Kaggle kernels published and verified (offline Kernel A full pass with internet disabled; live Kernel B mock path), and honest documentation of the Qwen3-TTS bf16 hardware blocker for German clips.
Verified before tagging: 125 offline tests, ruff clean, wheel-install smoke, live adapter/E2E/multilingual suites green.
voice-evals v0.2.0
What's new
The harness becomes a synthetic caller: TTS-generated caller lines from a scenario script (v2), a real WebSocket transport driven by declarative protocol maps, deliberate barge-ins, full session recording, and scoring with the unchanged v0.1 replay evaluator.
Highlights
- Scenario v2: bounded deterministic scripts with conditional branches, interrupt cues, fact placeholders/reveals; selection is seeded and order-independent. JSON Schema published, unknown fields rejected.
- Live probe: full-duplex paced sending, response correlation (provider IDs or serial), E2E/STT timing with explicit measurement bases, barge-in observation with right-censored "not stopped" results, honest NOT SCORED reporting with exclusion reasons.
- Transports & callers: default-v1 protocol map + real WebSocket client (httpx/websockets behind the optional
probeextra); ElevenLabs native PCM caller, generic HTTP open-model caller, checksummed fixture caller, and offline mocks. - Reproduction: internet-disabled Kaggle Kernel A (pinned wheelhouse, hashed locks, frozen 123-call regression corpus with hand-checked WER cases) and live-diagnostic Kernel B with a full mock fallback.
- Compatibility: replay API, demo metrics, and exit codes unchanged; base install still depends only on jiwer.
Verified
116 offline tests (no sockets, no providers); live validation 2026-10-04: real ElevenLabs PCM synthesis (16/24 kHz), scored WebSocket E2E, and a full real probe (ElevenLabs → Scribe STT) with measured WER 0.1379, turn E2E p50 1086 ms, barge-in stop 232.6 ms, replay-identical scores.
v0.1.1
Full Changelog: https://github.com/Gjusev/voice-evals/commits/v0.1.1