Releases: 927tanmay/react-ai-voice-avatar
Release list
v0.6.0 — voice mode on its own, cloud speech without the download, models kept between visits
Voice mode on its own becomes a first-class way to use the package: the
/headless hook, a page and an example for it, and an engine that no longer
charges a download for the setups production apps actually choose. Models are
kept between visits, so a returning visitor starts in seconds. The avatar
gestures while it talks and stands straight, speak() starts on its first
sentence, and an interruption, spoken or typed, now stops the voice cleanly.
No breaking changes. New: preloadLocalSpeech, onLocalSpeechReady,
isLocalSpeechReady, activeTtsEngine and onTtsEngineChange. One
behaviour change worth knowing: supplying onTranscribe or onSynthesize now
skips the local model it replaces instead of downloading it anyway.
Typing works on every page, and cuts the avatar off cleanly
- The homepage and
/voicehave a type box, as the scenario pages
already did. People who will not open a microphone for a page they just
found could watch the demo but never hear it answer. The reply is spoken
either way, and typing during a reply stops it first, like talking over it. - A sentence already being synthesised no longer plays after an
interrupt. Clearing a worker's queue does not stop the sentence it is
working on, and its audio arrived a second later. Talking over the avatar
hid this, because the late audio landed while the user was still speaking;
typed input starts the next turn at once, and the greeting's next sentence
played over the answer. Requests now carry a turn number that the workers
echo with their audio,interrupt()advances it, and audio from an older
turn is dropped. Both voices are covered: Kokoro, and MMS with the
in-browser model's replies.
Start on a provider, hand over to the browser
preloadLocalSpeechdownloads the in-browser hearing and voice behind
onTranscribeandonSynthesizewithout holding anything up, and
onLocalSpeechReadyfires when they could take over. Drop the adapters
then and the conversation carries on locally.isLocalSpeechReadyis on the
hook's return value. Measured with stored models: ready on the adapters in
0.2 s, local speech ready at 6 s, and the next reply spoken by Kokoro with no
call to the host's voice.examples/groq-voicedoes this on Groq's free tier with the developer's
own key, kept in a dev-server route: Whisper, GPT-OSS 20B and Orpheus to
start, then an offer to switch hearing and voice to the browser once they
have downloaded. It shows Groq's limits and what is left of them, and when
one runs out it says so and offers the in-browser models again.
The voice a phone gets is named, and kept for the right reasons
- Kokoro's download progress was wrong. It read 99% within two seconds
and stayed there for the whole 325 MB download: its three config files
finish before the model file starts, and the per-file estimate took those
3.6 KB for all of it. Only weight files count towards that estimate now. - Leaving during the download no longer counts as a crash. The
breadcrumb that switches a browser to the MMS voice after two failed Kokoro
loads was set as soon as loading began, so two abandoned downloads — easy on
a phone — downgraded the voice for good. It is now set once the download is
complete and the model is being built, which is where memory runs out. activeTtsEngineandonTtsEngineChangesay which voice is speaking:
'kokoro','mms'(every iPhone and iPad, and wherever Kokoro fails), or
'custom'foronSynthesize. MMS is a single plainer voice that ignores
ttsVoice, which is why a phone can sound like someone else. The demo now
names the voice under the button.
Cloud speech adapters no longer download the local models
onTranscribe and onSynthesize replaced Whisper and Kokoro at runtime, but
both were still downloaded — some 590 MB — and the engine stayed 'loading'
until they arrived, for models it never called. That is the configuration most
production apps choose.
- A model the host replaces is not loaded. With
onTranscribe,
onSubmitandonSynthesizeall supplied, the engine is ready in about two
seconds having requested no model at all. Each adapter skips only its own
model, soonSynthesizealone still loads Whisper. - Dropping an adapter later loads what it stood in for, so a host can fall
back from a cloud provider to the local model mid-session. - Kokoro's crash detector ignores hosts that synthesise. It counts loads
that started Kokoro and never finished it, and with Kokoro skipped, two
visits would have been read as two crashes and switched that browser to the
MMS voice for good. tests/cloud-adapters.spec.tsruns a whole turn through all three adapters
and fails on any model request. It fails against the previous engine.
Voice only, on the demo and as an example
Most apps want voice mode, not a character: talk to the app the way you talk to
ChatGPT or Gemini. The package has supported that since the
react-ai-voice-avatar/headless entry point, but nothing showed it. The one
example named "headless" still rendered the 3D avatar.
- The demo has a "3D avatar | Voice only" switch. Voice only is an orb that
moves with whoever is speaking, with live captions, running on
useAiVoiceAvatarfrom/headlessand nothing else. It never requests the
avatar file. Switching reloads the engine from the models the browser already
keeps, so nothing downloads twice, and the hosted conversation carries across. - It has a page of its own, at
/voice, linked from the demo's header: a
link to share for anyone who only wants voice mode. It is a separate entry
rather than a route, so it loads React and the hook and nothing else — 0.23
MB against the homepage's 10.3 MB with its two avatars. On a 10 Mbit/s line
it is usable in 0.2 s, where the homepage takes 8–9 s to show its avatar,
and with the models already stored it speaks 2–3 s sooner after the click. examples/voice-onlyis the same thing as an app to copy, with no
backend: the in-browser language model answers unless you passonSubmit.
Its production build contains no WebGL code.
The avatars stand symmetrically
Both bundled avatars stood with the right hand about 8 cm further from the body
than the left, and the right fingertips about 13 cm off where the left ones
were. scripts/convert-rocketbox.py bent every elbow and knuckle 10° about the
bone's own X axis on both sides. On a mirrored skeleton that axis is reflected,
so the same angle bent the two sides opposite ways: the left elbow in toward
the body, the right one out. Measured on the files, every right bone from the
forearm down sat exactly 20° off its mirror.
- Published avatars corrected in place.
scripts/mirror-arm-rest-pose.mjs
rewrites the 16 right-arm rotations (forearm and finger joints) as the mirror
of the left. Nothing else in either file changed: mesh, skin, morph targets
and textures are byte-identical. - Why the left arm is the one kept: a relaxed arm with the palm facing the
thigh brings the hand forward and slightly in, which is what the left does. - Served from a new tag. The avatars are served from a new tag,
avatars-v2. Installs pinned toavatars-v1keep the avatar they were built
against. - The converter is fixed too. It now turns the right side the opposite way.
That change was worked out from the measurements, not run: the Rocketbox
source files were not at hand. scripts/test-avatar-symmetry.mjsguards it. It compares every left
bone with its right twin, checks both elbows bend in rather than out, and
refuses a rig whose upper arms are not mirrored rather than "fixing" it. It
fails on the old files, including when they are mirrored the wrong way round.
speak() starts talking on its first sentence
speak(), and a reply an onSubmit handler returns as a single string, went to
the voice whole, and the voice produces audio for everything it is given before
handing any back. So nothing was audible until the last sentence existed. On a
machine without WebGPU, that was the demo's greeting: 30 to 33 seconds of
silence straight after a long download, measured over four loads, while the
status read speaking from the first millisecond.
Text is now cut into sentences (src/lib/speechChunks.ts) and the voice returns
each as it is ready, so the first plays while the rest are being made. Same
greeting, same machine: audible after about 2 seconds, and finished after about
24 rather than 46, because the voice works on the next sentence while the
current one plays.
What that machine still costs: its voice runs about 2.5× slower than real
time, so on it the sentences have pauses between them, the longest about 5 s
after the opening "Hi!". A pause appears only where the next sentence takes
longer to make than the current one takes to say. With WebGPU, on a Mac, none
did: the greeting was audible after 0.75 s and each sentence started within
2 ms of the previous one ending.
speak()now reportsthinkinguntil its audio starts, thenspeaking,
like every other reply. A host showingisSpeakingno longer shows it over
silence. Callingspeak()while the avatar is already talking still stays
speaking.- Transcripts are unchanged:
onTranscriptUpdatestill receives the text
once, whole. - An
onSynthesizeadapter now receives sentences rather than the whole
text, as it already did for streamed replies. They are requested one after
another, so they cannot finish out of order, and each is fetched while the one
before it plays. MMS and cloud voices keep the 35-character minimum piece they
have always had on the streaming path. - Splitting avoids the obvious traps: it doesn't cut inside "0.5B",
"3.5 GB", "Dr." or a domain name, and it splits Hindi at the danda. A sentence
over 200 characters is cut at...
v0.3.0 — interruptible voice, turn-taking that ignores coughs, error reporting
The release that makes the voice loop good enough to put in front of real
users. Everything below 0.2.1 worked in a demo; this is the pass that fixed
what only shows up when someone actually talks to it.
The user can interrupt the avatar
Talking over the avatar now cuts it off mid-sentence, on by default. Waiting
for a reply to finish is the thing that makes a voice agent feel like a
walkie-talkie.
allowInterruptionturns it off for a kiosk or a noisy room, where the
avatar hearing itself through the speakers is worse than waiting.onUserInterruptfires when the user takes the floor.speechDetectionexposes every threshold the microphone uses to decide
someone is talking, because the right values depend on the room.
Turns are decided on sustain, not loudness. Voice detection scores every
96ms frame; a cough clears a loudness bar as easily as a word does. A sound now
has to keep scoring as speech before it counts as a turn. Below that bar it
only ducks the avatar's voice, reversibly — cough during a reply and the reply
resumes from where it paused, at the right place in the sentence and with the
mouth still in sync.
The conversation has one set of rules
The status used to be assigned from seventeen places. Each assignment was
locally correct and the set of them was not, because nothing said which
transitions were legal and the events do not arrive in the order the code is
written in. It is now a single transition table (src/lib/turnState.ts),
testable without a browser, a microphone or a model download.
Hindi
Speech, transcription and replies, via a Devanagari phonemiser and Kokoro's
undocumented Hindi voices. Code-switching works: English words inside a Hindi
sentence are routed to the English phonemiser.
Local Hindi is a demo, not a product, and the README says so. Models small
enough to run in a browser are far weaker in Hindi than in English. For
production Hindi, route hearing and thinking to an API; the voice and lip-sync
stay local and are genuinely good.
Failures are reportable
onError gives a host application a stage, a message and a severity of
degraded or fatal, so a survivable fallback can be told apart from a dead
microphone rather than arriving as a console message nobody can see.
Fixes
- Workers abandoned during startup no longer hijack the pipeline. Spawning
is asynchronous, so a React StrictMode double-mount left an orphan worker
that finished loading and then claimed the reference — after which every
reply it sent was ignored, and requests vanished into it silently. - The microphone no longer goes stale. Voice detection stops the stream's
tracks when paused and reacquires privately on resume, which left the engine
holding a dead device: the level meter read a stopped microphone for the rest
of the session. Hitpush-to-talkandallowInterruption: falseevery turn. - Speech is scheduled against the audio clock, removing the audible seam
between chunks that main-thread jitter used to put in the middle of sentences. - The microphone opens on first use, not on mount. A visitor is no longer
asked for permission before clicking anything. - Stopped probing for a local avatar file in applications that have none, which
put a 404 in every Next.js dev console. - Avatars are pinned to a release tag, so a merge cannot change them underneath
a deployed site. - Replaced the avatar meshes with MIT-licensed Microsoft Rocketbox models,
relaxed the rest pose so the arms stop dangling, and relit for skin.
API
- Added:
allowInterruption,speechDetection,onUserInterrupt,onError,
AiVoiceAvatarError,AiVoiceAvatarErrorStage. - Added the
react-ai-voice-avatar/headlessentry point — the full
conversational loop with no three.js in the module graph. 117 kB first-load
against 389 kB for the 3D component, measured on the same Next.js build. - Changed:
onAudioLevelChange'ssourceis typed'mic' | 'tts' | 'idle'.
It has always emitted'idle'between turns; the type said otherwise. No
runtime change, but an exhaustiveswitchwill now want the third case.
Known limits
- Raising
positiveSpeechThresholdtrades quiet speakers for quiet rooms. A
speaker below it produces no events at all. - An "hmm" sustained past
minSpeechMsstill reaches transcription. A denylist
catches the common ones after the fact. - A session is one language at a time. Browser Whisper cannot detect language,
so someone who switches mid-conversation will be mistranscribed. @react-three/fiber@9.7caps React below 19.3. Install with
--legacy-peer-deps; the combination works, the peer range is over-cautious.
npm install react-ai-voice-avatar@0.3.0