-
Notifications
You must be signed in to change notification settings - Fork 0
Voice
As of 3.19.0 your companion has a real voice. Hold the push-to-talk key (default V) and speak — your words are transcribed with Whisper large-v3-turbo and delivered to your own companion only. Everything the agent says is also spoken out loud with a natural text-to-speech voice — privately: by default only you hear your agent, and nobody else hears it unless you choose to share.
- Stand in the world (no menu open) and hold V. The action bar shows
🎤 Listening…. - Say what you want — "follow me", "stay here", "go chop those trees and build a little hut" — and release the key.
- The recording is transcribed on your own machine's connection (audio never goes to the game server — only the resulting text does) and handled exactly like a chat message addressed to your bot by name: quick intents (come / follow / stay / stop / where are you) are instant and free, anything else becomes an AI task.
Voice input is always private: it never appears in public chat, and it can
only ever command your own companion — not a friend's, not a trusted bot.
You'll see a private You → Ethan 🎤 "…" line confirming what was heard.
-
Transcription engine: with an API key set, Whisper large-v3-turbo
(
sttModel) on HuggingFace's inference API (sttApiUrl) — the same key you already use for the chat models, and a free HF token is enough. With no key it tries the free voice service instead — but see the honesty note below: the free service currently has no speech model, so for now voice wants a key (the text AI stays fully keyless). -
Rebind the key with
/aivoice key <GLFW code>(86 = V, 66 = B, 71 = G). The key is only read while no GUI is open, so typing "v" in chat is safe.
Every line your companion says in chat is also synthesized and played through your speakers, in its own voice. This is client-side and private — the audio plays only for players the server addressed, never for bystanders.
- Turn it off/on for yourself:
/aivoice off//aivoice on. - Interrupt a long line:
/aivoice stop. -
Pick your default voice:
/aivoice voice <id>(alloy,echo,fable,onyx,nova,shimmer,coral, …). -
Give a specific bot its own voice: stand near it and
/ai voice set <id>— per-bot and saved with the bot, so linked companions sound like different people. - Hear a sample:
/aivoice test Hello! I'm your companion.
Synthesis sends an OpenAI-style audio-out chat request to the freeApiUrl
endpoint, asking for WAV audio the game can play natively. Current status: the
free service dropped its audio model (verified live, July 2026 — it now serves
text only), so agent speech stays silent until either that service restores audio
or freeApiUrl points at an audio-capable OpenAI-compatible endpoint (a local
server, or a keyed voice service). When synthesis 404s, Blockpal logs one clear
explanation and pauses TTS attempts for 10 minutes instead of spamming.
By default you only hear your own agent. Sharing changes that:
| Command | Effect |
|---|---|
/ai voice |
Status: server voice on/off, your bot's voice, how it all works |
/ai voice share <player> |
Let an online player hear your agent too |
/ai voice unshare <player> |
Stop sharing with them |
/ai voice clear |
Stop sharing with everyone — private again |
/ai voice list |
Who hears your agent, and whose agents you can hear |
/ai voice set <id> |
Give your nearby bot its own TTS voice |
Sharing works in both directions independently: you hear a friend's agent only if they share with you.
When agents are shared and linked (a share in either direction links their owners into one conversation group), Blockpal coordinates their speech like people in a real conversation:
- All linked agents' lines go through one turn-taking queue on the server — while one agent is speaking, the others wait their turn instead of talking over it.
- Speaking time is estimated from the line's length at a natural speaking rate, with a cap so one rambling agent can't hog the conversation.
- Your own client additionally plays incoming speech strictly one utterance at a time — a second guarantee that voices never overlap on your machine.
Unlinked agents (different groups) are independent — your private conversation doesn't wait for someone else's.
The whole voice layer is server-gated:
-
/ai admin voice on|off— enable/disable voice for everyone (text command, works from Bedrock/vanilla clients and the console). - The same toggle lives in the panel: Settings → Behavior → "Allow agent voice".
- When off, the server never sends agent speech and refuses push-to-talk input with a friendly message.
- The FPS kill-switch and
/ai admin disablealso silence voice (no bots = no speech), and voice shares are in-memory, resetting on server restart (like parties).
| Setting | Default | Meaning |
|---|---|---|
allowVoice |
true |
(server) master gate for the whole voice layer |
voiceResponses |
true |
Speak the agent's replies out loud on this client |
voicePushToTalkKey |
86 (V) |
Push-to-talk key as a GLFW key code |
sttApiUrl |
HF Whisper endpoint | Where push-to-talk audio is transcribed |
sttModel |
openai/whisper-large-v3-turbo |
The transcription model |
ttsVoice |
alloy |
Default TTS voice when a bot has none of its own |
- Audio never reaches the game server. The mic recording is transcribed from your own machine straight to the transcription API; the server receives only the final text (same as if you'd typed it).
- The mic is only open while the key is held (hard 30-second cap), and nothing is recorded in menus.
- Voice output/transcription are network services — on an outage the agent simply stays text-only; chat and commands are unaffected. Failures show an accurate action-bar reason (service down vs. key rejected vs. too quiet).
-
Verified live (July 2026): the HF chat endpoint, the Whisper
large-v3-turbo route (model live on hf-inference, raw-WAV request +
{"text"}response per the official docs), and the free text fallback all answer as the code expects. The free service's audio model has been removed upstream — that's a service change, not a mod bug, and voice input works with a free HF key. - Like other recent features, the live audio path (microphone capture, the Whisper/TTS endpoints, playback) needs real-machine testing — it can't run in CI. If your setup has no microphone, push-to-talk fails with a clear message and everything else keeps working.
Blockpal — a Minecraft Fabric AI companion mod · default companion Ethan · Report an issue · MIT License
Start here
Using the mod
- Commands
- Talking to Your Assistant
- Client-Side Assistant
- Possession Mode
- Voice
- Settings
- Performance Presets
- Custom Skins
- Personalities
- Trust & Per-Bot Management
- Running Commands
- AI Actions
- Vision & Code
- PVT — learning by watching
- Speed & Combat
- Bedrock (Geyser)
- Bedrock Add-On (single-player)
Multiplayer & games
AI providers
- MCP server (Claude/ChatGPT/Grok/Gemini)
- Local AI on your GPU
- Local & Player2 AI
- Per-Player Keys & Models
Admin & safety
Reference