Skip to content

Speech and Video

Kenny Lasyone edited this page Oct 2, 2026 · 1 revision

Speech and video

Arcadia Studio can speak, listen and record video. Everything runs on your own computer: nothing you say, type or record is sent anywhere, and it all works offline.

  • Speech turns text into a voice. Use it to give characters their lines, add a narrator or a menu voice, or let your AI assistant talk you through things.
  • Speech recognition turns speech into text. Talk to your assistant instead of typing, or write out a recording or a video.
  • Recording makes a video of your game straight from Preview, with its sound.

Each one runs on its own thread, so the editor never stops while it speaks, listens or records.

Your assistant's voice

The first time you start MCP (or open Create audio from text), Arcadia Studio asks which voice your assistant should speak with. Press ▶ Try it to hear each one, then Use this voice. Decide later uses the standard voice and asks again next time.

There are over 50 voices in nine languages: American and British English, Spanish, French, Italian, Brazilian Portuguese, Hindi, Japanese and Mandarin. You can also mix two voices into a new one. Pick a second voice under Mix with, then slide between them to choose how much of each.

Change the voice any time in Advanced → Speech settings….

Create audio from text

Advanced → Create audio from text… turns text into speech.

  1. Type the text, or press Import text file….
  2. Choose a voice (or a mix) and a speed.
  3. Then use one of the buttons:
Button What it does
▶ Listen Speaks the text now. It starts as soon as the first sentence is ready (usually under half a second), while the rest is still being made. Press it again to stop.
Save as… Saves the speech as an Ogg, WAV or MP3 file.
Add to project… Adds it to the project's sounds under the name you give. It plays like any other sound: a Sound control, the Play sound action or ctx.client.playSound('myproject:guard_hello') in a script. It's saved as Ogg, so it plays in Minecraft too.

Sentences become short natural pauses, and blank lines become longer ones. To change how a word is said, spell it the way it sounds.

Create text from audio

Advanced → Create text from audio… writes down what was said.

  • Choose audio or video file… reads any recording or video: MP3, WAV, Ogg, M4A, FLAC, MP4, MKV, WebM, MOV and more.
  • 🎤 Record records from your microphone until you press Stop.
  • Language is detected by itself, or you can choose one.
  • Timestamps (subtitles) writes the text as SRT subtitles, each line with its time.

The text is editable. Copy it, Save as… a text or subtitle file, or press Add to screen as label or Add to screen as text box to put it on the open screen (one Undo step).

Talking to your assistant

The Ask Agent window has a 🎤 Speak button. Press it and say what you want. When you pause, what it heard is written into the box and asked straight away. Press it again to stop listening early.

Speech recognition already knows Arcadia Studio's own words (Toolbox, Preview, Properties) and the names of your screens and controls, so it gets them right.

Writing for the voice

The voice reads what it's given, so write the way it should sound. Plain sentences work best: leave out emoji, links and symbols, write abbreviations the way they're said ("W, A, S, D"), and spell an unusual name the way it sounds. Commas give a short pause and full stops a longer one. Curly apostrophes and quotes from word processors are fine. When an assistant's line has something that may be misread, it's told, so it can say it better.

Speech settings

Advanced → Speech settings…:

Setting What it does
Voice, Mix with, Speed, Volume Your assistant's voice. ▶ Try it plays a sample.
Read tutorial steps aloud When an assistant's tutorial has narration, each step is read out as it's shown. The tutorial window has its own 🔊 Read aloud box too.
Voice model Standard comes with Arcadia Studio. Full precision is a 325 MB download: the same voices, very slightly cleaner, at the same speed.
Speech recognition Tiny comes with Arcadia Studio and is the fastest. Base, Small, Medium and Large v3 Turbo are downloads that hear more accurately (especially names, accents, noise and other languages) but take longer. Install downloads one and checks it, and Remove deletes it. Only the one you choose is used.
Language Detect it, or the language you speak.
Microphone Which microphone to listen to. 🎤 Test writes down what it hears.

Downloads go to %LOCALAPPDATA%\Arcadia Studio\Speech, never into your project.

If the microphone hears nothing, check Windows Settings → Privacy & security → Microphone: Let desktop apps access your microphone must be on.

Recording your game

Preview has a ● Record button. It records only the game, exactly as players see it, without Preview's toolbars or console. Press it again (it shows ■ Stop and the time) to finish. Videos are saved as MP4 in Videos\Arcadia Studio, and Output and the Preview console say where.

  • Sound: the recording includes Arcadia Studio's own sound, which is the game and any narration, and nothing else your computer is playing (no music player, no notifications). Mute in Preview also mutes the game in the recording.
  • Covered windows: Preview is recorded even when another window covers it.
  • Speed: on computers with an NVIDIA, AMD or Intel graphics card the video is made by the graphics card, so the game keeps its frame rate. Every other Windows computer uses Windows' own video encoder.

Videos with your assistant

Over MCP, your assistant can make videos of your game for you. It opens Preview, starts recording, plays the game (walking the character with your game's own inputs, tapping buttons, setting up a scene with a script) and narrates as it goes. Afterwards it can add narration over a recording, trim it, make a GIF, pull out a still, or do anything else FFmpeg can.

It can also record the editor or a whole monitor, for a tutorial video of how to use Arcadia Studio.

What's inside

Part Made by Licence
Speech Kokoro 82M, run by KokoroSharp, with MisakiSharp working out how words are said Apache-2.0 (model), MIT and Apache-2.0 (code)
Speech recognition Whisper, run by whisper.cpp through Whisper.net MIT
Recording and video FFmpeg, an LGPL build, run as its own program from the FFmpeg folder LGPL-3.0

They are all in C# or native code that comes with the app: there's no Python and nothing to install. The full notices are in Docs\third-party beside the app.

Clone this wiki locally