-
Notifications
You must be signed in to change notification settings - Fork 0
Sound and Voice
English · 中文
An agent can't hear, so sound here is built to be computed and measured. Narration and music are made first, and their measured timings drive the picture. Every command reads and writes plain files in the project's audio/ folder; the Python dependencies are fetched on first use into uv's cache, never into a global environment. The full reference is playbook/04-audio.md (in Chinese).
bin/vh tts projects/<p> qwen Serena zh # audio/script.txt → voiceover.zh.wav + timeline.zh.json
bin/vh captions projects/<p> zh # captions.zh / en / bi .srt + captions.json
bin/vh music --example > projects/<p>/audio/score.json # a starter score (--example zh: Chinese instruments; --example list: all of them)
bin/vh music projects/<p>/audio/score.json projects/<p>/audio/music.wav # + music.beats.json
bin/vh sfx lib projects/<p>/audio/sfx # write the 21 built-in effects (plain versions) and sfx-lib.json into the project
bin/vh mix projects/<p>/audio/mix.wav profile=short voice=projects/<p>/audio/voiceover.zh.wav music=projects/<p>/audio/music.wav \
events=projects/<p>/audio/events.json lib=projects/<p>/audio/sfx timeline=projects/<p>/audio/timeline.zh.json stems=projects/<p>/audio/stems
bin/vh qa projects/<p>/audio/mix.wav projects/<p>/audio/music.beats.json projects/<p>/audio/events.json --stems projects/<p>/audio/stems
bin/vh mux projects/<p>/out/final.mp4 projects/<p>/audio/mix.wav projects/<p>/out/final-av.mp4 \
projects/<p>/audio/captions.zh.srt projects/<p>/audio/captions.en.srt # soft subtitle tracksThe showcase films 00–03 each have a tools/build_audio.sh that runs the whole chain from their committed sources and gives the same bytes every time; 02 (Chinese narration) and 03 (English narration) are ready-made examples of narration with word timings, captions and a mix profile together.
| Provider | What it is | Needs | Use it for | Tested |
|---|---|---|---|---|
qwen (default) |
Local open-source Qwen3-TTS via mlx-audio. Chinese: Serena, Vivian, Uncle_Fu, Dylan (Beijing), Eric (Sichuan). English: Aiden (default), Ryan | Apple Silicon; about 2 GB download on first use | free, offline narration on a Mac | yes, zh + en |
say |
macOS built-in voices | macOS | instant drafts | yes |
edge |
Microsoft Edge online voices through the unofficial edge-tts
|
network | free narration away from a Mac; may break at any time | yes, English |
dashscope |
Qwen3-TTS on Alibaba Cloud Model Studio | DASHSCOPE_API_KEY |
a cloud voice hosted in China | not yet |
elevenlabs |
ElevenLabs multilingual voices |
ELEVENLABS_API_KEY, a voice id |
if you already use ElevenLabs | not yet |
gemini, gemini-lite
|
Google Gemini 3.8 Flash TTS and Flash-Lite TTS; 30 studio voices (default Kore), each speaks Chinese too |
GEMINI_API_KEY; free and paid tiers |
expressive narration, per-line direction, two-speaker dialogue; gemini-lite for bulk single-voice narration |
gemini: yes, zh + en |
Known issue: the default Qwen3-TTS 0.6B model sometimes keeps mumbling after a short English line. Check every take with --align gemini (below), or at least look at the durations in timeline.<lang>.json; re-run a flagged line, switch to a 1.7B model (QWEN_TTS_MODEL), or use gemini. Aiden became the default English voice because Ryan went wrong far more often. Each provider's silence at the start and end of a line is trimmed (below −50 dBFS, keeping 30 ms before the voice and 80 ms after), so gaps are the gaps you set and beat-snapped lines start speaking on the beat.
Gemini audio carries an inaudible SynthID watermark: note AI narration in NOTES.md and label it where a platform requires it. The free tier's content may be used to improve Google's products. Prices and limits are in playbook/04.
audio/script.txt holds one line per caption:
@hook [like sharing an amazing find with a friend, quick, rising at the end] 一句话,能做出一支片子吗? || One sentence. Can it make a film?
@yes:downbeat [land the answer, short and bright] 能! || It can.
-
中文 || Englishkeeps both languages.--langpicks the side that is spoken; both sides go into the timeline for captions. A script with no Chinese is spoken in English. -
@idnames a line.[direction]directs that line only and never reaches the captions; the overall tone is the fifth argument ofbin/vh tts(--instruct). - Tags like
<short pause>,<breath>and<laugh>are performed bygeminionly and removed everywhere else, captions included. - Who follows directions:
geminibest; localqwenonly with a 1.7B instruct model;sayandedgejust read. - Write what is said, not what is shown: units, abbreviations and symbols go to the voice in their spoken form ("2 G赫兹", not "2 GHz"). When the caption must differ, synthesize the line from the spoken form, keep the caption text, and note the spoken form in
script.txt. Listen to numbers and units after synthesis; the transcriber's first pass is unreliable on them. - How much fits: a knowledge short speaks about 4.5–5.5 Chinese characters a second within a sentence, an explainer or paper 3.5–4.5, and a script budget takes 15 % off for pauses (a 45 s short is about 170–210 characters).
bin/vh readcheck --budget <s>gives the number for a span: on screen, as subtitles and spoken. The default delivery per video type is tabled in playbook/04.
Narration for motion pieces, launch films, MVs and memes rides the music:
bin/vh tts projects/<p> gemini Kore zh --beats projects/<p>/audio/music.beats.json --snap beatEach line starts on the next grid point (--snap beat|half|downbeat; --lead sets the earliest start), and @id:downbeat pins one line, say the answer on the drop. The loop: sketch the score, snap the narration, read where each line landed in the timeline, adjust section lengths in the score, then build the picture from the timeline. In a live test at 120 BPM the lines landed at exactly 0.50, 4.00, 5.00 and 10.00 s.
This works with any provider. Each line is transcribed by gemini-3.5-transcribe for word timestamps and compared with the script. A line is flagged when similarity falls below --min-sim (default 0.85) or there is more than 1 s of sound before the first word or after the last, which catches the Qwen run-on and skipped words. Timestamps come in 0.1 s steps. Gemini's Tier 1 allows 10 transcriptions a minute and about 100 a day: on a 429 the tool waits as long as the API asks and retries, while a 429 that asks for more than 90 s (a daily quota) fails at once.
A failed transcription doesn't lose the take. The voiceover and the timeline are written right after synthesis, a line the check couldn't finish keeps its measured timing, and the command exits 1 with the reason. Run the same command with --resume: it redoes only what is missing and doesn't pay again for lines that already succeeded.
@speakers A=Kore B=Puck
@q1 A: [curious] 你猜这支片子手写了几行代码? || Guess how many lines of this film were written by hand?
@a1 B: [lower, teasing] 一行都没有。 |嗯?| 全是它自己写的。 || Not a single one. |huh?| It wrote all of them.
Only labels declared in @speakers count as speakers. |嗯?| is the other speaker's backchannel, voiced only in a Gemini dialogue. A Gemini dialogue goes out as one conversational request per block (at most 2 speakers, library voices). --join block|all does the same for narration: the delivery flows across lines, but a line inside a block can't be re-run or beat-snapped on its own.
bin/vh captions <project> [zh|en] [zh-max] writes captions.zh.srt, captions.en.srt, a two-line captions.bi.srt, and captions.json for engines. For vertical video pass 11 as zh-max (at most 11 Chinese characters per line). With word timings, captions.json also holds per-word times, and captions.<lang>.lines.srt shows each wrapped line when its first word is spoken. Burn captions into the picture (drawn from captions.json, still a pure function of t), or add them as switchable tracks with bin/vh mux. When the picture already has captions burned in, bin/vh mux --subs-off stops a player from switching a second set on by itself.
A caption shorter than 1.8 s is lengthened into the silence after it, up to 1.8 s, the next caption's start or the end of the picture (--media-end S, else the project's index.html, media/final.mp4 or the timeline); its start never moves and nothing overlaps. Burned-in captions that repeat the narration are checked as subtitles: mark them data-read="subtitle" and bin/vh readcheck holds them to subtitle reading speed instead of the slower on-screen rule.
bin/vh music score.json music.wav renders a score written in JSON: tempo, key and mode; sections with chords, layers (kick, clap, hats, bass, pad, arp, lead, plus bell, zheng, dizi and taiko for Chinese colour) and energy; section effects riser, impact, fill and bend; and meters for changes of time signature. The same score always gives the same audio, and music.beats.json lists the exact sections, beats, downbeats and hits for the picture to land on. Choose tempos where one beat is a whole number of frames (at 30 fps: 90, 100, 112.5, 120, 150 or 180 BPM).
Instrument parts. layers is one synthesiser palette, so scores written with it alone sound alike. A score can also add parts: 91 instruments synthesised in code with no samples (keys, mallets, plucked, bowed and wind instruments, synths, drum kits, percussion including a 锣鼓经 kit, textures such as vinyl and rain), played from step grids, note lists with chord-relative pitches or ready figures (walking bass, oompah, waltz, strum, arpeggios), with stereo, reverb sends and optional lo-fi and tape. bin/vh music --instruments lists them all (102 with the 11 layers voices) and --example list the starter scores. A score with only layers renders the same bytes as before, and a part that is silent, or under −40 dBFS in the final file, stops the render and names the part.
Plucked strings, modelled. Fourteen more voices are opt-in: physical models of existing ones (guqin_pm, pipa_pm, harp_pm, nylon_pm, ukulele_pm, upright_pm, balalaika_pm, cimbalom_pm) and new instruments (guitar, koto, shamisen, banjo, kalimba, musicbox_pm). The strings are digital waveguides, so pitch can slide, bend and vibrate inside one note and a hard pluck starts sharp and settles; the kalimba and music box are modal bars. Switching a part to one is a one-word change of inst, which makes A/B easy; check gain_db afterwards, because levels differ from the older voices in the upper register. Their checks are measurements: pitch within about a cent over three octaves, level against a plucked note, identical bytes on every render.
Motifs and dynamics. A phrase is written once under motifs and called anywhere: shifted, inverted, reversed, stretched, cut into fragments or repeated as a sequence. Inside a section, dyn is a fader line, cresc and dim ride a part across a section, a section stop leaves only the parts you keep (a held breath), and a negative beat is a pickup. Not there yet: a tempo map (ritardando, accelerando). playbook/11 shows how to use all this for chapters, a theme you can hum and a dynamic arc.
A note map for pictures driven note by note. A part with "note_map": true writes every note into the beat map's notes (start, end, pitch, velocity, part), so keys can go down as their notes sound or a character can walk on a melody, with the timing read from the map and never typed into the scene.
Pictures of the music. Nobody on the agent's side can hear the score, so two flags make it visible. --roll writes a piano roll for each part with its loudness, and the mix loudness against the score's energy. --length S|auto ends the file exactly at S (auto: where the last bar ends) with a fade that starts on a beat, instead of a tail that the mux then has to cut.
Sources, in order: a track you own or licensed (bin/vh beats finds its beats, hits, kicks and snares); a local music model you have checked actually runs; code composition. Don't download background music of unknown origin or lift music from a reference video. Songs with vocals come from a service such as Suno; ElevenLabs Music and local song models are reserved interfaces.
audio/events.json lists {t, sfx, gain_db, pan, dist}, and each sound lands on the action's frame. The 21 built-in effects are original and synthesized in code:
- transitions: whoosh, swish_rev, whip, swoosh_tonal, air, paper, tape, shimmer;
- small actions: click, tick, pop, typing, shutter, glitch;
- build-ups and hits: riser, impact, boom;
- signals: ding, success, error, toggle.
Pick a transition family that suits the style (paper for hand-drawn and collage, air for calm and documentary, whip for cartoon and fast cuts, swoosh_tonal for sci-fi and UI, tape for retro). Every event of a built-in, apart from the signals, gets its own variant, a step along a walk for that sound through the film, so consecutive events are never the same and never far apart; "variant": n pins one. An event can also be shaped to its move: dur, pitch or center, dir, bright, tone, and pan_from / pan_to. bin/vh sfx audition <name|all> plays n variants of a sound in one file, with a line of measurements each and an optional spectrogram, and bin/vh qa warns when a sound comes back the same three or more times in a row. A profile mix places the events itself (events=… lib=…); bin/vh sfx place renders a standalone effects track. Work out pan from the picture, pan = 2·x / width − 1, scaled by 0.7–0.8. Licensed recordings come first; log their source and license in NOTES.md.
For levelling, each event is in a class: hero (an impact, a bell: the moment), detail (a click, a step), ambience (wind, a room) or signal (a readout, a tick that carries information). The class comes from the event's role, else from whole words in the sound's name. Write role when the name misleads: in the hand-drawn film a gust of wind is the gag itself, so it is a detail, not ambience.
bin/vh mix combines voice, music and SFX into one stereo track at −14 LUFS. Pick a profile for the video type, and every level is set relative to one anchor: the narration's median line, or the music when there is no narration.
| Profile | For | Music under the narration (target, normal range, in LU) |
|---|---|---|
explainer |
01 explainers, 06 papers | 13 (11–18) |
short |
02 knowledge shorts | 11.5 (10–16) |
promo |
03 launch films, 05 data stories, 08 fast cuts | 10 (8–16) |
cartoon |
07 hand-drawn shorts | 10 (8–16) |
mv |
04 music videos | 6 (4–12): the music leads |
swatch |
the style swatches | no narration; the music is the anchor |
That number is the VMR: how many loudness units the voice sits over the music. The mixer works in this order:
- the narration's lines are levelled toward their median;
- the music is ridden line by line, with look-ahead, until each line sits at the profile's VMR; a 1–4 kHz carve (plus 250 Hz–1 kHz for Chinese) goes only as deep as the words need;
- each sound effect moves half way to the range of its class, and a low-frequency hit is never raised;
- all effects share one short room;
- the master is one static gain on a BS.1770 meter and a true-peak limiter, so a cinematic score keeps its dynamics.
A=projects/<p>/audio
bin/vh mix $A/mix.wav profile=explainer voice=$A/voiceover.en.wav music=$A/music.wav \
events=$A/events.json lib=$A/sfx timeline=$A/timeline.en.json music_db=-5 stems=$A/stems
bin/vh mix $A/mix.wav profile=cartoon music=$A/music.wav events=$A/events.json lib=$A/sfx dur=12 fade=0.04 stems=$A/stems # no narration: the music is the anchorstems= writes each bus at its final gain, for bin/vh qa. Without profile= the older ffmpeg chain runs:
- Default
duck=voice: the music makes way under narration. With no voice bus the default isduck=off. - Default
duck_ratio=1.6. In a listening test with 0.25 s pauses between lines, ratios of 6 and 3 made the music drop out between lines; 1.6 didn't, and sounded best. - In motion pieces the music carries the rhythm: duck lightly (
music_db=-5 duck=voicewithduck_ratioat 1.5–2) or not at all. -
duck=onalso lets SFX push the music down; keep the ratio at 2–3 or the music pumps on every hit.
Measure the encode. A profile mix leaves the WAV at −1.65 dBTP, but an AAC encode adds to the true peak, by an amount that depends on the content and the bitrate (from −0.1 to +0.1 dB at 192k on the showcase films, up to about +1.4 dB at 128k on dense swatches). Measure the encoded file (ffmpeg -i final.mp4 -af ebur128=peak=true -f null -) and mix again with a lower tp= if it is over −1.5 dBTP. The swatch renderer does that by itself.
bin/vh mux puts the track on the picture: it writes the index first (faststart), so a browser starts playing before the whole file has loaded, and it fades the last 40 ms when it trims a longer track to the picture.
bin/vh qa <mix> [beats.json] [events.json] [--stems DIR] runs on the final mix and exits 1 on failure:
- digital silence: every channel below −60 dBFS for 20 ms or more, anywhere mid-film;
- dropouts: 0.1 s windows 12 dB below their section's median;
-
pumping: dips over 4 dB lasting 60 ms or more (give it
--stems, or--voice voiceover.wav, so the dip designed under narration and the score's own dynamics don't count); - clicks: a warning only; listen to the 10 worst by ear;
- repeats: a warning only, when the same sound comes back three or more times in a row (a rotation of takes is fine);
-
cue check: every beat-map hit and SFX event must have an onset within one frame. Each cue prints its margin, and a barely found one is marked
OK~. A near-pure tone (a tick, a ding) that the onset detector misses is confirmed by matching its own sound. On an AAC file the tolerance grows by 12 ms, and a whole encode that is off, which means a muxing offset, fails.
With --stems it also prints the mix report (bin/vh qa mix <stems dir> prints only that): the VMR of every line, the words at risk of being masked in the 1–4 kHz band, the music in each pause, every effect against the anchor and its class, and the depth of each bus. It fails when a line is under the profile's floor, a hero sound is louder than the voice during speech, too many words are at risk, a class median is more than 3 LU out of range, or a cue is found but flagged BURIED, which means it is there but can't be heard. How to read each flag and what to change is in playbook/04 ("混音").
The report levels things toward the middle; it can't judge which effect matters, and a low-frequency hit needs a 1–4 kHz layer rather than more gain. A dramatic stop is a held breath: keep a low bed, close the filter, drop the drums, swell into the next hit.
API keys live only in environment variables: GEMINI_API_KEY, ELEVENLABS_API_KEY, DASHSCOPE_API_KEY, and for generative video for example FAL_KEY. Nothing is written to the repo, and LOCAL.md lists variable names, never values. When a key is missing, the tool tells you which one it needs.