-
Notifications
You must be signed in to change notification settings - Fork 0
How It Works
English · 中文
The model doesn't paint pixels. It writes a program that, given a time t, draws the frame at t. A renderer (headless Chrome, Manim) calls that program once per frame, and FFmpeg joins the frames and the sound into an MP4. OpenVideoHarness is the craft around that loop: what to build, in what order, and how to check it. The rules themselves are in CLAUDE.md (the agent's entry point; AGENTS.md is an identical copy for Codex and other agents) and in the playbook.
Because every frame is a pure function of t, three things follow:
- It's deterministic. The same t always gives the same frame, so rendering can run in parallel, resume after a stop, or jump to any moment.
- It can check itself. Any moment can be rendered on its own, so the agent lays frames out on a sheet and looks at them. It can't watch video or hear audio: it sees still frames and reads transcripts, timestamps and loudness numbers.
- It can be forked. A new language, song or style is a code change and a re-render.
Engines: HyperFrames (HTML + GSAP) by default; Manim CE for exact math; the bundled ClaudeAnimationBase (p5.js + p5.brush) for hand-drawn work; Remotion if you prefer React; Blender for path-traced 3D; generative video with a code layer on top when you need real people or real physics. The selection table is in playbook/00.
Blender, driven by code. Some pictures a web engine can't give: real glass refraction, nebulae and volumetric light, real depth of field, motion blur on a million particles. For those the agent writes Python that drives Blender, and the frame is still a function of t: numpy works out where every point is at t, Cycles path-traces the plate, and the plate joins the web engine's type and UI in the same film. On macOS the bpy scripts run in a sandbox (no network, writes only to the output folder, none of your environment's API keys), and long renders run in resumable chunks after a 3–5 frame calibration. The intro film's 15.8 s opening and the tabletop-miniature style swatch were made this way. When it is worth it and what went wrong along the way: engines/blender.md, partly verified, with what has been run marked. Files that import bpy are GPL-3.0-or-later; the rest of the repo stays MIT.
| Stage | Output |
|---|---|
| 0 Route | video type and engine, from the routing table in CLAUDE.md
|
| 1 Material, concept, brief, outline | the material list in NOTES.md; 2–3 concept cards; BRIEF.md (concept, audience, platform, length, frame, frame rate, resolution, where it's watched, sound, acceptance) and a 3–7 part outline → gate ①
|
| 2 Style |
STYLE.md: palette, type, easing, safe area, bans, all following from the concept; any preset it borrows from is noted in DECISIONS.md
|
| 3 Script + storyboard |
SCRIPT.md when there is narration, STORYBOARD.md with the reads of every shot, keyframe pages per section, optionally an animatic → gate ②
|
| 4 Audio first | voiceover, word timings, beat map; measured durations written back into the storyboard, then the timeline is locked |
| 5 Engine | render script, shared helpers, one sample scene; the same frame rendered out of order must match |
| 6 Scenes | one file per scene; long films split into chapters for parallel subagents |
| 7 Review | the 20-item checklist, then a scored review by a fresh-context reviewer |
| 8 Render | a draft first → gate ③ → the final |
| 9 Deliver | MP4, source, render commands, asset log, LESSONS.md
|
flowchart LR
A["0 Route<br/>type + engine"] --> B["1 Concept, brief<br/>+ outline"]
B --> G1{{"Gate ①"}}
G1 --> C["2 Style<br/>3 Storyboard"]
C --> G2{{"Gate ②"}}
G2 --> D["4 Audio first"]
D --> E["5–6 Engine + scenes"]
E --> F{"7 Review<br/>20 items · 8 scores"}
F -- fails --> E
F -- passes --> H["8 Draft"]
H --> G3{{"Gate ③"}}
G3 -- notes --> E
G3 -- approved --> I["8 Final<br/>9 Deliver + lessons"]
Pass conditions for every stage: playbook/01-pipeline.md.
Before style, recipes or the type's defaults, the agent looks for the film's concept: a device you can say in one sentence that makes the form itself tell the content, such as "the 3 seconds after Enter, slowed down to 2 minutes, with a clock on screen". It lists what only this material has (numbers, quotes, objects) in NOTES.md, writes 5–8 ideas of its own before it opens the style library, the recipes or the case studies, and then narrows them to 2–3 concept cards. Each card says the look it leads to (its own, or borrowed from one or more presets) and the hook, so picking a card at gate ① settles all three. One test of a concept: it fails on another subject; if it still works, it is only a style. The method is in playbook/12, and cases/oneshot-five.md breaks down five community films that each start from one idea.
Floors and taste defaults. The checklist tags its floor items, the ones that keep a film legible, correct and safe: the safe area, reading time, text size, no text burnt out by light, reads that nothing covers, loop seams, beat hits the storyboard declares, no placeholders or invented facts, and no digital silence in a film with sound. They hold for every concept and every effort level. Everything else (the type docs' registers and beat orders, the BRIEF's motion defaults, the recipes, the style presets, the checklist's other items) is a default that a concept may override, with one line in DECISIONS.md written before the scene that uses it. Overrides written on a card you picked count as your decision; the reviewer only checks they were carried out.
The frame, resolution, frame rate, length and where the film is watched are set when the project is made, with defaults per type (bin/vh new … --aspect --watch --res), and gate ① lists them as decided for you. The agent asks only when no platform was named and the answer would change the frame or the text size. Where it is watched (Watch on: in the BRIEF) sets the smallest text, in 1080p composition pixels:
Watch on |
What it means | Titles | Labels | Captions |
|---|---|---|---|---|
phone |
a vertical phone (also 1:1 and 4:5) | ≥ 84 | ≥ 44 | ≥ 65 |
desktop |
a computer, or a phone turned sideways (Bilibili, YouTube, a website) | ≥ 84 | ≥ 44 | ≥ 48 |
feed |
a landscape film in a phone feed, phone held upright (Douyin, WeChat Channels, Xiaohongshu, Weibo, X) | ≥ 150 | ≥ 80 | ≥ 115 |
The readability check scales the contact sheet to that screen (360 px per frame for phone and feed, 640 for desktop). 4K (Resolution: 4k) is written at 1080p and rendered at twice the size, so text looks the same size on screen. Details: playbook/01 (规格) and playbook/03 §4.
Human judgment is worth most while changes are cheap. Changing a storyboard takes minutes; changing finished code takes hours. So the agent stops three times (concept and outline, storyboard with a keyframe sheet, first draft) and waits. It skips a gate only when you choose quick or say plainly in the chat that you don't want to review. Your exact words and the date go into REVIEW.md, and a fresh-context reviewer stands in for the gates you skipped.
Effort says how hard the agent checks its own work. Director mode says what you decide. Twelve decisions can be named: concept, spec, outline, style, main character, theme music, voice, script, hook, storyboard, edit rhythm, and title and cover. Each one is handled in one of three ways:
| What happens | |
|---|---|
own |
The agent shows options with a recommendation, then waits for you |
review |
It shows one result on the next page and carries on; if you say nothing, it stands |
delegate |
It decides, writes down why in DECISIONS.md, and you can overrule it any time |
Name the ones you care about in the chat, on the Director: line of BRIEF.md, or as a default in LOCAL.md: Director: hook=own, character=own, theme=own, packaging=own, rest=delegate. Decisions nobody named follow the effort level. At quick everything is delegated. At standard the concept is yours to pick (its card carries the style and the hook, unless you name those separately), the outline and storyboard are reviewed and the rest is delegated; studio is the same but reviews the rest too.
The three gates stay the floor for standard and studio: director mode adds stops and never removes one. It adds them at six checkpoints, in time order, and you only stop at the ones you ask for (stop=E3 adds one by name):
| Checkpoint | When | What you can decide |
|---|---|---|
| E0 style frames | before the storyboard | the main character, the theme melody (two short audio drafts and a score picture); the style, when all you can say is "it doesn't feel right" (look-dev variants) |
| E1 script | before the storyboard | the voice, chosen with the script so that line lengths are measured on it; the script; the hook |
| E2 sound | before the music is written | where the theme plays and where it stays quiet |
| E3 timing lock | after the audio is made | you hear the final voice and music once, then the timeline is locked; changing a line later costs a re-time |
| E4 sample chapter | start of scene writing, long films | the first chapter, built by hand as the standard for the rest |
| E5 picture lock | after gate ③ | after this only sound, colour and the ending change |
Each decision has a first checkpoint where it can be made, and a cost for changing it later (the page says "now: low · after the storyboard: high"). The table of all of this is in playbook/01; CLAUDE.md has the rules, and the README has a plain-language version.
Every stop is, by default, one local HTML page. bin/vh review <project> <gate> turns out/review/gate-<n>.json into gate-<n>.html (the latest is also index.html). From the top:
-
The decisions that need you, at most 3 by default. Each has a recommended option, the reason, what changing it later would cost, and how to reply (
H1 / H2), then one line that accepts every recommendation. - Pictures and options for each decision, side by side where possible. Sound comes as playable files, with a plain note on what the agent couldn't hear.
- For the storyboard, one page per section: keyframes, that section's stretch of the animatic, and the shots the agent is least sure about in red.
- A folded appendix for what doesn't change your decision.
It opens straight from disk, follows your light or dark theme, and the chat gets only a few lines: the decisions and the page's path. A choice or two in plain text can still go straight into the chat. Your replies are copied word for word into REVIEW.md.
Judging a storyboard or a style shouldn't mean reading a table, so small commands build the pictures:
| Command | What you get |
|---|---|
bin/vh storyboard <project> |
pages from shots.json: a keyframe per shot with its times, length and reads, least-sure shots in red, too-long shots flagged |
bin/vh rhythm <project> |
one time axis with shots, narration, captions, on-screen text and the music's sections and hits; shots longer than the type's "new payoff every N s" and texts too short to read in red |
bin/vh style compare a,b,c |
presets side by side, as posters or at the same moment of each swatch; bin/vh style apply attaches the ones you pick as references |
bin/vh cover-preview <image> |
a cover at the real size of each feed slot (Bilibili, Douyin, YouTube, Xiaohongshu) with the platform's overlays, and a small-text legibility check |
bin/vh music … --roll |
a piano roll and a loudness curve for each part: a picture of the music for a reviewer who can't hear it |
bin/vh readcheck <project> |
reads the on-screen text of a HyperFrames composition without a browser and checks that each text stays long enough; --budget <s> says how many characters fit in a span, on screen, as subtitles and spoken |
A read is one thing the viewer must understand. Each shot lists its reads in order, each with a start and end time. A read needs time to be found, understood and digested, and two important reads never overlap. If the reads don't fit, the shot gets longer or loses a read; nothing gets squeezed. Pacing is where AI video most often fails, which is why the storyboard comes before any code. (The idea comes from ClaudeAnimationBase's ANIMATION_GUIDE.md.)
When there is sound, sound sets the timing. The agent makes the narration or music first, measures when every line, word and beat lands (timeline.json, music.beats.json), and the picture follows. A silent video takes its length from the reads and must make sense with the sound off. See Sound and Voice.
The agent checks its work in layers, cheapest first (playbook/02):
- Does it run? Lint, compile, dry-run.
- Checks without looking. Bounding-box overlaps, black, frozen or silent stretches, frame counts, determinism spot checks.
- Looking at stills. A contact sheet for composition, a strip of consecutive frames for timing, a crop for faces and details; each cut ±0.5 s on its own; a sheet scaled to the target screen (360 px per frame for a phone or a phone feed, 640 for a computer).
- Critique as a choice. A vision model picks from grid positions or rendered variants instead of guessing pixel offsets.
-
A reviewer who didn't make it. A fresh-context agent plays a harsh motion director and sees only the frames, the checklist and the storyboard, plus the BRIEF's concept line and the taste overrides in
DECISIONS.md; never the author's notes. - Whole-film viewing (optional). Gemini can read a video, but it samples about one frame per second: fine for the story, blind to fast motion.
- You. The final call on taste and sound.
Every draft has to clear two bars (TASTE_CHECKLIST.md):
- 20 pass/fail items across composition, text, colour, motion, pacing and content. Any FAIL goes back for a fix; the floor items must pass whatever the concept.
- 8 scored dimensions (1–10): concept (did the chosen idea land on screen?), hook, target-screen readability, motion quality, variety, composition and polish, content accuracy, audio-visual sync. Each one must reach 8; the average doesn't count. The scorer is never the author: when the first 26 style swatches were reviewed, their makers scored themselves 1–2 points higher than the independent reviewer did.
The first seven dimensions, the bar of 8 and the three rounds come from community practice; the checklist notes that the original sources could not all be verified. The concept dimension was added on 2026-10-02: the five community films in cases/oneshot-five.md win on their idea, and none of the seven asked about it. A low concept score is fixed by making the picture serve the chosen concept, not by swapping it; a reviewer who finds the concept itself weak takes that to gate ③.
Sound is measured the same way. bin/vh qa scans the final mix for digital silence, dropouts, pumping and clicks, and checks that every cue lands within one frame. For a profile= mix it also writes a mix report: how far the music sits under each narration line, which words are at risk of being masked, and whether the sound effects fall in their class ranges.
One switch sets how much of all this happens (see "努力程度" in CLAUDE.md, or run bin/vh effort):
quick |
standard (default) |
studio |
|
|---|---|---|---|
| For | trying a direction, drafts | most real videos | launch films, flagship pieces |
| Gates | none (decisions you name still stop) | ① ② ③ | ① ② ③, plus a 10–20 s sketch per concept card and a full-length animatic |
| Concept | three one-line ideas, pick one | 2–3 concept cards at gate ①, one frame each | the same, each with a 10–20 s sketch |
| Self-review | a contact sheet of the whole film, one scaled to the target screen, and a strip of the first 2 s | every scene: sheet, strips, crops | the same, plus the target-screen test, loop seams, lossless determinism, a scan for silent failures |
| Scored review | none | 1 round, fix the worst 3 | at least 3 rounds; all 8 scores ≥ 8 |
| Sound | optional: none, or a score or voiceover with a basic mix | score or voice, key-action SFX, a mix profile and bin/vh qa
|
the same, plus foley and panning for every action and a mix report with no hard failures |
| A 30 s film takes about | 10–30 min | 1–2 h | 3 h or more |
Only you can lower the level; the agent may not downgrade to save time. Some things hold at every level: every frame a pure function of t, facts copied from the source, no API keys in files, no digital silence mid-film (in a film with sound), at most 3 full-screen white flashes per second, text no smaller than the floor for where the film is watched, recorded asset licenses, the checklist's floor items, and a bin/vh check before delivery.
- Every frame is a pure function of t: no
Math.random(),Date.now(), CSS transitions or@keyframes; use seeded hashes. - With sound, audio sets the timing.
- Storyboard before code.
- Self-review every scene.
- Facts are copied from the source; anything uncertain goes into
NOTES.md, not the video. - Work only inside
projects/; engines are templates and references are read-only. - No API keys in files; read them from environment variables.
When rules conflict: what you say in the chat, then the type doc, then the project's STYLE.md, then playbook/, then external references. That order is for rules; taste defaults (register, palette, beats, motion and transitions) can give way to the concept, with one line in DECISIONS.md.
Every project keeps a LESSONS.md. General lessons move into playbook/ or into a type doc's checks, so the next film starts better. Many rules in the repo began this way. For example, "a dramatic stop is a held breath, never digital silence" came from the intro film: its second version paused on true silence, and the user heard the music stall.
Some of the rules above came from measurements. docs/research/ has six lab notes, in Chinese and in English. Each gives the question, how it was measured, the numbers, what changed in the repo and what is still unclear:
- voice, music and sound-effect levels, behind the mix profiles;
- 28 soundtracks: how to measure "too similar";
- same code, different pixels: why the swatches now render on the CPU;
- how long on-screen text has to stay, behind
bin/vh readcheck; - chapters, motif and dynamic arc, behind playbook/11;
- concept first against floors only, an A/B on two one-line requests: behind the
quickpath's two extra self-checks, and the question that led to the text-size floors by screen.
They are records, not rules (the rules stay in the playbook), and mostly meter readings rather than listening verdicts.