Skip to content

How It Works

zhanglinghao edited this page Oct 2, 2026 · 3 revisions

English · 中文

The model doesn't paint pixels. It writes a program that, given a time t, draws the frame at t. A renderer (headless Chrome, Manim) calls that program once per frame, and FFmpeg joins the frames and the sound into an MP4. OpenVideoHarness is the craft around that loop: what to build, in what order, and how to check it. The rules themselves are in CLAUDE.md (the agent's entry point; AGENTS.md is an identical copy for Codex and other agents) and in the playbook.

OpenVideoHarness at a glance: a one-sentence request goes through type and effort selection, three human gates, sound-first code rendering and a self-review loop to a finished film; on the right, what the repository provides

Video as code

Because every frame is a pure function of t, three things follow:

  • It's deterministic. The same t always gives the same frame, so rendering can run in parallel, resume after a stop, or jump to any moment.
  • It can check itself. Any moment can be rendered on its own, so the agent lays frames out on a sheet and looks at them. It can't watch video or hear audio: it sees still frames and reads transcripts, timestamps and loudness numbers.
  • It can be forked. A new language, song or style is a code change and a re-render.

Engines: HyperFrames (HTML + GSAP) by default; Manim CE for exact math; the bundled ClaudeAnimationBase (p5.js + p5.brush) for hand-drawn work; Remotion if you prefer React; Blender for path-traced 3D; generative video with a code layer on top when you need real people or real physics. The selection table is in playbook/00.

Blender, driven by code. Some pictures a web engine can't give: real glass refraction, nebulae and volumetric light, real depth of field, motion blur on a million particles. For those the agent writes Python that drives Blender, and the frame is still a function of t: numpy works out where every point is at t, Cycles path-traces the plate, and the plate joins the web engine's type and UI in the same film. On macOS the bpy scripts run in a sandbox (no network, writes only to the output folder, none of your environment's API keys), and long renders run in resumable chunks after a 3–5 frame calibration. The intro film's 15.8 s opening and the tabletop-miniature style swatch were made this way. When it is worth it and what went wrong along the way: engines/blender.md, partly verified, with what has been run marked. Files that import bpy are GPL-3.0-or-later; the rest of the repo stays MIT.

The pipeline: 10 stages, 3 gates

Stage Output
0 Route video type and engine, from the routing table in CLAUDE.md
1 Material, concept, brief, outline the material list in NOTES.md; 2–3 concept cards; BRIEF.md (concept, audience, platform, length, frame, frame rate, resolution, where it's watched, sound, acceptance) and a 3–7 part outline → gate ①
2 Style STYLE.md: palette, type, easing, safe area, bans, all following from the concept; any preset it borrows from is noted in DECISIONS.md
3 Script + storyboard SCRIPT.md when there is narration, STORYBOARD.md with the reads of every shot, keyframe pages per section, optionally an animatic → gate ②
4 Audio first voiceover, word timings, beat map; measured durations written back into the storyboard, then the timeline is locked
5 Engine render script, shared helpers, one sample scene; the same frame rendered out of order must match
6 Scenes one file per scene; long films split into chapters for parallel subagents
7 Review the 20-item checklist, then a scored review by a fresh-context reviewer
8 Render a draft first → gate ③ → the final
9 Deliver MP4, source, render commands, asset log, LESSONS.md
flowchart LR
  A["0 Route<br/>type + engine"] --> B["1 Concept, brief<br/>+ outline"]
  B --> G1{{"Gate ①"}}
  G1 --> C["2 Style<br/>3 Storyboard"]
  C --> G2{{"Gate ②"}}
  G2 --> D["4 Audio first"]
  D --> E["5–6 Engine + scenes"]
  E --> F{"7 Review<br/>20 items · 8 scores"}
  F -- fails --> E
  F -- passes --> H["8 Draft"]
  H --> G3{{"Gate ③"}}
  G3 -- notes --> E
  G3 -- approved --> I["8 Final<br/>9 Deliver + lessons"]
Loading

Pass conditions for every stage: playbook/01-pipeline.md.

Concept first

Before style, recipes or the type's defaults, the agent looks for the film's concept: a device you can say in one sentence that makes the form itself tell the content, such as "the 3 seconds after Enter, slowed down to 2 minutes, with a clock on screen". It lists what only this material has (numbers, quotes, objects) in NOTES.md, writes 5–8 ideas of its own before it opens the style library, the recipes or the case studies, and then narrows them to 2–3 concept cards. Each card says the look it leads to (its own, or borrowed from one or more presets) and the hook, so picking a card at gate ① settles all three. One test of a concept: it fails on another subject; if it still works, it is only a style. The method is in playbook/12, and cases/oneshot-five.md breaks down five community films that each start from one idea.

Floors and taste defaults. The checklist tags its floor items, the ones that keep a film legible, correct and safe: the safe area, reading time, text size, no text burnt out by light, reads that nothing covers, loop seams, beat hits the storyboard declares, no placeholders or invented facts, and no digital silence in a film with sound. They hold for every concept and every effort level. Everything else (the type docs' registers and beat orders, the BRIEF's motion defaults, the recipes, the style presets, the checklist's other items) is a default that a concept may override, with one line in DECISIONS.md written before the scene that uses it. Overrides written on a card you picked count as your decision; the reviewer only checks they were carried out.

Spec up front

The frame, resolution, frame rate, length and where the film is watched are set when the project is made, with defaults per type (bin/vh new … --aspect --watch --res), and gate ① lists them as decided for you. The agent asks only when no platform was named and the answer would change the frame or the text size. Where it is watched (Watch on: in the BRIEF) sets the smallest text, in 1080p composition pixels:

Watch on What it means Titles Labels Captions
phone a vertical phone (also 1:1 and 4:5) ≥ 84 ≥ 44 ≥ 65
desktop a computer, or a phone turned sideways (Bilibili, YouTube, a website) ≥ 84 ≥ 44 ≥ 48
feed a landscape film in a phone feed, phone held upright (Douyin, WeChat Channels, Xiaohongshu, Weibo, X) ≥ 150 ≥ 80 ≥ 115

The readability check scales the contact sheet to that screen (360 px per frame for phone and feed, 640 for desktop). 4K (Resolution: 4k) is written at 1080p and rendered at twice the size, so text looks the same size on screen. Details: playbook/01 (规格) and playbook/03 §4.

Why the gates

Human judgment is worth most while changes are cheap. Changing a storyboard takes minutes; changing finished code takes hours. So the agent stops three times (concept and outline, storyboard with a keyframe sheet, first draft) and waits. It skips a gate only when you choose quick or say plainly in the chat that you don't want to review. Your exact words and the date go into REVIEW.md, and a fresh-context reviewer stands in for the gates you skipped.

Director mode

Effort says how hard the agent checks its own work. Director mode says what you decide. Twelve decisions can be named: concept, spec, outline, style, main character, theme music, voice, script, hook, storyboard, edit rhythm, and title and cover. Each one is handled in one of three ways:

What happens
own The agent shows options with a recommendation, then waits for you
review It shows one result on the next page and carries on; if you say nothing, it stands
delegate It decides, writes down why in DECISIONS.md, and you can overrule it any time

Name the ones you care about in the chat, on the Director: line of BRIEF.md, or as a default in LOCAL.md: Director: hook=own, character=own, theme=own, packaging=own, rest=delegate. Decisions nobody named follow the effort level. At quick everything is delegated. At standard the concept is yours to pick (its card carries the style and the hook, unless you name those separately), the outline and storyboard are reviewed and the rest is delegated; studio is the same but reviews the rest too.

The three gates stay the floor for standard and studio: director mode adds stops and never removes one. It adds them at six checkpoints, in time order, and you only stop at the ones you ask for (stop=E3 adds one by name):

Checkpoint When What you can decide
E0 style frames before the storyboard the main character, the theme melody (two short audio drafts and a score picture); the style, when all you can say is "it doesn't feel right" (look-dev variants)
E1 script before the storyboard the voice, chosen with the script so that line lengths are measured on it; the script; the hook
E2 sound before the music is written where the theme plays and where it stays quiet
E3 timing lock after the audio is made you hear the final voice and music once, then the timeline is locked; changing a line later costs a re-time
E4 sample chapter start of scene writing, long films the first chapter, built by hand as the standard for the rest
E5 picture lock after gate ③ after this only sound, colour and the ending change

Each decision has a first checkpoint where it can be made, and a cost for changing it later (the page says "now: low · after the storyboard: high"). The table of all of this is in playbook/01; CLAUDE.md has the rules, and the README has a plain-language version.

The review page

Every stop is, by default, one local HTML page. bin/vh review <project> <gate> turns out/review/gate-<n>.json into gate-<n>.html (the latest is also index.html). From the top:

  1. The decisions that need you, at most 3 by default. Each has a recommended option, the reason, what changing it later would cost, and how to reply (H1 / H2), then one line that accepts every recommendation.
  2. Pictures and options for each decision, side by side where possible. Sound comes as playable files, with a plain note on what the agent couldn't hear.
  3. For the storyboard, one page per section: keyframes, that section's stretch of the animatic, and the shots the agent is least sure about in red.
  4. A folded appendix for what doesn't change your decision.

It opens straight from disk, follows your light or dark theme, and the chat gets only a few lines: the decisions and the page's path. A choice or two in plain text can still go straight into the chat. Your replies are copied word for word into REVIEW.md.

Pictures to decide from

Judging a storyboard or a style shouldn't mean reading a table, so small commands build the pictures:

Command What you get
bin/vh storyboard <project> pages from shots.json: a keyframe per shot with its times, length and reads, least-sure shots in red, too-long shots flagged
bin/vh rhythm <project> one time axis with shots, narration, captions, on-screen text and the music's sections and hits; shots longer than the type's "new payoff every N s" and texts too short to read in red
bin/vh style compare a,b,c presets side by side, as posters or at the same moment of each swatch; bin/vh style apply attaches the ones you pick as references
bin/vh cover-preview <image> a cover at the real size of each feed slot (Bilibili, Douyin, YouTube, Xiaohongshu) with the platform's overlays, and a small-text legibility check
bin/vh music … --roll a piano roll and a loudness curve for each part: a picture of the music for a reviewer who can't hear it
bin/vh readcheck <project> reads the on-screen text of a HyperFrames composition without a browser and checks that each text stays long enough; --budget <s> says how many characters fit in a span, on screen, as subtitles and spoken

Reads: the unit of pacing

A read is one thing the viewer must understand. Each shot lists its reads in order, each with a start and end time. A read needs time to be found, understood and digested, and two important reads never overlap. If the reads don't fit, the shot gets longer or loses a read; nothing gets squeezed. Pacing is where AI video most often fails, which is why the storyboard comes before any code. (The idea comes from ClaudeAnimationBase's ANIMATION_GUIDE.md.)

Sound first

When there is sound, sound sets the timing. The agent makes the narration or music first, measures when every line, word and beat lands (timeline.json, music.beats.json), and the picture follows. A silent video takes its length from the reads and must make sense with the sound off. See Sound and Voice.

Self-review

The agent checks its work in layers, cheapest first (playbook/02):

  1. Does it run? Lint, compile, dry-run.
  2. Checks without looking. Bounding-box overlaps, black, frozen or silent stretches, frame counts, determinism spot checks.
  3. Looking at stills. A contact sheet for composition, a strip of consecutive frames for timing, a crop for faces and details; each cut ±0.5 s on its own; a sheet scaled to the target screen (360 px per frame for a phone or a phone feed, 640 for a computer).
  4. Critique as a choice. A vision model picks from grid positions or rendered variants instead of guessing pixel offsets.
  5. A reviewer who didn't make it. A fresh-context agent plays a harsh motion director and sees only the frames, the checklist and the storyboard, plus the BRIEF's concept line and the taste overrides in DECISIONS.md; never the author's notes.
  6. Whole-film viewing (optional). Gemini can read a video, but it samples about one frame per second: fine for the story, blind to fast motion.
  7. You. The final call on taste and sound.

Every draft has to clear two bars (TASTE_CHECKLIST.md):

  • 20 pass/fail items across composition, text, colour, motion, pacing and content. Any FAIL goes back for a fix; the floor items must pass whatever the concept.
  • 8 scored dimensions (1–10): concept (did the chosen idea land on screen?), hook, target-screen readability, motion quality, variety, composition and polish, content accuracy, audio-visual sync. Each one must reach 8; the average doesn't count. The scorer is never the author: when the first 26 style swatches were reviewed, their makers scored themselves 1–2 points higher than the independent reviewer did.

The first seven dimensions, the bar of 8 and the three rounds come from community practice; the checklist notes that the original sources could not all be verified. The concept dimension was added on 2026-10-02: the five community films in cases/oneshot-five.md win on their idea, and none of the seven asked about it. A low concept score is fixed by making the picture serve the chosen concept, not by swapping it; a reviewer who finds the concept itself weak takes that to gate ③.

Sound is measured the same way. bin/vh qa scans the final mix for digital silence, dropouts, pumping and clicks, and checks that every cue lands within one frame. For a profile= mix it also writes a mix report: how far the music sits under each narration line, which words are at risk of being masked, and whether the sound effects fall in their class ranges.

Effort

One switch sets how much of all this happens (see "努力程度" in CLAUDE.md, or run bin/vh effort):

quick standard (default) studio
For trying a direction, drafts most real videos launch films, flagship pieces
Gates none (decisions you name still stop) ① ② ③ ① ② ③, plus a 10–20 s sketch per concept card and a full-length animatic
Concept three one-line ideas, pick one 2–3 concept cards at gate ①, one frame each the same, each with a 10–20 s sketch
Self-review a contact sheet of the whole film, one scaled to the target screen, and a strip of the first 2 s every scene: sheet, strips, crops the same, plus the target-screen test, loop seams, lossless determinism, a scan for silent failures
Scored review none 1 round, fix the worst 3 at least 3 rounds; all 8 scores ≥ 8
Sound optional: none, or a score or voiceover with a basic mix score or voice, key-action SFX, a mix profile and bin/vh qa the same, plus foley and panning for every action and a mix report with no hard failures
A 30 s film takes about 10–30 min 1–2 h 3 h or more

Only you can lower the level; the agent may not downgrade to save time. Some things hold at every level: every frame a pure function of t, facts copied from the source, no API keys in files, no digital silence mid-film (in a film with sound), at most 3 full-screen white flashes per second, text no smaller than the floor for where the film is watched, recorded asset licenses, the checklist's floor items, and a bin/vh check before delivery.

The seven hard rules

  1. Every frame is a pure function of t: no Math.random(), Date.now(), CSS transitions or @keyframes; use seeded hashes.
  2. With sound, audio sets the timing.
  3. Storyboard before code.
  4. Self-review every scene.
  5. Facts are copied from the source; anything uncertain goes into NOTES.md, not the video.
  6. Work only inside projects/; engines are templates and references are read-only.
  7. No API keys in files; read them from environment variables.

When rules conflict: what you say in the chat, then the type doc, then the project's STYLE.md, then playbook/, then external references. That order is for rules; taste defaults (register, palette, beats, motion and transitions) can give way to the concept, with one line in DECISIONS.md.

Lessons flow back

Every project keeps a LESSONS.md. General lessons move into playbook/ or into a type doc's checks, so the next film starts better. Many rules in the repo began this way. For example, "a dramatic stop is a held breath, never digital silence" came from the intro film: its second version paused on true silence, and the user heard the music stall.

Research notes

Some of the rules above came from measurements. docs/research/ has six lab notes, in Chinese and in English. Each gives the question, how it was measured, the numbers, what changed in the repo and what is still unclear:

  1. voice, music and sound-effect levels, behind the mix profiles;
  2. 28 soundtracks: how to measure "too similar";
  3. same code, different pixels: why the swatches now render on the CPU;
  4. how long on-screen text has to stay, behind bin/vh readcheck;
  5. chapters, motif and dynamic arc, behind playbook/11;
  6. concept first against floors only, an A/B on two one-line requests: behind the quick path's two extra self-checks, and the question that led to the text-size floors by screen.

They are records, not rules (the rules stay in the playbook), and mostly meter readings rather than listening verdicts.

Clone this wiki locally