Releases: ZLHad/OpenVideoHarness
Release list
Showcase media (videos)
The videos of OpenVideoHarness's showcase films and style library. They live here instead of in git, so cloning the repo stays small. The READMEs play short clips inline and link here for the full files.
本仓库样片和风格库的视频文件。它们放在这里而不是 git 里,clone 仓库不用下载它们。README 里能直接播放短片段,完整文件从这里下载。
| File | What |
|---|---|
intro-film-1080p.mp4 |
04 · Intro film v5, 103 s, 1920×1080, 30 fps, with sound (297 MB) |
00-launch-short.mp4 |
00 · Launch short, 20 s |
01-handdrawn-short.mp4 |
01 · Hand-drawn character short, 12 s |
02-vertical-science-short.mp4 |
02 · Vertical science short, 24.8 s, 1080×1920 |
03-math-explainer.mp4 |
03 · 3b1b-style math explainer, 25 s |
styles-gallery.mp4 |
The 31 style samples in one reel |
intro-film-v3-81s.mp4 |
The previous intro film (v3, 81 s) |
intro-film-gate1-animatic.mp4 |
The intro's gate ① grey-box animatic with the score sketch |
SHA256SUMS.txt |
Checksums of the files here (tools/fetch_media.sh checks the copies in the repo's tools/media.txt) |
To put them back where the build scripts expect them: tools/fetch_media.sh (from the repo root).
v0.2.1 — lively narration, Gemini alignment and dialogue, 28 styles
Narration with feeling and rhythm (user: "声音是对的,但不够活泼,太僵硬")
-
bin/vh ttsaccepts a per-line direction in[brackets], added to the overall--instructfor that line only and kept out of captions. -
--beats map.json [--snap beat|half|downbeat] [--lead s]starts every line on the next grid point instead of a fixed gap.@id:downbeatpins one line, e.g. the answer on the drop. Tested live with Gemini at 120 BPM: lines landed at 0.50 / 4.00 / 5.00 / 10.00 s exactly. -
playbook/04"让声音有表情、有节奏":- how to direct narration, with a default delivery per video type;
- riding the beat;
- frame-aligned tempos at 30 and 24 fps;
- rhythm density per type;
- mixing so music stays present (a measured
duck_ratio=3hole vs a clean1.6); - writing sound as prompts.
Types 02, 03, 04 and 08 point to it.
Gemini narration, further
- Gemini 3.8 Flash TTS is now tested live in Mandarin and English (it was only mock-tested in v0.2.0). Also documented: the default local Qwen3-TTS 0.6B sometimes runs on after short English lines.
--align gemini: every line, whatever the provider, is transcribed bygemini-3.5-transcribefor word timestamps and compared with the script. A line below--min-sim(default 0.85) is flagged, which catches the Qwen 0.6B run-on and skipped words.--vocabpasses terms to the transcriber.- Two-speaker dialogue: an
@speakers A=Kore B=Puckheader, thenA: …lines. Labels only count when declared, so an ordinary "注意:" line stays narration. --join block|allsends consecutive lines as one request, so the delivery flows across sentences; the per-line files are then cut at word timestamps. A Gemini dialogue defaults toblock.bin/vh voices list | design | deletemanages Gemini designed voices.bin/vh captionswrites per-word timing intocaptions.jsonand acaptions.<lang>.lines.srtwith one cue per wrapped line, each starting on its first spoken word.- The key is read from
GEMINI_API_KEYonly; nothing is written to the repo.
Reading time and motion rules
tools/readcheck.py/bin/vh readcheck: on-screen text needs max(2.5 s, CJK ÷ 4.5 + other ÷ 15 + 1.5 s), adapted from lemo-opuscar (MIT); subtitles follow the Netflix ceilings (CJK ≤ 9 chars/s, English ≤ 20) with a 1.8 s floor. Checklist item #5 andplaybook/03use it.playbook/03: four spring registers by element (buttons, cards and camera, big type and logos, mascots), from xilo-opus-video, plus the NaN trap at ζ = 1.playbook/08: sub-frame motion blur must not sample across a cut, and per-subframe sub-pixel jitter gives free anti-aliasing (from abstract-algebra-promo).playbook/02: CJK web fonts on canvas arrive as unicode-range subsets, sodocument.fonts.readycan pass with glyphs missing. Calldocument.fonts.loadonce per weight with every character the film uses, and fail loudly on rejection, an empty result or a timeout.playbook/05: green-screen character + code scene: key and despill with ffmpeg, keep alpha (PNG or ProRes 4444), let code own the scene, type and beat.
Styles: 28
- Two new presets, each with a swatch, QA and determinism check, and two rounds of independent review:
pixel-16bit(Chrono Trigger, A Link to the Past; 320 × 180 integer scaling, a fixed 24-color palette, mosaic transitions; lowest score 7) andisotype(Neurath and Arntz; one symbol = a fixed quantity, added or removed on the beat; lowest score 6). - The deferred fixes from v0.2.0 are done: monumental-scifi has a hook, cutout-jazz and swiss-grid-type no longer share a motif, and the shadow-puppet click is gone.
gallery.jpg/gallery.mp4rebuilt for 28. - lemo-opuscar is now 43 styles and MIT for the whole repo (since 2026-09-29; older snapshots were CC BY 4.0). Presets adapted from the older snapshot keep their CC BY attribution.
References
cases/opus55-gallery.md§7: a deep read of abstract-algebra-promo (sub-frame blur, cut protection, math promo structure); §5 adds the dsxzai catalog.references/fetch.sh: text-only clones now actually pull on update (the empty checkout used to short-circuit it).
Full history: CHANGELOG.md
v0.2.0 — style library, effort hub, intro film
Effort hub: one switch for how hard the agents work
- Three levels:
quick,standard(default) andstudio, defined in one table inCLAUDE.md. The table sets:- the human gates;
- how many styles are offered;
- storyboard depth and self-check depth;
- rounds of independent scoring;
- sound;
- deliverables;
- draft count, subagents, research, and a suggested reasoning effort.
- A floor that never drops at any level: hard rules 1, 5 and 7, no digital silence, flash safety, type minimums, licenses,
bin/vh check. bin/vh new … --effort <level>writesEffort:into BRIEF.md and, forquick, notes the gate waiver in REVIEW.md.bin/vh effort [level]prints the rules.- Only the user can lower the level; an agent may not downgrade to save time.
- The checklist, pipeline, verification and review templates now say what each level does.
Intro film and a rewritten README
showcase/04-intro-film/: the project's own 81 s intro film.- One continuous 3D shot built with HyperFrames + Three.js, and a score composed in code. It shows the architecture, the workflow with its three gates, the features and the case library.
- It was made by agents following this repo and revised after two rounds of human notes: the music stuttered, and it wasn't cool enough. The notes, the look-dev in three intensities and every fix are in the folder.
- README (en / zh) rewritten in plain language, with the intro film as the hero.
- WeChat appreciation code at the end of both READMEs.
Style swatches, reviewed and reworked
- An independent "harsh motion director" scored all 26 swatches on the 7-dimension layer, twice. The median lowest score rose from about 4.5 to 6, but none has all seven scores at 8 or above yet: the library is honest about being work in progress (see
styles/README.md). - Makers' self-scores ran 1–2 points above the independent reviewer's, which is why the checklist insists the scorer is not the author.
- A pre-release fix round after the second review: editorial-data and dark-math got real hooks; ink-wash, silhouette-papercut and brutalist-meme had small rule violations fixed. Deferred to the next version: monumental-scifi (hook), the cutout-jazz / swiss-grid-type look-alike motif, and a shadow-puppet click warning.
styles/_swatch/custom_sfx.pyrebuilds the few custom foley WAVs byte for byte, so nothing shipped is hand-made or downloaded.- Three rounds of fixes followed:
- a hook at 0.1 s;
- no freezes after the motif lands;
- each style draws "draft" in its own language instead of ▶;
- clones separated (cutout vs Swiss, CRT vs FUI, shadow puppet vs paper-cut, three "growing circle" endings);
- the Dunhuang flying apsaras redrawn with proper bodies;
- type sizes raised.
- Foley:
styles/<slug>/events.jsonis mixed under the score.styles/_swatch/foley.mjsgenerates it from aFOLEYexport inswatch.js, so picture and sound share one timing table. The foley track fades out with the score. - The review's systemic findings became rules in the swatch content spec (
styles/_swatch/README.md).
Style library styles/: 26 tastes instead of one
- 26 presets in 6 families (film titles, brand, data, illustration and print, Chinese aesthetics, retro). Each is distilled from famous works: Saul Bass titles, Se7en, Blade Runner 2049, Wes Anderson symmetry, Wong Kar-wai step-printing, Ken Burns, Müller-Brockmann, film FUI, 3Blue1Brown, NYT / The Pudding, Gapminder, Spider-Verse, risograph, CRT terminals, 水墨, 敦煌, 皮影, 国潮, Reiniger's silhouettes, watercolor backgrounds, synthwave, and more.
- Each preset has
STYLE.md(study works, visual / motion / sound grammar, bans, prompt block, engine recipe, self-check),tokens.json, and a real 5 s swatch with its own code-composed score. - All 26 swatches show the same content, so the only difference is the style. Overview:
styles/gallery.jpgandgallery.mp4. - Swatch renderer
styles/_swatch/(HyperFrames): scene API, sharedlib.js(easing, springs, decode text, textures, transitions), per-slug staging, a watchdog, and automatic audio QA. Swatches are deterministic, with lossless frames compared across worker counts. bin/vh style list | <preset> | gallery | check;bin/vh new <type> <slug> --style <preset>seeds the project with the preset.- Gate ① now proposes 2–3 contrasting presets instead of defaulting to one look. Principle: learn the grammar, never copy characters, logos or shots.
Review and verification
- A scored layer on top of the 20-item checklist: a harsh-director reviewer scores 7 dimensions, each must reach ≥ 8, over at least 3 rounds. It catches "correct but not exciting".
- Phone-size (360 px) readability sheets, a loop-seam check, and determinism compared on lossless PNG frames rather than mp4.
- Silent-failure detection: frozen frames (renderAt exceptions), empty frames (YAVG), and render watchdogs.
- Look-dev (2–3 variants of one segment) when feedback is vague; optional full-length animatic at gate ②; authorized skips recorded verbatim.
- Motion: superposed springs, leading and trailing edge stiffness, layout functions for multi-aspect renders, decode-text rules.
Sound
bin/vh tts gemini/gemini-lite: Google Gemini 3.8 Flash TTS and Flash-Lite TTS as cloud providers (GEMINI_API_KEY), with a plain-language delivery direction as the 5th argument. Inline tags such as<short pause>are voiced by Gemini and stripped for every other provider and from captions. Request and parsing tested against a mocked response; not yet called for real.bin/vh beats: addshitsandkick/snareaccents with strength (HPSS plus band-split onsets, within ±1 frame on test loops).bin/vh sfx place: stereo, with per-eventpananddist, so sound follows on-screen position. The built-inerrorSFX loses its hard edges; the library is still 15 sounds.bin/vh music:- Chinese colour layers
bell编钟,zheng古筝 (Karplus–Strong withbend),dizi竹笛 andtaiko大鼓; - pentatonic modes;
- variable
meterswith abarsmap; --example zh.
- Chinese colour layers
bin/vh mix: stereo, two-pass linear loudness (keeps LRA),duck=voiceandduck_ratio. Fix: music was truncated to the voice length when ducking.- New
bin/vh qa: final-mix QA covering digital silence, dropouts, pumping, click warnings to re-listen, and a cue check against the beat map and events. - Rules learned from the intro film: dramatic stops are held breaths, never digital silence; don't key the ducker on every SFX.
Robustness
bin/vh hf-initvendors GSAP locally and warns about any remaining CDN links, because an offline render with a CDN import hangs silently. Showcase 00 and 02 were updated the same way.references/fetch.shneutralizes agent files (.claude,.agents,CLAUDE.md,CLAUDE.local.md,AGENTS.md,.mcp.json) at every depth, not only the repo root. Newfetch.sh <dir>and--neutralizeoptions.bin/vhhelp lists every command.
Knowledge and references
playbook/08: FX preset stack (fx= A/B/C) with a reason per effect, flash and readability limits, true sub-frame motion blur, SFX pan from screen position, and one-take 3D world techniques.engines/README: Three.js + HyperFrames pitfalls.video-types/03: a promo shows the product itself; asset inventory before animating; look-dev step; up to ~90 s.video-types/07: Chinese characters written in true stroke order (Make Me a Hanzi data fetched per project, never vendored).cases/opus55-gallery.md: a second catalog (962 works) and a deep-dive into the Battle of Austerlitz 5-minute WebGL film.cases/community-prompts.md: new community data points.- References grow from 24 to 30 repos: athemeroy's research catalog, claude-animation-skill, product-film-skill, procedural-film, the 962-work catalog, and Battle-of-Austerlitz-Film.
Full history: CHANGELOG.md
v0.1.0 — first public release
Harness
CLAUDE.md/AGENTS.mdrouter over 8 video types, with 7 hard rules and a rule-precedence section.AGENTS.mdis generated fromCLAUDE.md(bin/vh sync-agents).- Three mandatory human review gates (outline → storyboard + keyframe preview sheet → first draft), recorded in
templates/REVIEW.md. playbook/00–08: paradigm and engine choice, 10-stage pipeline with "reads" timing, 7-layer verification, motion-design numbers, audio, hybrid generative video, research mechanisms, reverse-engineering a reference video, VFX and motion sources.templates/: BRIEF, STORYBOARD, STYLE, REVIEW, NOTES, LESSONS, TASTE_CHECKLIST (20 items).cases/: 11 case studies +opus55-gallery.md(curated from 389 community videos).references/community-skills.md: curated community skills, a 39-style library and a license table.
Sound (music · SFX · voice · songs · captions)
bin/vh tts: local open-source Qwen3-TTS by default (Chinese voices Serena, Vivian, Uncle_Fu, Dylan, Eric; English voices Ryan, Aiden), plussay/edge/dashscope/elevenlabs. Bilingual scripts (中文 || English) with per-language timelines.bin/vh captions: zh / en / bilingual SRT +captions.jsonfor engines, with CJK-aware wrapping;bin/vh muxadds soft zh/en subtitle tracks.bin/vh music: deterministic code-composed soundtrack (score.json → WAV + exact beat/section/hit map).bin/vh sfx: 15 original synthesized SFX; events placed so each sound lands on its action.bin/vh mix: voice + music + SFX, music ducks under voice and SFX, −14 LUFS.- Songs: Suno-import workflow; ElevenLabs Music and local song-model interfaces reserved.
CLI bin/vh
doctor,setup,types,new <type> <slug>(templates, prompt block pre-filled, engine scaffolded),hf-init(HyperFrames without global skill installs),sync-agents,install-skill.install.sh: one-line install (clone, deps, references, skill registration for Claude Code and Codex).- QA:
sheet(timestamped contact sheets, written into the owning project),check(black/freeze/silence, 3 s threshold),gif.
References
references/fetch.shfetches 23 read-only repos; media-heavy ones text-only; upstreamCLAUDE.md/.claude//AGENTS.mdrenamed to_upstream_*so they are never auto-loaded as instructions.
Showcase (made by agents following only this harness)
- 00 launch film (HyperFrames, 20 s), 01 hand-drawn short (p5.brush, 12 s), 02 vertical science short (HyperFrames, 24.8 s), 03 Fourier explainer (Manim, 25 s).
Known gaps
- Tested end to end:
qwen(zh + en),say, captions, music, sfx, mix, mux. Implemented from official docs but not yet tested (no keys):edge,dashscope,elevenlabs. Word-level forced alignment and song-generation providers are reserved interfaces. - Workflow docs are Chinese-first; English translations welcome.
Full history: CHANGELOG.md