Skip to content

Roadmap

zhanglinghao edited this page Oct 3, 2026 · 4 revisions

English · 中文

This is the maintainer's plan (@ZLHad) as of 2026-10-03. It changes as work lands, and this page is edited to match. To suggest a change, open an issue. "Later" items are ideas, not promises.

What already works is on Project Status; what changed in each version is in the CHANGELOG.

Now: in progress

  • Try the experimental pieces on the real thing. Type 09 has never cut real footage and nothing it exports has been opened in an editor. The Blender guide is partly verified: Blender 5.2.2 rendered a style swatch and the intro film's opening, but a few of its smoke tests and its own build and render commands haven't been run. No video model has been called for the generated-clip route. A first real run of each will correct their numbers.
  • Listen to the new sound. The 31 swatch scores and the soundtracks of films 00–03 pass every measurement the agents can make. The maintainer has heard the first candidate mixes and the flagged narration spots in 02; nobody has yet judged the final mixes by ear.

Done recently

  • Remake the intro film (v5). 103 s: the first 15.8 s are path-traced in Blender, driven by code (a galaxy of 1.18 million stars in which every star is a film), then v3's one-take world with the counts brought up to date and a wall of the 31 style samples. v3 is kept in showcase/04-intro-film/v3/.
  • Concept before style (#44): playbook 12, concept cards at gate ①, floor items tagged and the other taste rules made defaults, and 立意 as the eighth score; with a case study of five community films (#43), five full-film skeletons that start from a concept (#45) and research note 06, an A/B against floors only (#46).
  • The spec up front (#47): Watch on and Resolution in the BRIEF, text-size floors by where the film is watched, 4K as a render setting.
  • 31 styles: Blender swatch scenes and tabletop-miniature (#42), pastel-ui and y2k-chrome (#49); presets attach as references, not templates (#48).
  • Sound effects that vary: every event its own variant, six new transitions, sfx audition, a repetition warning (#39), and the 28 swatches of the time and films 00–03 re-rendered on them (#40); a note map for pictures driven note by note (#41).
  • Word timings for the narration of showcase films 02 and 03, with 02's "2 GHz" line re-taken (#36); spoken captions checked as subtitles and short cues lengthened (#38).
  • The README's sample films, one per row: the preview, a link to the original with sound, the request, a suggested workflow and the sound (#34).
  • bin/vh <command> -h prints the usage instead of running the command (#50); hf-init takes fonts from the machine (#35); bin/vh check reports the colour format (#37).
  • Genre-true soundtracks for all 28 swatches, on a new music engine: 77 synthesised instruments, a pattern language, stereo and reverb buses (#13), plucked strings as physical models (#20), motifs and dynamics (#24), and the re-scored, re-rendered swatches (#31).
  • Sound for showcase films 00–03: scores, sound effects, and Chinese and English narration for 02 and 03 (#32).
  • Mix profiles and the mix report (#21): the level hierarchy between voice, music and effects, and a harder cue check.
  • Director mode and decision-first review pages (#19), and the pictures to decide from: storyboard pages, a rhythm map, style comparison, cover preview, a piano roll (#25).
  • Shot recipes: 24 recipes, three pacing skeletons, bin/vh recipes (#27).
  • Three new playbook docs (narrative, hooks and packaging, composition; #23); playbook 12, on the concept, followed in #44, so the playbook now runs 00–12.
  • Type 09, editing your own footage (experimental; #30), the Blender guide (experimental; #29) and generated clips reorganised by task (#26).
  • Research on similar projects (agent video production, editing tools, sound studios): done on 2026-09-30. Its findings have been landing as small PRs; type 09 was written from it.
  • Install docs: what gets downloaded and how much, mirror settings for mainland China (#22).
  • Research notes: five lab notes on things measured along the way (#33).
  • Swatch renders are byte-identical again (#17).
  • ElevenLabs word timing (#7): per-character timestamps are grouped into words, the same units --align gemini produces. Not yet tested against the live API.
  • Fixes from a second review of everything merged since v0.2.0 (#14, #15).
  • A one-picture overview in the README (#12).
  • This wiki, refreshed on 2026-10-03 to match main.

Next: v0.3

  • A release that carries what has piled up on main since v0.2.1: the CHANGELOG's Unreleased section is long.
  • Every swatch reaches 8 or more on every review score (today none does; most were scored on seven dimensions, before 立意 was added; see Style Library).
  • Tagged releases: v0.1.0, v0.2.0 and v0.2.1 are published. v* tags are protected by two rulesets: only admins create them, and nobody can move or delete them (CONTRIBUTING.md).
  • 9:16 variants of the showcase films.
  • A Chinese-narration variant of the intro film.
  • The tools type 09's docs describe but that don't exist yet: a gap probe, seam metrics, an edit-list compiler.
  • A type doc for 3D and shader films. Today they follow the promo type plus playbook/08, the Blender guide (partly verified) covers path-traced shots, and the intro film's opening is a worked example; the README roadmap calls this type 10.
  • Linux install polish. CI covers Ubuntu and a first round of fixes is in; the install has been tried on Ubuntu 24.04 only.
  • Community style submissions: a submission template and a review process.

Later: ideas

  • More engines and video types as the community asks for them. Editing your own footage is already started as type 09.
  • A tempo map for bin/vh music (ritardando, accelerando).
  • A hosted screening page for the samples.
  • Packaging the whole workbench as an installable skill or plugin. Today's skill is a thin pointer that installs the repo on first use.
  • Languages beyond Chinese and English.
  • Two more items from the README roadmap: word-by-word highlighted subtitles (per-word timing already lands in captions.json when you use --align gemini), and English versions of the workflow docs.

How to help

Pick an item, say so in an issue so work isn't duplicated, and follow Contributing.

Clone this wiki locally