Repository navigation
Releases: Edakuban/VocaVid
Releases · Edakuban/VocaVid
Release list
v1.0.0
VocaVid 0.1.0 – First public release
VocaVid is a local, storyboard-based workflow for AI music videos using ComfyUI.
This first public release covers the entire process:
- Create a project with WAV files, lyrics, style, and references
- Analyze lyrics using Whisper and automatically segment them
- Manually correct timing and segments as needed
- Generate a scene plan as well as image and motion prompts
- Generate images, avatar variants, and video clips
- Review, approve, and selectively re-render clips
- Export a Kdenlive project and the final MP4
- Analyze reel candidates and export them vertically
Important
- VocaVid runs locally and requires Python, ComfyUI, FFmpeg, and a CUDA-capable GPU.
- Models and custom nodes depend on the workflow used. The included workflow JSONs serve as the authoritative reference.
- The project is still under active development. Feedback, bug reports, and suggestions for improvement are welcome.
More information and documentation:
https://vocavid.de/