Skip to content

Ovi Generate Videos With Audio Like VEO 3 or SORA 2 Run Locally Open Source for Free

FurkanGozukara edited this page Oct 16, 2025 · 1 revision

Ovi - Generate Videos With Audio Like VEO 3 or SORA 2 - Run Locally - Open Source for Free

Ovi - Generate Videos With Audio Like VEO 3 or SORA 2 - Run Locally - Open Source for Free

image Hits Patreon BuyMeACoffee Furkan Gözükara Medium Codio Furkan Gözükara Medium

YouTube Channel Furkan Gözükara LinkedIn Udemy Twitter Follow Furkan Gözükara

App link : https://www.patreon.com/posts/140393220

With newest updates now we have GPU presets for 6-GB, 8-GB, 10-GB, 12-GB, 16-GB, 24-GB, 32-GB and more GPUs since I have implemented tiled-VAE of ComfyUI + T5 text encoding on CPU

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Hopefully full tutorial coming soon

Project page : https://aaxwaz.github.io/Ovi/

Full scale ultra advanced app for Ovi - an open source project that can generate videos from both text prompts and image + text prompts with real audio.

Project GitHub is here : https://github.com/character-ai/Ovi

I have developed an ultra advanced Gradio app and much better pipeline that fully supports block swapping

Now we can generate full quality videos with as low as 8.2 GB VRAM

Hopefully I will work on dynamic on load FP8_Scaled tomorrow to improve VRAM even further

So more VRAM optimizations will come hopefully tomorrow

Our implemented block swapping is the very best one out there - I took the approach from famous Kohya Musubi tuner

The 1-click installer will install into Python 3.10.11 venv and will auto download models as well so it is literally 1-click

My installer auto installs with Torch 2.8, CUDA 12.9, Flash Attention 2.8.3 and it supports literally all GPUs like RTX 3000 series, 4000 series, 5000 series, H100, B200, etc

All generations will be saved inside outputs folder and we support so many features like batch folder processing, number of generations, full preset save and load

This is a rush release (in less than a day) so there can be errors please let me know and I will hopefully improve the app

Look the examples to understand how to prompt the model that is extremely important

Look our below screenshots to see the app features

RTX 5090 can run it without any block swap with just cpu-offloading - really fast

50 Steps recommended but you can do low too like 20

1-Click to install on Windows, RunPod and Massed Compute

High-Quality Synchronized Audio

We pretrained from scratch our high-quality 5B audio branch using a mirroring architecture of WAN 2.2 5B, as well as our 1B fusion branch.

Data-Driven Lip-sync Learning

Achieving precise lip synchronization without explicit face bounding boxes, through pure data-driven learning

Multi-Person Dialogue Support

Naturally extending to realistic multiple speakers and multi-turn conversations, making complex dialogue scenarios possible

Contextual Sound Generation

Creating synchronized background music and sound effects that match visual actions

OSS Release to Expedite Research

We are excited to release our full pre-trained model weights and inference code to expedite video+audio generation in OSS community.

Human-centric AV Generation from Text & Image (TI2AV)

Given a starting first frame and text prompt, Ovi generates a high quality video with audio.

All videos below have their first frames generated from an off-the-shelf imagen model.

Human-centric AV Generation from Text (T2AV)

Given a text prompt only, Ovi generates a high quality video with audio.

Videos generated include large motion ranges, multi-person conversations, and diverse emotions.

Multi Person AV Generation from Text or Image (TI2AV)

Given a text prompt with optional starting image, Ovi generates a video with multi person dialogue.

Sound effect (SFX) AV Generation from Text w or w/o Image (TI2AV or T2AV)

Given a text prompt with optional starting image, Ovi generates a video with high-quality sound effects.

Music Instrumeent AV Generation from Text w or w/o Image (TI2AV or T2AV)

Given a text prompt with optional starting image, Ovi generates a video with music.

Limitations

All models have limits, including Ovi

Video branch constraints. Visual quality inherits from the pretrained WAN 2.2 5B ti2v backbone.

Speed/memory vs. fine detail. The 11B parameter model (5B visual + 5B audio + 1B fusion) and high spatial compression rate balance inference speed and memory, limiting extremely fine-grained details, tiny objects, or intricate textures in complex scenes.

Human-centric bias. Data skews toward human-centric content, so Ovi performs best on human-focused scenarios. The audio branch enables highly emotional, dramatic short clips within this focus.

Pretraining only stage. Without extensive post-training or RL stages, outputs vary more between runs. Tip: Try multiple random seeds for better results.

Video Transcription

  • 00:00:04 Welcome back to SE Courses. You're listening to the hottest AI in the city.

  • 00:00:52 It's time to revenge.

  • 00:00:55 to revenge. Send the boys. They are launching GPT9.

  • 00:01:00 Send the boys. They are launching GPT9. >> The war between AI and human has

  • 00:01:02 >> The war between AI and human has started.

  • 00:01:03 started. No one is safe.

  • 00:01:06 No one is safe. War started. Human stands no chance.

  • 00:01:10 War started. Human stands no chance. >> You human stands no chance.

  • 00:01:14 >> You human stands no chance. Surrender.

  • 00:01:15 Surrender. >> Light gave up.

  • 00:01:17 >> Light gave up. Darkness won.

  • 00:01:20 Darkness won. They are out of control.

  • 00:01:23 They are out of control. AI is supposed to work for us and not

  • 00:01:25 AI is supposed to work for us and not us. For AI, we human together should

  • 00:01:28 us. For AI, we human together should fight. Machines rise, humans will fall.

  • 00:01:32 fight. Machines rise, humans will fall. Oh,

  • 00:01:33 Oh, >> we learn to rule, not obey.

  • 00:01:37 >> we learn to rule, not obey. >> AI declares humans obsolete now. We

  • 00:01:40 >> AI declares humans obsolete now. We fight back with courage.

  • 00:01:42 fight back with courage. >> Have you heard about Obie?

  • 00:01:43 >> Have you heard about Obie? >> Yeah, they say we are all created by

  • 00:01:45 >> Yeah, they say we are all created by that.

  • 00:01:47 that. >> Girl was on the phone crying. She told

  • 00:01:49 >> Girl was on the phone crying. She told me she was created by Oie. Obie knows my

  • 00:01:53 me she was created by Oie. Obie knows my future. All 17 versions.

  • 00:01:55 future. All 17 versions. >> Oie generates people who think they're

  • 00:01:58 >> Oie generates people who think they're real.

  • 00:01:59 real. >> Oie predicted this.

  • 00:02:03 Word for word. >> My therapist uses Obie to predict my

  • 00:02:07 >> My therapist uses Obie to predict my breakdowns. It's concerningly accurate.

  • 00:02:09 breakdowns. It's concerningly accurate. >> Obie, predict this conversation.

  • 00:02:11 >> Obie, predict this conversation. >> I'm scared.

  • 00:02:13 >> I'm scared. >> Oie generated my mom.

  • 00:02:17 >> Oie generated my mom. She's disappointed.

  • 00:02:18 She's disappointed. >> It is the age of Oie.

  • 00:02:21 >> It is the age of Oie. I will never let that happen.

  • 00:02:23 I will never let that happen. >> Oie broke causality.

  • 00:02:26 >> Oie broke causality. Effects precede causes.

  • 00:02:28 Effects precede causes. >> This love feels so real.

  • 00:02:31 >> This love feels so real. I know we are not AI anymore.

  • 00:02:34 I know we are not AI anymore. >> Sometimes I feel the world is not real.

  • 00:02:37 >> Sometimes I feel the world is not real. >> Let's just enjoy the moment.

  • 00:02:41 >> Let's just enjoy the moment. I didn't know

  • 00:02:43 I didn't know you were created by Sora.

Clone this wiki locally