Skip to content

Repository files navigation

AudioVisGen: Video Generation Using Audio

Abstract

This project explores the development of an audio-to-video generation model that integrates Con- trastive Audio-Visual Pretraining (CAVP) with existing frameworks to enhance video generation from direct audio inputs. Motivated by the need to improve semantic coherence and temporal synchronization in generated videos, our approach leverages this integration to address the challenges of translating com- plex audio cues into visually coherent outputs. While the model showed improvements in aligning audio with video content, significant discrepancies were observed between quantitative metrics and qualitative assessments. These discrepancies highlight not only the limitations of current evaluation metrics but also underscore the inherently subjective nature of art and visual content, which can vary widely in individual perception. This project points to the necessity for developing more comprehensive evaluation metrics that better align with human perception and acknowledge the subjective experience of art. The findings lay the groundwork for future research in enhancing audio-to-video synthesis models, emphasizing the need for larger datasets, expanded computational resources, and new evaluation methods to truly gauge the effectiveness and appeal of generated video content.

Model Structure

Screenshot 2024-05-05 at 18 19 23

Example Outputs

volleyball mix_dog_prompt dog_prompt

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages