Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

voicestyle

Guess the appearance of an unseen person's face from their unheard voice. A voice-based generation model. Corresponding paper: TOMM-VoiceStyle

Inference

Install CMPC
In order to make it more convenient to use CMPC, we referred to the approach of CLIP: added some methods and provided external APIs.
Finally, we encapsulated it in the form of a Python package cmpc2
pip install git+https://github.com/audio-visual/cmpc2.git

Download checkpoints and necessary files
https://drive.google.com/drive/folders/1U77L9kx95HrpMPBnJGMVb4ShUr3iTR26?usp=drive_link
stylegan's checkpoint 550000.pt: https://github.com/rosinality/stylegan2-pytorch
id loss releated checkpoint model_ir_se50.pth: https://drive.google.com/file/d/1KW7bjndL3QG3sxBbZxreGHigcCCpsDgn/view
put these files to the weights folder

Full pipeline

voice --> (1) find k nearest image prototypes --> (2) get the corresponding image paths
  ↓                                                        ↓
(7) compute CMPC similarity loss               (3) project images to StyleGAN's latent space                                 
  ↑                                                        ↓
(6) generated image <-- (5) feed init to generator <-- (4) aggregate these latent code as init code 

Generate

> cd generation
# we uploaded two examples and their releated files, so you can go through the full pipeline
> python pipleline.py

Notice that the first stage is training and testing on images with a resolution of 224x224. However, the StyleGAN(pytorch version) was trained based on 256x256 images. Therefore, in order to obtain better results, we chose to retrain a 256-size version of the CMPC model. (Of course, you can also use the checkpoint of the original 224-size CMPC model and perform upsampling operations in the code)

Important things about GAN inversion(step3):
Regarding the GAN inversion, we have conducted experiments in advance to test its performance. For the testing process, please refer to: https://github.com/audio-visual/investigate_stylegan.
The test results show that the low-quality, excessively large head angle, and misaligned images in the voxceleb dataset (although they can be aligned through post-processing, but due to low resolution, the post-processing results are very bad) are not suitable for gan inversion.

Therefore, we adopted a compromise solution, that is, replacing the faces in voxceleb videos with the same faces in vggface(filterd by https://github.com/cmu-mlsp/reconstructing_faces_from_voices, named VGG_ALL_FRONTAL). However, this also has limitation, that is, vggface does not have video-level tags, and the appearance, makeup, and expressions of the same person will vary greatly (as shown in the below). This may reduce model performance.

Examples from VGG_ALL_FRONTAL 'Luke_Mitchell' image

input voice example1:

00001.mp4

You can try to imagine it first :)

generated face real face

|


input voice example2:

00001.mp4
generated face real face

|

Citation

@article{10.1145/3671002,
author = {Chen, Wuyang and Zhu, Boqing and Xu, Kele and Dou, Yong and Feng, Dawei},
title = {VoiceStyle: Voice-Based Face Generation via Cross-Modal Prototype Contrastive Learning},
year = {2024},
issue_date = {September 2024},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
volume = {20},
number = {9},
issn = {1551-6857},
url = {https://doi.org/10.1145/3671002},
doi = {10.1145/3671002},
journal = {ACM Trans. Multimedia Comput. Commun. Appl.},
month = sep,
articleno = {279},
numpages = {23},
keywords = {Cross-modal representation learning, contrastive learning, StyleGAN, face generation}
}

About

Guess the appearance of a person's face from their voice. A voice-based image generation Model

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages