Skip to content
Discussion options

You must be logged in to vote

You can definitely build your own avatar and integrate it with your existing STT, LLM, and TTS APIs without relying on expensive avatar platforms.

The important thing to understand is that lip-sync isn't controlled by a specific audio file format. It's usually handled through audio analysis or phoneme-to-viseme mapping (converting speech sounds into mouth shapes).

Here's how I'd approach it:

  1. Create your avatar

For 3D, use Blender and export as .glb, .gltf, or .vrm with facial blendshapes.

For 2D, consider Live2D or a custom rigged avatar with mouth animations.

  1. Implement lip-sync

You have two main options:

Audio amplitude analysis: Analyze the TTS audio and animate the mouth based on v…

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by karthiksreenivasanp
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
Question Ask and answer questions about GitHub features and usage Programming Help Discussions around programming languages, open source and software development source:ui Discussions created via Community GitHub templates
2 participants