Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
-
Updated
Aug 5, 2026 - Python
Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
Open-source, local-first video editor where creators and AI agents edit the same real timeline.
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · Whisper · Chatterbox · MuseTalk.
[CVPR-2025] The official code of HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
DialogLab is an authoring tool for configuring and running Human-AI multi-party conversations.
[NeurlPS-2024] The official code of MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
Explore the power of Azure Text-to-Speech with interactive talking avatar, Lisa 👩🏻🦱. Choose from multiple languages and avatar styles to bring your text to life.
ComfyUI custom nodes for LongCat Video Avatar 1.5 audio-driven human video generation; a macOS inference branch, macOS-MLX, is now available for Apple Silicon MLX testing.
Interactable AI that have control over your frontend website, It guides your user walk around your website its a salesman / supports.
Talking Avatar: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
Budget-aware AI agent content studio for creating videos, images, carousels, voiceovers, music, captions, and content calendars with consistent brand assets and cost control.
AI Avatar/Anchor: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
The text to speech avatar system is a text to speech feature with vision capabilities, that allow customers to create synthetic videos of a 2D photorealistic avatar speaking. The Neural text to speech Avatar models are trained by deep neural networks based on the human video recording samples, and the voice of the avatar .
Animated Characters: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
The official main page of "EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion".
Audio-driven talking avatar and lip-sync video generation from a reference image and speech audio.
Give your AI assistant a face — a WebGL particle talking-head (~800k dots) that speaks with your own local Claude. Unofficial community project · CC0.
These are output demos of different models for Talking Avatar Generation (TAG).
Talking avatars created using Leonardo.ai, VoiceOverMaker, and D-ID.
Add a description, image, and links to the talking-avatar topic page so that developers can more easily learn about it.
To associate your repository with the talking-avatar topic, visit your repo's landing page and select "manage topics."