A curated, visually-organized list of Text-to-Video (T2V) products, open-source models, research papers, datasets, and benchmarks.
- What's New
- Commercial Text-to-Video Products
- Open-Source Models & Toolkits
- Research Papers
- Datasets & Benchmarks
- Contributing
- License
- References
2026 Update: The T2V landscape has shifted dramatically. OpenAI discontinued the consumer Sora app in early 2026, while open-source models (Wan 2.2, HunyuanVideo 1.5, LTX-2.3) and commercial alternatives (Runway Gen-4.5, Seedance 2.0, Kling 3.0, Veo 3.1) now dominate.
Data collected as of June 2026. Prices, features, and availability change quickly — always check the official website before making a decision.
Key trends:
- Native 4K and synchronized audio are becoming standard.
- Open-source models now rival closed commercial systems.
- Multi-shot storytelling and character consistency are the new battlegrounds.
- Inference acceleration (sliding tile attention, token carving, caching) makes local generation feasible.
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🎬 Runway Gen-4 / Gen-4.5 | RunwayML | Professional creators | 1080p (4K upscale) | @reference consistency, world-class physics, Aleph editing |
runwayml.com |
| 🐉 Kling 3.0 | Kuaishou | Motion-heavy / cinematic | 1080p, 15 s | Best-in-class motion, generous free tier | klingai.com |
| 🔍 Veo 3 / Veo 3.1 | Google DeepMind | 4K broadcast production | Native 4K | Scene extension, native audio + lip-sync | deepmind.google |
| 🌱 Seedance 2.0 | ByteDance | Multi-shot storytelling | 2K, 60 s multi-shot | 12 mixed inputs, native audio-video joint gen | seedance.tv |
| ✨ Luma Dream Machine | Luma AI | Action / sports / physics | 4K | Realistic motion blur, fluid dynamics | lumalabs.ai |
| 🎭 Pika 2.5 / 3.0 | Pika Labs | Social / stylized content | 2K | Fast, cheap, strong style transfer | pika.art |
| 🌀 Hailuo AI | MiniMax | Realistic humans / prompt adherence | 1080p, 10 s | Strong physical realism, #1 in China | hailuoai.video |
| 🎬 PixVerse V6 | PixVerse | Anime / stylized content | 1080p, 15 s | Character consistency engine, 20+ camera controls, native audio | pixverse.ai |
| 🎥 Vidu Q1/Q2 | Shengshu / Tsinghua | Highly consistent T2V | 1080p, 16 s | U-ViT backbone, subject consistency, 1080p generation | vidu.com |
| 🔬 Lumiere | Google DeepMind | Research T2V / I2V / editing | 720p | Space-time U-Net, single-pass temporal generation | lumiere-video.github.io |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🧑💼 Synthesia | Synthesia | Corporate training / avatars | 1080p | 100+ avatars, 130+ languages | synthesia.io |
| 🎤 DeepBrain AI | DeepBrain AI | Hyper-realistic avatars | 1080p | PPT-to-video, chroma key, native AI anchors | deepbrain.io |
| 🎥 HeyGen | HeyGen | AI avatars / translation | 1080p | 100+ avatars, multilingual voice clone, video translation | heygen.com |
| 🏢 Colossyan | Colossyan | Corporate training / e-learning | 1080p | SCORM export, branching quizzes, dual-avatar scenes | colossyan.com |
| ⚡ Elai.io | Elai | Document-to-video automation | 1080p | URL/PPTX-to-video, 100+ languages, ~1.5 min rendering | elai.io |
| 🖼️ D-ID | D-ID | Talking photos / short outreach | 1080p | Animate any photo, live streaming API, cheapest entry | d-id.com |
| 📺 Hour One | Hour One | Enterprise presentations | 1080p | Broadcast-ready virtual presenters, clean UI | hourone.ai |
| 🧑🏫 Vidnoz AI | Vidnoz | Budget avatar videos / training | 1080p | Cheap avatar generation, multilingual voiceovers | vidnoz.com |
| 🎨 Steve.AI | Steve.AI | Animated explainer videos | 4K | Multiple animation styles, character library | steve.ai |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🛠️ Krea AI | Krea | Model-agnostic access | Varies | Unified UI for Kling, Hailuo, Luma, Runway, Pika | krea.ai |
| 🌀 Morphic / Morph Studio | Morph Studio | All-in-one creative platform | Varies | 15+ models, canvas editing, custom model training | morphic.com |
| ✂️ InVideo AI | InVideo | Social media / marketing | 1080p | 5000+ templates, AI script-to-video workflow | invideo.io |
| 🎬 YumCut | YumCut | Self-hosted vertical shorts | 9:16 | Script, voice, visuals, captions, and automation API | yumcut.com · GitHub |
| 🎩 Magic Hour | Magic Hour | Multi-format creative suite | 1080p | Face swap, talking photos, headshots, clothes swapper | magichour.ai |
| 🛒 Creatify | Creatify | UGC-style ad generation | 1080p | E-commerce focused, ad performance tracking | creatify.ai |
| 🎞️ Vivideo | Vivideo | Model-agnostic short-video creation | Varies | Unified access to multiple T2V/I2V models, synced audio, text- and image-to-video | vivideo.ai |
| 🎞️ cv.cm/v (Cloud Clipboard AI Studio) | Cloud Clipboard | Seedance-based video workflows | Varies | Queue-free Seedance 2.0, image generation, canvas, and short-drama agent | cv.cm/v |
| 🎞️ Pixo | Pixo | Story-driven / storyboard-to-film | Varies | Story idea → storyboard → scene-by-scene generation → finished video, agent-assisted editing | pixo.video |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| Consumer app shut down March 2026 | |||||
| Consumer product shut down in 2026 |
|
Ranked #1 on Artificial Analysis T2V benchmark (1,247 Elo). Native text-to-video with character-locking |
Best-in-class motion quality, 66 free credits/day, and strong temporal consistency for cinematic content. Try Kling → |
|
First true native 4K T2V model with scene extension to 60+ seconds and synchronized lip-sync audio. Try Veo → |
ByteDance's multi-shot native audio-video generator. Up to 12 mixed inputs, 60 fps, timeline prompting. Try Seedance → |
|
SCORM-compliant corporate training with branching quizzes, dual-avatar conversations, and 80+ languages. Try Colossyan → |
Creative multi-format suite: face swap, talking photos, headshots, and clothes swapper for social content. Try Magic Hour → |
|
Specialized in anime and stylized content with character consistency, 20+ camera controls, and native audio. Try PixVerse → |
All-in-one creative platform with 15+ models, canvas-based editing, and custom model training for teams. Try Morphic → |
| Tool | What it does | Link |
|---|---|---|
| Podframes | Generates two-host AI podcast videos with scripted dialogue, voices, lip-sync, and captions. | github.com/Jellypod-Inc/podframes |
| TubePrompter | Converts existing videos into optimized text-to-video prompts for Sora, Veo, Runway, etc. | tubeprompter.com |
| Vadoo AI | AI shorts automation platform for faceless channels and social clips. | vadoo.tv |
| Omni-Rewriter | Open agentic prompt-expansion harness for image/video model dialects (schema + validation + bounded repair; expand ≠ generate). | github.com/WayneJin0918/Omni-Rewriter |
| Your Goal | Recommended Tool | Why |
|---|---|---|
| Professional filmmaking / VFX | Runway Gen-4.5 | Best creative control, reference consistency, Aleph editing |
| Best motion realism on a budget | Kling 3.0 | Top-tier motion, generous free tier, native audio |
| 4K broadcast / cinematic scenes | Veo 3.1 | Native 4K, scene extension, lip-sync audio |
| Multi-shot storytelling | Seedance 2.0 | Unified audio-video generation, timeline prompting |
| Anime / stylized content | PixVerse V6 | Character consistency, camera controls, native audio |
| Corporate training / LMS | Synthesia / Colossyan | SCORM, avatars, quizzes, multilingual |
| Avatar marketing videos | HeyGen / D-ID | Expressive avatars, fast generation, API access |
| Real-time creative exploration | Krea AI | 64+ models, sub-50ms feedback |
| Local / open-source deployment | Wan 2.2 / HunyuanVideo 1.5 | Apache 2.0; Wan 2.2 offers a 5B model for 24 GB GPUs |
Get started with the most popular open-source models:
Wan 2.2
git clone https://github.com/Wan-Video/Wan2.2
cd Wan2.2
pip install -r requirements.txtHunyuanVideo 1.5
git clone https://github.com/Tencent-Hunyuan/HunyuanVideo
cd HunyuanVideo
# Follow the official guide for sampling and I2V inferenceCogVideoX
git clone https://github.com/THUDM/CogVideo
cd CogVideo
pip install -r requirements.txtFor detailed inference configs, LoRA fine-tuning, and ComfyUI workflows, see each repository's README.
- Efficient Video Diffusion Models: Advancements and Challenges, Shitong Shao. [Paper]
- First comprehensive survey focused on efficient video diffusion: step distillation, efficient attention, model compression, caching/trajectory optimization.
- Controllable Video Generation: A Survey [Paper]
- Systematic review of pose-guided, structure-controlled, and other conditional video generation methods.
- Survey of Video Diffusion Models: Foundations, Implementations, and Applications, Yimu Wang et al. [Paper]
- Comprehensive taxonomy of diffusion-based video generation, evaluation metrics, industry solutions, and training engineering.
- UNIC: Unified In-Context Video Editing [Paper]
- BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios [Paper]
- Runway Gen-4.5 Technical Report — Runway (2026) [Runway Help]
- ByteDance Seedance 2.0 — Global launch April 2026 [Overview]
- LTX-2.3: Native 4K Video + Audio Generation — Lightricks (March 2026) [Project]
- Bridging Text and Video Generation: A Survey, Nilay Kumar et al. [Paper]
- Comprehensive survey on T2V evolution, DiT architectures, datasets, training recipes, and benchmarks.
- Wan 2.2: Mixture-of-Experts Video Generation — Alibaba (July 2025) [Project]
- Wan: Open and Advanced Large-Scale Video Generative Models, Alibaba. [Paper] [Code]
- Training a Commercial-Level Video Generation Model in $200k (Open-Sora 2.0) [Paper] [Code]
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model, Guoqing Ma et al. [Paper] [Code]
- HunyuanVideo: A Systematic Framework For Large Video Generative Models, Weijie Kong et al. [Paper] [Code]
- Fast Video Generation with Sliding Tile Attention, Peiyuan Zhang et al. [Paper]
- Training-Free Efficient Video Generation via Dynamic Token Carving [Paper]
- FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge [Paper]
- On-Device Sora: Enabling Training-Free Diffusion-Based Text-to-Video Generation for Mobile Devices [Paper]
- SkyReels-A2: Compose Anything in Video Diffusion Transformers [Paper]
- MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation [Paper]
- EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer [Paper]
- MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models [Paper]
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness [Paper] [Code]
- Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation [Paper]
- ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation [Paper]
- T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models [Paper]
- A Survey of AI-Generated Video Evaluation [Paper]
- Movie Gen: A Cast of Media Foundation Models, Meta. [Paper]
- Lumiere: A Space-Time Diffusion Model for Video Generation, Google DeepMind. [Paper] [Project] [Code]
- Vidu: A Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models, Shengshu / Tsinghua. [Paper] [Code]
- Allegro: Open the Black Box of Commercial-Level Video Generation Model, Rhymes AI. [Paper] [Code]
- DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors, CUHK / Tencent AI Lab. [Paper] [Code]
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models, Yixin Liu et al. [Paper]
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models, Haoxin Chen et al. [Paper] [Code]
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. [Paper] [Code]
- Latte: Latent Diffusion Transformer for Video Generation, Xin Ma et al. [Paper] [Code]
- VideoTetris: Towards Compositional Text-to-Video Generation [Paper]
- MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators [Paper]
- T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-Video Generation [Paper]
- CPA: Camera-pose-awareness Diffusion Transformer for Video Generation [Paper]
- Open-Sora: Democratizing Efficient Video Production for All [Project]
- Open-Sora-Plan: Open-Source Reproduction of Sora [Project]
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. [Paper]
- ModelScope Text-to-Video Technical Report, Jiayu Wang et al. [Paper]
- I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models, Alibaba / DAMO. [Paper] [Code]
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. [Paper] [Code]
- Scalable Diffusion Models with Transformers (DiT), William Peebles & Saining Xie. [Paper] — Foundation for diffusion-transformer architectures used in Sora and successors.
- Video Generation Models as World Simulators, OpenAI (Sora technical report). [Report]
- Text-To-4D Dynamic Scene Generation, Uriel Singer et al. [Paper] [Project]
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation, Jay Zhangjie Wu et al. [Paper] [Project] [Code]
- MagicVideo: Efficient Video Generation With Latent Diffusion Models, Daquan Zhou et al. [Paper] [Project]
- Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. [Paper]
- Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. [Paper] [Project]
- Text-driven Video Prediction, Xue Song et al. [Paper]
- Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. [Paper] [Project] [Code]
- StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation, Adyasha Maharana et al. [Paper] [Code]
- Word-Level Fine-Grained Story Visualization, Bowen Li et al. [Paper] [Code]
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers, Wenyi Hong et al. [Paper] [Code]
- Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning, Yogesh Balaji et al. [Paper] [Code] [Project]
- Video Diffusion Models, Jonathan Ho et al. [Paper] [Project]
- Transcript to Video: Efficient Clip Sequencing from Texts, Ligong Han et al. [Paper] [Project]
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions, Chenfei Wu et al. [Paper]
- Text2Video: Text-driven Talking-head Video Synthesis with Phonetic Dictionary, Sibo Zhang et al. [Paper]
- TiVGAN: Text to Image to Video Generation With Step-by-Step Evolutionary Generator, Doyeon Kim et al. [Paper]
- Conditional GAN with Discriminative Filter Generation for Text-to-Video Synthesis, Yogesh Balaji et al. [Paper] [Code]
- IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation, Kangle Deng et al. [Paper]
- StoryGAN: A Sequential Conditional GAN for Story Visualization, Yitong Li et al. [Paper] [Code]
- Video Generation From Text, Yitong Li et al. [Paper]
- To create what you tell: Generating videos from captions, Yingwei Pan et al. [Paper]
| Dataset | Scale | Highlights | Link |
|---|---|---|---|
| WebVid-10M | 10.7M clips | Large-scale text-video pairs scraped from the web | m-bain/webvid |
| InternVid | 7M+ clips | High-quality video-text dataset with multimodal annotations | OpenGVLab/InternVid |
| HD-VILA-100M | 100M clips | High-resolution long-form videos with dense captions | microsoft/HQD |
| Panda-70M | 70M clips | Large dataset of high-quality video-caption pairs | snap-research/Panda-70M |
| Vript | 420K clips | Ultra-detailed captions (~145 words), shot type and camera movement | Vript |
| MiraData | 798K clips | Long-duration videos (avg 72s), 1080p, 318-word captions | MiraData |
| OpenVid-1M | 1M clips | High-quality, diverse scenarios, ~98-word captions | OpenVid-1M |
| CI-VID | 1M clips / 717K seqs | Coherent interleaved text-video dataset | CI-VID |
| HD-VG-130M | 130M clips | High-definition, watermark-free video-text pairs | HD-VG-130M |
| VidProM | — | Large prompt-gallery dataset for video generation | VidProM |
| Benchmark | Focus | Link |
|---|---|---|
| VBench / VBench-2.0 | Comprehensive video generation evaluation suite | Vchitect/VBench |
| T2V-CompBench | Compositional text-to-video generation | T2V-CompBench |
| T2VTextBench | Human evaluation of textual control in video generation | T2VTextBench |
| ChronoMagic-Bench | Metamorphic / time-lapse text-to-video evaluation | ChronoMagic-Bench |
| BRITE | Reliable T2V evaluation on implausible scenarios | arXiv:2605.00873 |
| VideoEval | Low-cost evaluation of video foundation models | VideoEval |
| VideoScore2 | Think-before-you-score reward model / metric for generated video: visual quality, T2V alignment & physical consistency with chain-of-thought (ships VideoScore-Bench-v2) | Paper · Code |
Contributions are welcome! Please open a pull request to add new products, papers, models, datasets, or benchmarks. Keep entries concise and include both a paper link and a code/project link when available.
If you have any questions, feel free to contact jianzhnie.
Awesome-Text-To-Video is released under the Apache 2.0 license.
This list was compiled and updated using information from the following sources.
- OpenAI Sora
- Google DeepMind Veo
- Runway • Runway Help: Creating with Gen-4 Video • Runway Review 2026
- Luma Dream Machine
- Kling AI
- Hailuo AI
- Pika
- Krea AI
- ByteDance Seedance • Seedance 2.0 Global Release Guide • Seedance 2.0 Review
- PixVerse
- Morphic / Morph Studio
- Hedra
- Vidnoz AI
- Steve.AI
- Vidu
- Lumiere Project Page
- Synthesia
- HeyGen
- DeepBrain AI
- Colossyan
- Elai.io
- D-ID
- Hour One
- Creatify
- InVideo AI
- Magic Hour
- Wan 2.1
- HunyuanVideo
- CogVideo
- Open-Sora
- Open-Sora-Plan
- LTX-Video
- Mochi
- Step-Video-T2V
- SkyReels-V2
- MAGI-1
- Waver
- VideoCrafter
- VGen / ModelScope T2V
- Allegro
- DynamiCrafter
- I2VGen-XL
- VideoComposer
- Lumiere PyTorch
- AnimateDiff
- Latte
- VBench
- T2V-CompBench
- T2VTextBench
- ChronoMagic-Bench
- VideoEval
- arXiv papers cited in the Research Papers section.