JoyAI-Video-Edit is a real-time, instruction-guided video editing system for open-ended video streams. Given a live camera stream or uploaded video and a natural-language edit instruction, it edits frames causally as they arrive, without waiting for the full video, requiring a predefined video length, or revisiting future frames. In our deployment benchmark, the full end-to-end pipeline reaches 30.19 FPS at 720x1280, pushing video editing from offline batch processing toward interactive streaming generation.
The system combines an MLLM-based condition encoder, a causal video VAE, and a 16B-parameter multimodal diffusion transformer. It is trained and deployed as an autoregressive diffusion editor, then accelerated with aligned autoregressive distribution matching distillation, long-horizon optimization, bounded KV-state inference, and deployment-oriented scheduling to sustain high-throughput 720p editing while reducing train-inference mismatch and accumulated temporal drift.
- 2026.08.05: 🎉 We release the deployment code, technical report, online demo, and JoyAI-Video-Edit checkpoints. Please check the links above for details.
- Real-time open-ended editing. Edits live or uploaded videos as frames arrive, without requiring the full sequence upfront.
- Diverse instruction control. Supports subject edits, local edits, background changes, style transfer, motion changes, and reference-guided editing.
- Autoregressive diffusion design. Combines an MLLM condition encoder, causal video VAE, and MMDiT backbone for streaming video editing.
- High-throughput 720p deployment. Reaches 30.19 FPS end-to-end throughput at 720x1280 with bounded KV-state inference and stable per-chunk compute.
- Consumer GPU support. Optimize deployment for consumer-grade GPUs such as GeForce RTX 5090.
- Stronger model version in progress. A more powerful version is under active development, with a particular focus on advancing reference-image-guided video editing (RV2V) capabilities.
- Release full training and data pipelines. Open-source the complete training framework and data generation pipeline.
JoyAI-Video-Edit is designed for broad video editing tasks, including global appearance changes, local object edits, subject add/remove/replace, background replacement, style transfer, and reference-guided edits.
demo.mp4
Download the released JoyAI-Video-Edit weights from Hugging Face, then place them under:
deploy/deps/checkpoints/JoyAI-Video-Edit/
|-- dit/
| `-- joyai_video_edit_dit_0804.pth
`-- vae/
|-- config.json
`-- diffusion_pytorch_model.safetensors
conda create -n joyai-video-edit python=3.10 -y
conda activate joyai-video-edit
python -m pip install -r requirements.txtDownload the released weights from the Hugging Face link above. MiMo-VL and the ONNX detector files are external runtime dependencies; see DEPLOYMENT.md for deployment details.
cd deploy
bash run_server.shThen open:
http://localhost:8080
For remote machines, bind the server to 0.0.0.0 and open the selected port, or use SSH port forwarding.
For custom deployment, edit deploy/run_server.sh to set checkpoint paths, CUDA device placement, host, and port. The default script also sets persistent TorchInductor, Triton, and CUDA cache directories so compile artifacts are reused across launches.
TORCHINDUCTOR_AUTOGRAD_CACHE is not required for inference-only serving.
If JoyAI-Video-Edit is useful for your research or product prototype, please cite:
@article{xiao2026joyai,
title={JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion},
author={Xiao, Yicheng and Dai, Wenxun and Qin, Xinran and Song, Lin and Zhang, Maoquan and Xu, Hang and Chen, Yukang and Li, Yitong and Zhang, Guohui and Zhang, Yuan and Zhang, Xuying and Zhang, Tommy and Yuan, Jianlong and Li, Peihao and Lu, Shuai and Fu, Siming and Zhao, Chuyang and Han, Xin and Huang, Jie and Li, Wenbo and Ma, Guoqing and Huang, Wei and Qi, Xiaojuan and Huang, Haoyang and Duan, Nan},
journal={arXiv preprint arXiv:2608.03974},
year={2026}
}JoyAI-Video-Edit is licensed under Apache 2.0.










