Welcome to Video Generation papers!
Table of Contents
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-08-20 | 4DAnyone: Create Anyone in 4D from a Casual Monocular Video | Yudong Jin et.al. | 2608.20335 | null |
| 2026-08-20 | DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery | Yufei Liu et.al. | 2608.20308 | null |
| 2026-08-20 | AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures | Guoxing Sun et.al. | 2608.19900 | null |
| 2026-08-20 | VGI-BENCH: Probing Visual Intelligence in Video Generation Models | Xuan He et.al. | 2608.19583 | null |
| 2026-08-20 | VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation | Yinming Huang et.al. | 2608.18607 | null |
| 2026-08-19 | Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models | Pardis Taghavi et.al. | 2608.18484 | null |
| 2026-08-18 | SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation | Keyu Tu et.al. | 2608.17426 | null |
| 2026-08-18 | SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering | Gen Li et.al. | 2608.17420 | null |
| 2026-08-17 | DriveCache: Action-Aware Caching for Driving World Model Inference | Jianchun Yang et.al. | 2608.16354 | null |
| 2026-08-17 | AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model | Kwan Yun et.al. | 2608.16143 | null |
| 2026-08-16 | RigidBench: Evaluating Rigid-Body Physics in Video Generation Models | Swarnim Jain et.al. | 2608.15555 | null |
| 2026-08-16 | Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention | Shengchuan Gao et.al. | 2608.15522 | null |
| 2026-08-17 | Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation | Lunjie Zhu et.al. | 2608.13602 | null |
| 2026-08-13 | SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation | Jisoo Jeong et.al. | 2608.13460 | null |
| 2026-08-13 | HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models | Jiazi Bu et.al. | 2608.13205 | null |
| 2026-08-13 | H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models | Dingyi Rong et.al. | 2608.13049 | null |
| 2026-08-13 | From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion | Xichen Ye et.al. | 2608.13043 | null |
| 2026-08-12 | Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence | Aman Tyagi et.al. | 2608.12290 | null |
| 2026-08-11 | VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics | Bowei Liu et.al. | 2608.11201 | null |
| 2026-08-11 | Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning | Liangyu Fu et.al. | 2608.11013 | null |
| 2026-08-12 | Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation | Guixu Lin et.al. | 2608.10479 | null |
| 2026-08-11 | Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation | Yueting Zhu et.al. | 2608.10439 | null |
| 2026-08-10 | Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning | Haodong Li et.al. | 2608.09926 | null |
| 2026-08-10 | GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction | Khang Minh Le et.al. | 2608.09493 | null |
| 2026-08-10 | CodecArena: Codec Quality Assessment via Visual Reinforcement Learning | Jiaye Fu et.al. | 2608.09139 | null |
| 2026-08-10 | RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement | Ziheng Jia et.al. | 2608.09111 | null |
| 2026-08-07 | MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation | Youjun Zhao et.al. | 2608.07463 | null |
| 2026-08-07 | CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition | Novendra Setyawan et.al. | 2608.06691 | null |
| 2026-08-06 | GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions | Chenghao Gu et.al. | 2608.06332 | null |
| 2026-08-06 | HOPE: Hand-Object Pressure Estimation from Monocular Videos | Subin Jeon et.al. | 2608.06192 | null |
| 2026-08-06 | Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training | Rui Li et.al. | 2608.06125 | null |
| 2026-08-06 | Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features | Sining Ang et.al. | 2608.06008 | null |
| 2026-08-06 | Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model | Haoning Yang et.al. | 2608.05976 | null |
| 2026-08-07 | Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models | Haodong Yan et.al. | 2608.05903 | null |
| 2026-08-06 | Vorch-Omni: Multi-Task Orchestration of Sight and Sound | Vorch Team et.al. | 2608.05803 | null |
| 2026-08-05 | In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion | Lingxiao Yang et.al. | 2608.05237 | null |
| 2026-08-05 | HelloWorld: Enabling Socially Interactive Characters in Video World Models | Liangyang Ouyang et.al. | 2608.05070 | null |
| 2026-08-05 | UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models | Haiyang Zhou et.al. | 2608.04701 | null |
| 2026-08-05 | RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation | Shuliang He et.al. | 2608.03387 | null |
| 2026-08-04 | SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference | Shanghao Liu et.al. | 2608.03335 | null |
| 2026-08-04 | FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection | Pei Li et.al. | 2608.03096 | null |
| 2026-08-03 | WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity | Yuxue Yang et.al. | 2608.02603 | null |
| 2026-08-03 | Faster-WAM: Do World Action Models Need Deep Action Modules? | Liheng Ma et.al. | 2608.02365 | null |
| 2026-08-03 | UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation | Liming Tan et.al. | 2608.01944 | null |
| 2026-08-02 | InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos | Chi Zhang et.al. | 2608.01157 | null |
| 2026-08-04 | MiniWorld: Democratizing the Training of Video World Models from Scratch | Yian Zhao et.al. | 2608.01127 | null |
| 2026-08-01 | DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents | Tongsheng Ding et.al. | 2608.00486 | null |
| 2026-08-04 | Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh | Junhao Chen et.al. | 2608.00094 | null |
| 2026-07-30 | Mirror Learning | Yunpeng Liu et.al. | 2607.28737 | null |
| 2026-07-30 | Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion | Henglin Liu et.al. | 2607.28058 | null |
| 2026-07-30 | Articulated Object Reconstruction from Rest-State Observation | Daeun Lee et.al. | 2607.27749 | null |
| 2026-07-29 | VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System | Haodong Li et.al. | 2607.27380 | null |
| 2026-08-03 | FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring | Jiatong Li et.al. | 2607.27110 | null |
| 2026-07-29 | Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory | Yanbo Ding et.al. | 2607.26818 | null |
| 2026-07-29 | TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models | Taewon Kang et.al. | 2607.26706 | null |
| 2026-07-29 | Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation | Yongxin Su et.al. | 2607.26646 | null |
| 2026-07-29 | ContactFlow: A video action conditioning that transfers across embodiments | Sami Azirar et.al. | 2607.26579 | null |
| 2026-07-29 | CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling | Yuyang Huang et.al. | 2607.26529 | null |
| 2026-07-28 | WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models | Jiamin Xu et.al. | 2607.26203 | null |
| 2026-07-28 | Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics | Timy Phan et.al. | 2607.25984 | link |
| 2026-07-28 | Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography | Hyunkyung Han et.al. | 2607.25748 | null |
| 2026-07-29 | I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models | Yimao Guo et.al. | 2607.25522 | null |
| 2026-07-28 | Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision | Ruijie Su et.al. | 2607.25321 | null |
| 2026-07-27 | MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention | Jianlin Yu et.al. | 2607.24377 | null |
| 2026-07-29 | FilmBench: A Film-Grade Benchmark for Cinematic Video Generation | Shengyi Wang et.al. | 2607.24241 | null |
| 2026-07-27 | DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning | Mengqi Zhang et.al. | 2607.24159 | link |
| 2026-07-27 | ViDS: Video Diffusion Shader using 3D Face Tracking | Wenbo Ji et.al. | 2607.24124 | null |
| 2026-07-26 | OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models | Zhaoyuan He et.al. | 2607.23844 | null |
| 2026-07-26 | VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation | Tianxiao Chen et.al. | 2607.23472 | null |
| 2026-07-25 | CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion | Shreshth Saini et.al. | 2607.23159 | null |
| 2026-07-24 | Generative Video Compression with Adaptive Score Distillation | Naifu Xue et.al. | 2607.22772 | null |
| 2026-07-17 | MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion | Jiacheng Liu et.al. | 2607.22696 | null |
| 2026-07-24 | AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment | Ziyao Huang et.al. | 2607.22241 | null |
| 2026-07-27 | Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering | Wenchao Ma et.al. | 2607.21848 | null |
| 2026-07-23 | Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers | Sicheng Mo et.al. | 2607.21594 | null |
| 2026-07-23 | GraphVid: Interactive Graph-Controllable Video Generation | Vedant Shah et.al. | 2607.21580 | null |
| 2026-07-23 | Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention | Zekun Li et.al. | 2607.20940 | null |
| 2026-07-22 | HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation | Jinliang Shen et.al. | 2607.20125 | null |
| 2026-07-21 | Geospatial Diffusion-based Evolution Synthesis (GeoDES) for Storm-Centered Weather Augmentation | Sonia Cromp et.al. | 2607.19522 | null |
| 2026-07-21 | FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling | Jialong Zuo et.al. | 2607.19038 | null |
| 2026-07-21 | Learning Explicit Physical Parameter Control and Benchmarking for Video Generation | Yanxun Li et.al. | 2607.18924 | null |
| 2026-07-21 | DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking | Yunyi Li et.al. | 2607.18664 | null |
| 2026-07-20 | AniGS: Bridging Rendering and Diffusion Prior for 3D Scene Animation | Yen-Chi Cheng et.al. | 2607.18539 | null |
| 2026-07-20 | ShotPlan: Cinematic Video Generation with Learnable Planning Token | Su Guo et.al. | 2607.17675 | null |
| 2026-07-20 | Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection | Amir Hosein Fadaei et.al. | 2607.17625 | null |
| 2026-07-20 | Thinking in Video: Can Video Generators Really Reason About the Real World? | Yongheng Zhang et.al. | 2607.17523 | null |
| 2026-07-19 | Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models | Xingkai Peng et.al. | 2607.17279 | null |
| 2026-07-19 | The generator is the tracker: Multi-object tracking by painting persistent identity colours | Haiyu Yang et.al. | 2607.17120 | null |
| 2026-07-17 | Apple- |
Runmao Yao et.al. | 2607.16401 | link |
| 2026-07-17 | PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation | Qirui Li et.al. | 2607.16355 | null |
| 2026-07-17 | Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation | Dimitrios Karageorgiou et.al. | 2607.15849 | null |
| 2026-07-17 | PE-Field 4D: Video Generation Models as Canvas | Yunpeng Bai et.al. | 2607.15667 | null |
| 2026-07-16 | From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting | Zizhao Chen et.al. | 2607.14976 | null |
| 2026-07-16 | CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation | Yuanpeng Zhang et.al. | 2607.14908 | null |
| 2026-07-16 | FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers | Minguk Kang et.al. | 2607.14898 | null |
| 2026-07-16 | ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships | Xinyu Liu et.al. | 2607.14681 | null |
| 2026-07-16 | MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation | Yinhan Zhang et.al. | 2607.14595 | null |
| 2026-07-15 | VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders | Zhihao Xie et.al. | 2607.14088 | null |
| 2026-07-15 | From Pixels to States: Rethinking Interactive World Models as Game Engines | Zhen Li et.al. | 2607.14076 | null |
| 2026-07-15 | Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data | Thang-Anh-Quan Nguyen et.al. | 2607.13927 | null |
| 2026-07-15 | VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation | Songyu Xu et.al. | 2607.13527 | null |
| 2026-07-14 | Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization | Yuxin Huang et.al. | 2607.13336 | null |
| 2026-07-14 | Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation | Ruize Xia et.al. | 2607.13164 | null |
| 2026-07-14 | The Seriality Gap in Video Diffusion Models | Jorge Diaz Chao et.al. | 2607.13031 | null |
| 2026-07-16 | ACID: Adaptive Caching for vIDeo generation | Om Agrawal et.al. | 2607.12358 | null |
| 2026-07-13 | Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model | Xinghang Li et.al. | 2607.11643 | null |
| 2026-07-13 | Video Transformer for Remote Identity Document Hologram Detection | Joris Voerman et.al. | 2607.11419 | null |
| 2026-07-10 | 4D Human-Scene Reconstruction from Low-Overlap Captures | Minhyuk Hwang et.al. | 2607.09125 | link |
| 2026-07-10 | Video Generation Models are General-Purpose Vision Learners | Letian Wang et.al. | 2607.09024 | null |
| 2026-07-09 | LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models | Cheng-De Fan et.al. | 2607.08770 | link |
| 2026-07-09 | OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators | Hongyu Liu et.al. | 2607.08766 | null |
| 2026-07-09 | OpenCoF: Learning to Reason Through Video Generation | Xinyan Chen et.al. | 2607.08763 | link |
| 2026-07-09 | HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales | Wenbo Xu et.al. | 2607.08705 | null |
| 2026-07-09 | Native Video-Action Pretraining for Generalizable Robot Control | Qihang Zhang et.al. | 2607.08639 | null |
| 2026-07-15 | LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting | Zixin Guo et.al. | 2607.08016 | null |
| 2026-07-08 | Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence | Shuailei Ma et.al. | 2607.07675 | null |
| 2026-07-07 | Gen4U: Unifying Video Generation and Understanding via Diffusion | Michael King et.al. | 2607.06856 | null |
| 2026-07-07 | Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation | Yu Cheng et.al. | 2607.06631 | null |
| 2026-07-07 | ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation | Ruihang Zhang et.al. | 2607.06555 | null |
| 2026-07-07 | FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration | Jin Jiang et.al. | 2607.06389 | null |
| 2026-07-07 | MobileWan: Closing the Quality Gap for Mobile Video Diffusion | Mohsen Ghafoorian et.al. | 2607.06173 | null |
| 2026-07-07 | RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures | Hanan Gani et.al. | 2607.06018 | null |
| 2026-07-06 | MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing | Gal Fiebelman et.al. | 2607.05376 | null |
| 2026-07-06 | InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization | Haoxiang Ma et.al. | 2607.04988 | null |
| 2026-07-06 | Video Generation Models Are Inherent Lighting Estimators | Ziqi Cai et.al. | 2607.04674 | null |
| 2026-07-08 | Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising | Guangting Zheng et.al. | 2607.04653 | null |
| 2026-07-05 | Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption | Nidhal Jegham et.al. | 2607.04553 | null |
| 2026-07-04 | Reward Lightning: Fast Video Generation via Homologous Preference Distillation | Jiaxiang Cheng et.al. | 2607.03960 | null |
| 2026-07-04 | ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics | Zanwei Zhou et.al. | 2607.03732 | null |
| 2026-07-03 | Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model | Xinyin Ma et.al. | 2607.03509 | null |
| 2026-07-03 | Vidu S1: A Real-Time Interactive Video Generation Model | Jintao Zhang et.al. | 2607.03118 | null |
| 2026-07-03 | Natural Language Camera Movement Understanding | Yuwen Tan et.al. | 2607.03043 | null |
| 2026-07-02 | Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation | Long Vu et.al. | 2607.02798 | null |
| 2026-07-01 | RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation | Yaofu Liu et.al. | 2607.02584 | null |
| 2026-07-02 | QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers | Kyobin Choo et.al. | 2607.01869 | link |
| 2026-07-02 | ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning | Xuanhua He et.al. | 2607.01677 | null |
| 2026-07-01 | Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models | Yue Han et.al. | 2607.01222 | link |
| 2026-07-06 | Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences | Zhenjia Li et.al. | 2607.00832 | null |
| 2026-07-01 | RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail | Amirreza Rouhi et.al. | 2607.00310 | null |
| 2026-06-29 | Vertigo Vertigo: Reconstructing a Cinematic Ideal through its Predictive AI Double | Adam Cole et.al. | 2607.00047 | null |
| 2026-06-30 | MemLearner: Learning to Query Context memory for Video World Models | Jiwen Yu et.al. | 2606.31734 | null |
| 2026-06-30 | InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving | Xiaoyu Ye et.al. | 2606.31109 | null |
| 2026-06-30 | Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors | Zixing Wang et.al. | 2606.31101 | null |
| 2026-06-29 | Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding | Seongro Yoon et.al. | 2606.30611 | null |
| 2026-06-29 | The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction | Yuxi Wang et.al. | 2606.30308 | null |
| 2026-06-26 | Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors | Youzhi Liu et.al. | 2606.28237 | null |
| 2026-06-26 | PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation | Peiwen Zhang et.al. | 2606.28128 | link |
| 2026-07-02 | TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL | Jing Wang et.al. | 2606.28016 | null |
| 2026-06-26 | SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models | Ruoyu Wang et.al. | 2606.27741 | null |
| 2026-07-02 | MemoBench: Benchmarking World Modeling in Dynamically Changing Environments | Haoyu Chen et.al. | 2606.27537 | null |
| 2026-06-25 | DnA: Denoising Attention for Visual Tasks | Ron Campos et.al. | 2606.27372 | null |
| 2026-06-25 | PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation | Kexu Cheng et.al. | 2606.26916 | null |
| 2026-06-25 | NaviCache: Test-Time Self-Calibration Caching for Video Generation | Zheqi Lv et.al. | 2606.26795 | link |
| 2026-06-23 | Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover | Théo Biardeau et.al. | 2606.26151 | null |
| 2026-06-24 | MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation | JoungBin Lee et.al. | 2606.26087 | null |
| 2026-06-24 | PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling | Zhijie Zheng et.al. | 2606.25430 | null |
| 2026-06-24 | Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation | Atin Pothiraj et.al. | 2606.25306 | link |
| 2026-06-23 | FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation | Orest Kupyn et.al. | 2606.24876 | null |
| 2026-06-30 | TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration | Yang Zhou et.al. | 2606.24336 | link |
| 2026-06-23 | Trimming the Long-Tail of Visual World Modeling Evaluation | Bingxuan Li et.al. | 2606.24256 | null |
| 2026-06-23 | Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models | Xin Wang et.al. | 2606.24152 | null |
| 2026-06-24 | Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation | Yitong Li et.al. | 2606.23743 | null |
| 2026-06-22 | Vera: A Layered Diffusion Model for Content-Preserving Video Editing | Hongkai Zheng et.al. | 2606.23610 | null |
| 2026-06-22 | SteerVTE: Seamless Video Text Editing with Style and Glyph Control | Kai Zeng et.al. | 2606.23254 | null |
| 2026-06-22 | Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation | Yu Cao et.al. | 2606.22918 | null |
| 2026-06-21 | Generative Relightable Avatars | Kunwar Maheep Singh et.al. | 2606.22718 | null |
| 2026-06-21 | Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars | Ramesh Nandipalli et.al. | 2606.22494 | null |
| 2026-06-21 | Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition | Prajwal Gatti et.al. | 2606.22416 | link |
| 2026-06-20 | CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation | Wenhu Zhang et.al. | 2606.21982 | null |
| 2026-06-18 | World Action Models: A Survey | Qiuhong Shen et.al. | 2606.20781 | null |
| 2026-06-15 | GEOPHYS: The Geometry of Physical Plausibility | Christian Internò et.al. | 2606.20707 | null |
| 2026-06-18 | DataMagic: Transforming Tabular Data into Data Insight Video | Yupeng Xie et.al. | 2606.20388 | null |
| 2026-06-18 | Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models | Haoxuan Wu et.al. | 2606.20310 | null |
| 2026-06-17 | Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models | Tianyi Xiang et.al. | 2606.20233 | null |
| 2026-06-17 | LooseControlVideo: Directorial Video Control using Spatial Blocking | Shariq Farooq Bhat et.al. | 2606.19495 | null |
| 2026-06-17 | Physics-IQ Verified | Tim Rädsch et.al. | 2606.18943 | link |
| 2026-06-17 | UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation | Lin Zhang et.al. | 2606.18702 | link |
| 2026-06-17 | APT: Atomic Physical Transitions for Causal Video-Language Understanding | Shang Wu et.al. | 2606.18586 | null |
| 2026-06-16 | Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation | Siyi Chen et.al. | 2606.18478 | null |
| 2026-06-16 | MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model | Lichen Bai et.al. | 2606.17800 | null |
| 2026-06-15 | SierpinskiCam: Camera-Controlled Video Retaking with Sierpinski Triangle Pattern Cues | Suttisak Wizadwongsa et.al. | 2606.17310 | null |
| 2026-06-15 | Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering | Rohit Kundu et.al. | 2606.17257 | null |
| 2026-06-15 | Revealing Artifacts via Noise Amplification: A Novel Perspective for AI-Generated Video Detection | Renxi Cheng et.al. | 2606.16742 | null |
| 2026-06-16 | PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory | Shuai Yang et.al. | 2606.16449 | link |
| 2026-06-14 | Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation | Jingyu Li et.al. | 2606.15869 | link |
| 2026-06-13 | CausalDrive: Real-time Causal World Models for Autonomous Driving | Tianyi Yan et.al. | 2606.15341 | null |
| 2026-06-12 | Toward Richer Material Generation via Procedural Data Enhancement | Yunchen Yu et.al. | 2606.14988 | null |
| 2026-06-12 | CausalMotion: Structured Physical Reasoning as Keyframe and Trajectory Guidance for Training-Free Video Generation | Sihan Zhuang et.al. | 2606.14317 | link |
| 2026-06-12 | VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling | Xunzhi Xiang et.al. | 2606.14162 | null |
| 2026-06-16 | CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation | Sharath Girish et.al. | 2606.13768 | null |
| 2026-06-13 | RepWAM: World Action Modeling with Representation Visual-Action Tokenizers | Junke Wang et.al. | 2606.13674 | link |
| 2026-06-13 | Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction | Jen-Hao Cheng et.al. | 2606.13655 | null |
| 2026-06-11 | ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance | Salaheldin Mohamed et.al. | 2606.13304 | null |
| 2026-06-11 | TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment | Yu Meng et.al. | 2606.13035 | link |
| 2026-06-10 | Making Foresight Actionable: Repurposing Representation Alignment in World Action Models | Lu Qiu et.al. | 2606.12217 | null |
| 2026-06-10 | World Model Self-Distillation: Training World Models to Solve General Tasks | Sebastian Stapf et.al. | 2606.12072 | null |
| 2026-06-10 | VICX: Generalizable Robot Manipulation via Video Generation and In-Context Operator Network | Song Chen et.al. | 2606.12028 | null |
| 2026-06-09 | AnimaSpark: A Feed-Forward Method for Animating Arbitrary 3D Objects | Yiming Zhao et.al. | 2606.10988 | null |
| 2026-06-10 | BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression | Shaohao Rui et.al. | 2606.10135 | null |
| 2026-06-11 | CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation | Yuheng Chen et.al. | 2606.09639 | null |
| 2026-06-08 | CP4D: Compositional Physics-aware 4D Scene Generation | Hanxin Zhu et.al. | 2606.09187 | null |
| 2026-06-15 | Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions | Luxury et.al. | 2606.09150 | link |
| 2026-06-08 | MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation | Ishaan Preetam Chandratreya et.al. | 2606.09056 | null |
| 2026-06-05 | CULTURESCORE: Evaluating Cultural Faithfulness in Video Generation Models | Anku Rani et.al. | 2606.07311 | null |
| 2026-06-05 | MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models | Yifan Xu et.al. | 2606.06853 | null |
| 2026-06-04 | Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them | Woojung Han et.al. | 2606.06361 | null |
| 2026-06-04 | RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling | Chensheng Dai et.al. | 2606.06309 | link |
| 2026-06-03 | The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show | Parsa Esmati et.al. | 2606.05328 | null |
| 2026-06-04 | Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? | Rui Zhao et.al. | 2606.04811 | link |
| 2026-06-03 | Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control | Jihoon Hong et.al. | 2606.04775 | null |
| 2026-06-03 | Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment | Cong Wang et.al. | 2606.04737 | null |
| 2026-06-03 | DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation | Thanh-Tung Le et.al. | 2606.04432 | null |
| 2026-06-02 | Video-Mirai: Autoregressive Video Diffusion Models Need Foresight | Yonghao Yu et.al. | 2606.03971 | null |
| 2026-06-02 | PointAction: 3D Points as Universal Action Representations for Robot Control | Mutian Tong et.al. | 2606.03943 | link |
| 2026-06-02 | Inference-Time Scaling for Joint Audio-Video Generation | Jaemin Jung et.al. | 2606.03183 | link |
| 2026-06-05 | Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting Reproduction | Yufan Zhang et.al. | 2606.02919 | null |
| 2026-06-01 | RoboDream: Compositional World Models for Scalable Robot Data Synthesis | Junjie Ye et.al. | 2606.02577 | link |
| 2026-06-01 | VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization | Junhao Cheng et.al. | 2606.02564 | null |
| 2026-06-01 | Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization | Jingyun Liang et.al. | 2606.02000 | null |
| 2026-06-01 | Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs | Sicheng Xu et.al. | 2606.01620 | null |
| 2026-06-01 | Effective Multi-sensor Conditioning for Street-view Novel-view Synthesis | Zhengfei Kuang et.al. | 2606.01590 | null |
| 2026-06-01 | MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics | Žiga Kovačič et.al. | 2606.01538 | null |
| 2026-05-31 | AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance | Xilong Zhou et.al. | 2606.01362 | null |
| 2026-05-31 | Knowledge-Intensive Video Generation | Chenxu Wang et.al. | 2606.01285 | link |
| 2026-05-31 | ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning | Xuchen Liu et.al. | 2606.01205 | null |
| 2026-05-30 | SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models | Ziheng He et.al. | 2606.00664 | null |
| 2026-05-30 | Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models | Jinyang Du et.al. | 2606.00658 | null |
| 2026-05-30 | OptiWorld: Optimal Control for Video World Generation under Physical Constraints | Yu Yuan et.al. | 2606.00499 | null |
| 2026-05-29 | Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion | Jiayi Wu et.al. | 2606.00299 | null |
| 2026-05-29 | Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction | Mungyeom Kim et.al. | 2605.31595 | link |
| 2026-05-29 | DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory | Zhenhao Yang et.al. | 2605.31336 | link |
| 2026-05-29 | SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation | Weijia Dou et.al. | 2605.31033 | null |
| 2026-05-28 | DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution | Hidir Yesiltepe et.al. | 2605.30431 | null |
| 2026-05-28 | AdaState: Self-Evolving Anchors for Streaming Video Generation | Yusuf Dalva et.al. | 2605.30349 | null |
| 2026-05-28 | YoCausal: How Far is Video Generation from World Model? A Causality Perspective | You-Zhe Xie et.al. | 2605.30346 | link |
| 2026-05-28 | Benchmarking Single-Factor Physical Video-to-Audio Generation | Tingle Li et.al. | 2605.30339 | null |
| 2026-05-28 | Veda: Scalable Video Diffusion via Distilled Sparse Attention | Shihao Han et.al. | 2605.30325 | null |
| 2026-05-28 | minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models | Min Zhao et.al. | 2605.30263 | null |
| 2026-05-28 | LiveSVG: Zero-Shot SVG Animation via Video Generation | Matan Levy et.al. | 2605.30174 | null |
| 2026-05-28 | SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation | Zhuguanyu Wu et.al. | 2605.30116 | link |
| 2026-05-28 | LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation | Mohammad Khoshnazar et.al. | 2605.29864 | null |
| 2026-05-27 | HarmoVid: Relightful Video Portrait Harmonization | Jun Myeong Choi et.al. | 2605.28811 | null |
| 2026-05-27 | OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning | Yunyang Ge et.al. | 2605.28691 | null |
| 2026-05-27 | DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving | Chen Shi et.al. | 2605.28544 | null |
| 2026-05-27 | Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization | Gaurav Rai et.al. | 2605.28394 | null |
| 2026-05-27 | Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation | Mariam Hassan et.al. | 2605.28230 | null |
| 2026-05-27 | Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models | Haozhan Shen et.al. | 2605.28132 | null |
| 2026-05-27 | SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control | Zhida Zhang et.al. | 2605.27891 | null |
| 2026-05-27 | Turning Video Models into Generalist Robot Policies | Sizhe Lester Li et.al. | 2605.27817 | link |
| 2026-05-26 | What-If World: A Causal Benchmark for General World Models in Embodied Scenarios | Kunlin Cai et.al. | 2605.27589 | null |
| 2026-05-26 | Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation | Unggi Lee et.al. | 2605.26918 | null |
| 2026-05-25 | Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion | Tuna Tuncer et.al. | 2605.26266 | null |
| 2026-05-28 | Paris 2.0: A Decentralized Diffusion Model for Video Generation | Ali Rouzbayani et.al. | 2605.26064 | null |
| 2026-05-25 | Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models | Yiwei Xie et.al. | 2605.25941 | null |
| 2026-05-25 | Full-4D: Generating Full-Scope 4D Scenes from a Single-View Video | Tingxi Chen et.al. | 2605.25500 | null |
| 2026-05-24 | DeltaCam: Differential Intrinsic Camera Modeling for Video Generation | Debabrata Mandal et.al. | 2605.25266 | null |
| 2026-05-24 | Tempered Self-Similarity Alignment for Physically Plausible Video Generation | Manjin Kim et.al. | 2605.24962 | null |
| 2026-05-23 | AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models | Jialiang Yang et.al. | 2605.24652 | null |
| 2026-05-23 | DexSIM: Real-time Dexterous Simulation with Unified Causal Video Diffusion | Adam Lee et.al. | 2605.24630 | null |
| 2026-05-23 | Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation | Ofir Abramovich et.al. | 2605.24509 | null |
| 2026-05-22 | LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation | Bo Jiang et.al. | 2605.23878 | null |
| 2026-05-22 | SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models | Zizhao Tong et.al. | 2605.23345 | link |
| 2026-05-22 | SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion | Xinyu Chen et.al. | 2605.23245 | null |
| 2026-05-21 | MotiMotion: Motion-Controlled Video Generation with Visual Reasoning | Lee Hsin-Ying et.al. | 2605.22818 | link |
| 2026-05-21 | WorldKV: Efficient World Memory with World Retrieval and Compression | Jung Yi et.al. | 2605.22718 | link |
| 2026-05-20 | Q-ARVD: Quantizing Autoregressive Video Diffusion Models | Siao Tang et.al. | 2605.21072 | null |
| 2026-05-20 | Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning | Zhangchi Hu et.al. | 2605.20961 | link |
| 2026-05-20 | FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching | Jangho Park et.al. | 2605.20910 | null |
| 2026-05-20 | What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing | Hangyu Lin et.al. | 2605.20795 | null |
| 2026-05-20 | Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models | Taesung Kwon et.al. | 2605.20624 | null |
| 2026-05-19 | CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition | Hongji Yang et.al. | 2605.19995 | null |
| 2026-05-19 | Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls | Abdul Mohaimen Al Radi et.al. | 2605.19728 | null |
| 2026-05-19 | Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention | Wenhu Zhang et.al. | 2605.19726 | link |
| 2026-05-19 | SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution | Yiren Song et.al. | 2605.19319 | link |
| 2026-05-19 | PhyWorld: Physics-Faithful World Model for Video Generation | Pu Zhao et.al. | 2605.19242 | null |
| 2026-05-18 | Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos | Yuqi Tang et.al. | 2605.18984 | null |
| 2026-05-20 | Spectral Progressive Diffusion for Efficient Image and Video Generation | Howard Xiao et.al. | 2605.18736 | link |
| 2026-05-19 | NEWTON: Agentic Planning for Physically Grounded Video Generation | Yuxiang Feng et.al. | 2605.18396 | null |
| 2026-05-18 | GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation | Jan Ackermann et.al. | 2605.18365 | null |
| 2026-05-18 | Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos | X. Feng et.al. | 2605.18233 | null |
| 2026-05-18 | AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training | Yucheng Guo et.al. | 2605.17923 | null |
| 2026-05-20 | Temporal Aware Pruning for Efficient Diffusion-based Video Generation | Sheng Li et.al. | 2605.17837 | null |
| 2026-05-17 | SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection | Zixi Wei et.al. | 2605.17311 | null |
| 2026-05-16 | 3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation | Hwidong Kim et.al. | 2605.16795 | link |
| 2026-05-15 | AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling | Ziyang Mai et.al. | 2605.16649 | null |
| 2026-05-15 | Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion | Kunyang Li et.al. | 2605.16579 | null |
| 2026-05-15 | SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation | Ssharvien Kumar Sivakumar et.al. | 2605.16530 | link |
| 2026-05-14 | Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance | Stelio Bompai et.al. | 2605.16420 | null |
| 2026-05-15 | Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation | Mingqiang Wu et.al. | 2605.16003 | link |
| 2026-05-15 | Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization | Xiaoxuan He et.al. | 2605.15980 | link |
| 2026-05-14 | Video Models Can Reason with Verifiable Rewards | Tinghui Zhu et.al. | 2605.15458 | null |
| 2026-05-14 | Sound Sparks Motion: Audio and Text Tuning for Video Editing | AmirHossein Naghi Razlighi et.al. | 2605.15307 | null |
| 2026-05-14 | RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO | Yanzuo Lu et.al. | 2605.15190 | link |
| 2026-05-14 | Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video | Yifan Wang et.al. | 2605.15182 | link |
| 2026-05-14 | Compositional Video Generation via Inference-Time Guidance | Ariel Shaulov et.al. | 2605.14988 | null |
| 2026-05-14 | SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer | Zheng Hui et.al. | 2605.14894 | null |
| 2026-05-14 | MechVerse: Evaluating Physical Motion Consistency in Video Generation Models | Rahul Jain et.al. | 2605.14843 | link |
| 2026-05-14 | Probing into Camera Control of Video Models | Chen Hou et.al. | 2605.14815 | null |
| 2026-05-14 | Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity | Jiahao Tian et.al. | 2605.14487 | null |
| 2026-05-14 | CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL | Zhenyang Ni et.al. | 2605.14274 | null |
| 2026-05-13 | TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion | Nurislam Tursynbek et.al. | 2605.14136 | null |
| 2026-05-13 | RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data | Harold Haodong Chen et.al. | 2605.13775 | null |
| 2026-05-13 | AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation | Yuchao Gu et.al. | 2605.13724 | null |
| 2026-05-13 | GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion | Hanxin Zhu et.al. | 2605.12957 | null |
| 2026-05-12 | GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization | Huiran Duan et.al. | 2605.12431 | null |
| 2026-05-12 | From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation | Yajie Li et.al. | 2605.12167 | link |
| 2026-05-12 | Single-Shot HDR Recovery via a Video Diffusion Prior | Chinmay Talegaonkar et.al. | 2605.11628 | link |
| 2026-05-11 | PhyGround: Benchmarking Physical Reasoning in Generative World Models | Juyi Lin et.al. | 2605.10806 | link |
| 2026-05-11 | Progressive Photorealistic Simplification | Adi Rosenthal et.al. | 2605.10409 | null |
| 2026-05-11 | SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation | Liangyang Ouyang et.al. | 2605.10079 | link |
| 2026-05-10 | Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models | Yicheng Ji et.al. | 2605.09681 | link |
| 2026-05-10 | Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study | Yuhan Wang et.al. | 2605.09622 | null |
| 2026-05-10 | SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation | Shanwen Tan et.al. | 2605.09442 | link |
| 2026-05-09 | CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models | Joowon Kim et.al. | 2605.08735 | link |
| 2026-05-09 | Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation | Shihao Cheng et.al. | 2605.08729 | null |
| 2026-05-08 | SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models | Jiesong Lian et.al. | 2605.07800 | link |
| 2026-05-08 | Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers | Jingyuan Zhu et.al. | 2605.07503 | null |
| 2026-05-08 | Do Joint Audio-Video Generation Models Understand Physics? | Zijun Cui et.al. | 2605.07061 | link |
| 2026-05-07 | ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation | Omar El Khalifi et.al. | 2605.06667 | link |
| 2026-05-07 | Relit-LiVE: Relight Video by Jointly Learning Environment Video | Weiqing Xiao et.al. | 2605.06658 | link |
| 2026-05-07 | FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction | Fangda Chen et.al. | 2605.06509 | null |
| 2026-05-07 | Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models | Nilaksh et.al. | 2605.06388 | null |
| 2026-05-07 | EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields | Zhaoyang Yang et.al. | 2605.06192 | null |
| 2026-05-07 | CFE-PPAR: Compression-friendly encryption for privacy-preserving action recognition leveraging video transformers | Haiwei Lin et.al. | 2605.05692 | null |
| 2026-05-05 | Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation | Bin Wu et.al. | 2605.03849 | null |
| 2026-05-05 | Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback | Edoardo Bianchi et.al. | 2605.03848 | null |
| 2026-05-11 | AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics | Tencent HY Team et.al. | 2605.03652 | null |
| 2026-05-05 | Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing | Zhiyuan Li et.al. | 2605.03637 | null |
| 2026-05-04 | Video Generation with Predictive Latents | Yian Zhao et.al. | 2605.02134 | link |
| 2026-05-03 | Exploring Data-Free LoRA Transferability for Video Diffusion Models | Yuchen Wang et.al. | 2605.01929 | link |
| 2026-05-06 | SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages | Sen Fang et.al. | 2605.01720 | link |
| 2026-05-03 | Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models | Martin Q. Ma et.al. | 2605.01662 | null |
| 2026-05-02 | Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion | Jeffrin Sam et.al. | 2605.01477 | null |
| 2026-04-25 | Latent Space Probing for Adult Content Detection in Video Generative Models | Alizishaan Khatri et.al. | 2605.00874 | null |
| 2026-05-01 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors | Houyuan Chen et.al. | 2605.00658 | link |
| 2026-05-01 | Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling | Sen Cui et.al. | 2605.00412 | null |
| 2026-04-30 | PhyCo: Learning Controllable Physical Priors for Generative Motion | Sriram Narayanan et.al. | 2604.28169 | link |
| 2026-04-30 | MotuBrain: An Advanced World Action Model for Robot Control | MotuBrain Team et.al. | 2604.27792 | link |
| 2026-04-30 | ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control | Yanghao Zhou et.al. | 2604.27711 | null |
| 2026-05-07 | Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising | Jun Guo et.al. | 2604.26694 | link |
| 2026-04-29 | DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation | Junhu Fu et.al. | 2604.26232 | null |
| 2026-04-28 | A Systematic Post-Train Framework for Video Generation | Zeyue Xue et.al. | 2604.25427 | null |
| 2026-04-28 | HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation | Bingzi Zhang et.al. | 2604.25361 | null |
| 2026-04-27 | OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer | Boyang Wang et.al. | 2604.24762 | null |
| 2026-04-26 | Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling | Zhen Ye et.al. | 2604.23586 | link |
| 2026-04-24 | Video Analysis and Generation via a Semantic Progress Function | Gal Metzer et.al. | 2604.22554 | link |
| 2026-04-23 | VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis | Songen Gu et.al. | 2604.21914 | null |
| 2026-04-23 | Grounding Video Reasoning in Physical Signals | Alibay Osmanli et.al. | 2604.21873 | null |
| 2026-04-26 | Building a Precise Video Language with Human-AI Oversight | Zhiqiu Lin et.al. | 2604.21718 | link |
| 2026-04-23 | WorldMark: A Unified Benchmark Suite for Interactive Video World Models | Xiaojie Xu et.al. | 2604.21686 | link |
| 2026-04-23 | Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation | Boxun Xu et.al. | 2604.21221 | link |
| 2026-04-22 | Agentic AI for Personalized Physiotherapy: A Multi-Agent Framework for Generative Video Training and Real-Time Pose Correction | Abhishek Dharmaratnakar et.al. | 2604.21154 | null |
| 2026-04-22 | DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation | Hyeonwoo Kim et.al. | 2604.20841 | link |
| 2026-04-21 | AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model | Yutian Chen et.al. | 2604.19747 | link |
| 2026-04-21 | CityRAG: Stepping Into a City via Spatially-Grounded Video Generation | Gene Chou et.al. | 2604.19741 | null |
| 2026-04-21 | ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis | Zhengwentai Sun et.al. | 2604.19720 | null |
| 2026-04-21 | MultiWorld: Scalable Multi-Agent Multi-View Video World Models | Haoyu Wu et.al. | 2604.18564 | link |
| 2026-04-20 | Training and Agentic Inference Strategies for LLM-based Manim Animation Generation | Ravidu Suien Rammuni Silva et.al. | 2604.18364 | null |
| 2026-04-22 | ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes | Honglin Chen et.al. | 2604.17623 | null |
| 2026-04-19 | Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation | Zhijiang Tang et.al. | 2604.17428 | link |
| 2026-04-19 | DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior | Junjia Huang et.al. | 2604.17195 | link |
| 2026-04-18 | LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing | Weicheng Wang et.al. | 2604.17021 | null |
| 2026-04-14 | Motif-Video 2B: Technical Report | Junghwan Lim et.al. | 2604.16503 | null |
| 2026-04-12 | Latent-Compressed Variational Autoencoder for Video Diffusion Models | Jiarui Guan et.al. | 2604.16479 | null |
| 2026-04-17 | Efficient Video Diffusion Models: Advancements and Challenges | Shitong Shao et.al. | 2604.15911 | null |
| 2026-04-16 | TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation | Xiangyu Liu et.al. | 2604.14580 | null |
| 2026-04-15 | Geometrically Consistent Multi-View Scene Generation from Freehand Sketches | Ahmed Bourouis et.al. | 2604.14302 | null |
| 2026-04-15 | Seedance 2.0: Advancing Video Generation for World Complexity | Team Seedance et.al. | 2604.14148 | null |
| 2026-04-15 | DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer | Hengye Lyu et.al. | 2604.13509 | null |
| 2026-04-15 | VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning | Yifan Li et.al. | 2604.13425 | null |
| 2026-04-14 | ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models | Xinliang Wang et.al. | 2604.12251 | link |
| 2026-04-14 | Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation | Wentai Zhang et.al. | 2604.12219 | null |
| 2026-04-13 | Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale | Dongxu Wei et.al. | 2604.11331 | null |
| 2026-04-13 | AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps | Liaoyuan Fan et.al. | 2604.11135 | null |
| 2026-04-12 | TAPNext++: What's Next for Tracking Any Point (TAP)? | Sebastian Jung et.al. | 2604.10582 | null |
| 2026-04-14 | Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models | Dehui Wang et.al. | 2604.10578 | null |
| 2026-04-11 | VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation | Longteng Jiang et.al. | 2604.10127 | null |
| 2026-04-11 | Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation | Ruibin Li et.al. | 2604.10103 | null |
| 2026-04-11 | Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation | Gordon Chen et.al. | 2604.10030 | null |
| 2026-04-10 | Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories | Wonbong Jang et.al. | 2604.09429 | link |
| 2026-04-10 | CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation | Haoyu Zhao et.al. | 2604.09201 | null |
| 2026-04-09 | InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation | Zhefan Rao et.al. | 2604.08646 | null |
| 2026-04-09 | When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models | Zhengyang Sun et.al. | 2604.08546 | link |
| 2026-04-09 | Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics | Ying Shen et.al. | 2604.08503 | null |
| 2026-04-09 | Novel View Synthesis as Video Completion | Qi Wu et.al. | 2604.08500 | link |
| 2026-04-09 | DiV-INR: Extreme Low-Bitrate Diffusion Video Compression with INR Conditioning | Eren Çetin et.al. | 2604.08329 | null |
| 2026-04-09 | ViVa: A Video-Generative Value Model for Robot Reinforcement Learning | Jindi Lv et.al. | 2604.08168 | link |
| 2026-04-09 | Lighting-grounded Video Generation with Renderer-based Agent Reasoning | Ziqi Cai et.al. | 2604.07966 | null |
| 2026-04-08 | Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations | Chao Tang et.al. | 2604.07517 | null |
| 2026-04-08 | Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity | Yucheng Zhou et.al. | 2604.07402 | null |
| 2026-04-08 | Controllable Generative Video Compression | Ding Ding et.al. | 2604.06655 | null |
| 2026-04-10 | DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models | Zhengming Yu et.al. | 2604.06161 | link |
| 2026-04-07 | SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation | Hiba Dahmani et.al. | 2604.06113 | null |
| 2026-04-07 | HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation | Tao Hu et.al. | 2604.05961 | null |
| 2026-04-06 | Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale | Zhengcen Li et.al. | 2604.04634 | null |
| 2026-04-06 | Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? | Zhongru Zhang et.al. | 2604.04502 | null |
| 2026-04-06 | Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse | Hao Liu et.al. | 2604.04451 | null |
| 2026-04-06 | UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining | Pei Yang et.al. | 2604.04402 | null |
| 2026-04-05 | DriveVA: Video Action Models are Zero-Shot Drivers | Mengmeng Liu et.al. | 2604.04198 | null |
| 2026-04-05 | ATSS: Detecting AI-Generated Videos via Anomalous Temporal Self-Similarity | Hang Wang et.al. | 2604.04029 | link |
| 2026-04-04 | Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation | Binyuan Huang et.al. | 2604.03738 | link |
| 2026-04-04 | CRAFT: Video Diffusion for Bimanual Robot Data Generation | Jason Chen et.al. | 2604.03552 | null |
| 2026-04-03 | Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation | Xingtong Ge et.al. | 2604.03118 | link |
| 2026-04-03 | Not All Frames Deserve Full Computation: Accelerating Autoregressive Video Generation via Selective Computation and Predictive Extrapolation | Hanshuai Cui et.al. | 2604.02979 | null |
| 2026-04-03 | HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits | Leyang Jin et.al. | 2604.02867 | null |
| 2026-04-03 | NavCrafter: Exploring 3D Scenes from a Single Image | Hongbo Duan et.al. | 2604.02828 | null |
| 2026-04-03 | MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling | Shubo Lin et.al. | 2604.02817 | null |
| 2026-04-02 | ActionParty: Multi-Subject Action Binding in Generative Video Games | Alexander Pondaven et.al. | 2604.02330 | link |
| 2026-04-02 | VOID: Video Object and Interaction Deletion | Saman Motamed et.al. | 2604.02296 | null |
| 2026-04-02 | Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion | Edoardo A. Dominici et.al. | 2604.01761 | link |
| 2026-04-02 | Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation | Lingyu Liu et.al. | 2604.01700 | link |
| 2026-04-02 | From Understanding to Erasing: Towards Complete and Stable Video Object Removal | Dingming Liu et.al. | 2604.01693 | link |
| 2026-04-02 | DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data | Wonjoon Jin et.al. | 2604.01666 | null |
| 2026-04-01 | ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation | Hao Zhang et.al. | 2604.01129 | null |
| 2026-04-01 | PHASOR: Anatomy- and Phase-Consistent Volumetric Diffusion for CT Virtual Contrast Enhancement | Zilong Li et.al. | 2604.01053 | null |
| 2026-04-01 | HICT: High-precision 3D CBCT reconstruction from a single X-ray | Wen Ma et.al. | 2604.00792 | null |
| 2026-03-31 | OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation | Yuheng Liu et.al. | 2603.30045 | link |
| 2026-03-31 | Video Models Reason Early: Exploiting Plan Commitment for Maze Solving | Kaleb Newman et.al. | 2603.30043 | null |
| 2026-03-30 | Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos | Yujin Ham et.al. | 2603.29036 | null |
| 2026-03-30 | Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas | Felix Wimbauer et.al. | 2603.28980 | null |
| 2026-03-30 | Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms | Muyang He et.al. | 2603.28489 | link |
| 2026-03-30 | FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation | Liuzhou Zhang et.al. | 2603.27915 | link |
| 2026-03-29 | Wan-R1: Verifiable-Reinforcement Learning for Video Reasoning | Ming Liu et.al. | 2603.27866 | null |
| 2026-03-29 | TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets | Zhixuan Liu et.al. | 2603.27520 | link |
| 2026-03-28 | LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model | Quankai Gao et.al. | 2603.27449 | link |
| 2026-03-27 | Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling | Ruixing Zhang et.al. | 2603.26610 | null |
| 2026-03-27 | VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward | Zhaochong An et.al. | 2603.26599 | null |
| 2026-04-02 | Generation Is Compression: Zero-Shot Video Coding via Stochastic Rectified Flow | Ziyue Zeng et.al. | 2603.26571 | null |
| 2026-03-26 | THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond | Letian Wang et.al. | 2603.25892 | null |
| 2026-03-26 | PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference | Xiaofeng Mao et.al. | 2603.25730 | null |
| 2026-04-01 | Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training | Xiangyang Luo et.al. | 2603.25527 | null |
| 2026-03-26 | Free-Lunch Long Video Generation via Layer-Adaptive O.O.D Correction | Jiahao Tian et.al. | 2603.25209 | link |
| 2026-03-25 | DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation | Junyi Ouyang et.al. | 2603.24835 | link |
| 2026-04-01 | DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving | Pengxuan Yang et.al. | 2603.24587 | null |
| 2026-03-25 | Anti-I2V: Safeguarding your photos from malicious image-to-video generation | Duc Vu et.al. | 2603.24570 | null |
| 2026-03-25 | Toward Physically Consistent Driving Video World Models under Challenging Trajectories | Jiawei Zhou et.al. | 2603.24506 | null |
| 2026-03-25 | OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning | Kaihang Pan et.al. | 2603.24458 | link |
| 2026-03-24 | VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs | Haoran Yuan et.al. | 2603.23481 | link |
| 2026-03-24 | RealMaster: Lifting Rendered Scenes into Photorealistic Video | Dana Cohen-Bar et.al. | 2603.23462 | link |
| 2026-03-24 | ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images | Yunfeng Wu et.al. | 2603.23326 | link |
| 2026-03-24 | GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models | Zekai Gu et.al. | 2603.23246 | link |
| 2026-03-23 | P-Flow: Prompting Visual Effects Generation | Rui Zhao et.al. | 2603.22091 | link |
| 2026-03-23 | Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation | Yuyang You et.al. | 2603.21864 | link |
| 2026-03-23 | Climate Prompting: Generating the Madden-Julian Oscillation using Video Diffusion and Low-Dimensional Conditioning | Sulian Thual et.al. | 2603.21856 | null |
| 2026-03-23 | PROBE: Diagnosing Residual Concept Capacity in Erased Text-to-Video Diffusion Models | Yiwei Xie et.al. | 2603.21547 | null |
| 2026-03-22 | Respiratory Status Detection with Video Transformers | Thomas Savage et.al. | 2603.21349 | null |
| 2026-03-22 | Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows | Janne Perini et.al. | 2603.21210 | link |
| 2026-03-20 | MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints | Yu Qi et.al. | 2603.20194 | null |
| 2026-03-24 | Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis | Weisheng Xu et.al. | 2603.19709 | null |
| 2026-03-20 | Making Video Models Adhere to User Intent with Minor Adjustments | Daniel Ajisafe et.al. | 2603.19672 | link |
| 2026-03-20 | OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis | Jinglin Liang et.al. | 2603.19613 | null |
| 2026-03-20 | Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning | Qin Zhang et.al. | 2603.19607 | null |
| 2026-03-19 | Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding | Xianjin Wu et.al. | 2603.19235 | link |
| 2026-03-19 | V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors | Songjia He et.al. | 2603.18811 | null |
| 2026-03-19 | 6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models | Rundong Su et.al. | 2603.18742 | null |
| 2026-03-19 | PhysVideo: Physically Plausible Video Generation with Cross-View Geometry Guidance | Cong Wang et.al. | 2603.18639 | null |
| 2026-03-19 | Training-Free Sparse Attention for Fast Video Generation via Offline Layer-Wise Sparsity Profiling and Online Bidirectional Co-Clustering | Jiayi Luo et.al. | 2603.18636 | null |
| 2026-03-19 | Improving Joint Audio-Video Generation with Cross-Modal Context Learning | Bingqi Ma et.al. | 2603.18600 | null |
| 2026-03-19 | 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model | Hyun-kyu Ko et.al. | 2603.18524 | link |
| 2026-03-18 | ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation | Dmitriy Rivkin et.al. | 2603.17812 | null |
| 2026-03-18 | EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards | Ruixiang Wang et.al. | 2603.17808 | link |
| 2026-03-18 | TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos | Yan Zeng et.al. | 2603.17735 | null |
| 2026-03-18 | SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning | Xi Ye et.al. | 2603.17426 | null |
| 2026-03-21 | GigaWorld-Policy: An Efficient Action-Centered World--Action Model | Angen Ye et.al. | 2603.17240 | link |
| 2026-03-17 | MosaicMem: Hybrid Spatial Memory for Controllable Video World Models | Wei Yu et.al. | 2603.17117 | null |
| 2026-03-17 | Demystifing Video Reasoning | Ruisi Wang et.al. | 2603.16870 | link |
| 2026-03-17 | DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models | Emily Yue-Ting Jia et.al. | 2603.16860 | null |
| 2026-03-18 | World Reconstruction From Inconsistent Views | Lukas Höllein et.al. | 2603.16736 | null |
| 2026-03-17 | VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment | Tengjiao Yin et.al. | 2603.16271 | null |
| 2026-03-16 | Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion | Zhenghong Zhou et.al. | 2603.15614 | null |
| 2026-03-16 | ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer | Ruonan Yu et.al. | 2603.15478 | link |
| 2026-03-18 | Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors | Yunuo Chen et.al. | 2603.15129 | link |
| 2026-03-16 | GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis | Minjun Kang et.al. | 2603.14965 | link |
| 2026-03-16 | MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model | Jinguang Tong et.al. | 2603.14686 | null |
| 2026-03-15 | Early Failure Detection and Intervention in Video Diffusion Models | Kwon Byung-Ki et.al. | 2603.14320 | link |
| 2026-03-15 | Seeking Physics in Diffusion Noise | Chujun Tang et.al. | 2603.14294 | link |
| 2026-03-15 | CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control | Zhiyi Kuang et.al. | 2603.14241 | null |
| 2026-03-14 | Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation | Stefan Ainetter et.al. | 2603.13910 | null |
| 2026-03-14 | PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment | Zhexiao Xiong et.al. | 2603.13770 | null |
| 2026-03-14 | UniVid: Pyramid Diffusion Model for High Quality Video Generation | Xinyu Xiao et.al. | 2603.13739 | null |
| 2026-03-13 | Draft-and-Target Sampling for Video Generation Policy | Qikang Zhang et.al. | 2603.13438 | null |
| 2026-03-12 | Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion | Yang Yang et.al. | 2603.13405 | null |
| 2026-03-13 | V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration | Shenghe Zheng et.al. | 2603.13089 | link |
| 2026-03-12 | VQQA: An Agentic Approach for Video Evaluation and Quality Improvement | Yiwen Song et.al. | 2603.12310 | null |
| 2026-03-12 | EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation | Tianwei Xiong et.al. | 2603.12267 | link |
| 2026-03-12 | DVD: Deterministic Video Depth Estimation with Generative Priors | Hongfei Zhang et.al. | 2603.12250 | link |
| 2026-03-12 | OSCBench: Benchmarking Object State Change in Text-to-Video Generation | Xianjing Han et.al. | 2603.11698 | link |
| 2026-03-10 | When to Lock Attention: Training-Free KV Control in Video Diffusion | Tianyi Zeng et.al. | 2603.09657 | null |
| 2026-03-10 | Chain of Event-Centric Causal Thought for Physically Plausible Video Generation | Zixuan Wang et.al. | 2603.09094 | link |
| 2026-03-09 | HECTOR: Hybrid Editable Compositional Object References for Video Generation | Guofeng Zhang et.al. | 2603.08850 | link |
| 2026-03-09 | SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution | Chao Wang et.al. | 2603.08536 | null |
| 2026-03-11 | SPIRAL: A Closed-Loop Framework for Self-Improving Action World Models via Reflective Planning Agents | Yu Yang et.al. | 2603.08403 | null |
| 2026-03-09 | Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades | Ashkan Taghipour et.al. | 2603.08028 | null |
| 2026-03-02 | Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index | Chao Yuan et.al. | 2603.06664 | null |
| 2026-03-06 | DreamToNav: Generalizable Navigation for Robots via Generative Video Planning | Valerii Serpiva et.al. | 2603.06190 | null |
| 2026-03-06 | GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection | Xuan Huang et.al. | 2603.06048 | link |
| 2026-03-06 | VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction | Muhua Zhu et.al. | 2603.05851 | link |
| 2026-03-06 | Training-free Latent Inter-Frame Pruning with Attention Recovery | Dennis Menn et.al. | 2603.05811 | null |
| 2026-03-05 | EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation | Gehao Zhang et.al. | 2603.05757 | null |
| 2026-03-05 | FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning | Weijie Lyu et.al. | 2603.05506 | link |
| 2026-03-05 | RealWonder: Real-Time Physical Action-Conditioned Video Generation | Wei Liu et.al. | 2603.05449 | null |
| 2026-03-05 | Orthogonal Spatial-temporal Distributional Transfer for 4D Generation | Wei Liu et.al. | 2603.05081 | null |
| 2026-03-05 | FC-VFI: Faithful and Consistent Video Frame Interpolation for High-FPS Slow Motion Video Generation | Ganggui Ding et.al. | 2603.04899 | null |
| 2026-03-04 | Helios: Real Real-Time Long Video Generation Model | Shenghai Yuan et.al. | 2603.04379 | link |
| 2026-03-04 | ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors | Zihao Huang et.al. | 2603.04338 | link |
| 2026-03-03 | PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation | Shang Wu et.al. | 2603.03505 | null |
| 2026-03-06 | Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion | Haoran Lu et.al. | 2603.03485 | link |
| 2026-03-03 | SIGMark: Scalable In-Generation Watermark with Blind Extraction for Video Diffusion | Xinjie Zhu et.al. | 2603.02882 | link |
| 2026-03-03 | Compositional Visual Planning via Inference-Time Diffusion Scaling | Yixin Zhang et.al. | 2603.02646 | null |
| 2026-03-03 | Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild | Seunguk Do et.al. | 2603.02619 | null |
| 2026-02-27 | CamDirector: Towards Long-Term Coherent Video Trajectory Editing | Zhihao Shi et.al. | 2603.02256 | link |
| 2026-03-02 | LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation | Hualiang Wei et.al. | 2603.02129 | null |
| 2026-03-02 | WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories | Yisu Zhang et.al. | 2603.02049 | link |
| 2026-03-06 | FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters | Shitong Shao et.al. | 2603.01685 | null |
| 2026-03-02 | Adaptive Spectral Feature Forecasting for Diffusion Sampling Acceleration | Jiaqi Han et.al. | 2603.01623 | link |
| 2026-03-02 | UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation | Hebeizi Li et.al. | 2603.01418 | null |
| 2026-03-01 | EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization | Zhaoxin Fan et.al. | 2603.00978 | null |
| 2026-03-03 | PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation | Jiangshan Wang et.al. | 2603.00976 | null |
| 2026-02-28 | MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation | Rongsheng Wang et.al. | 2603.00585 | link |
| 2026-02-27 | SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models | Yang Yang et.al. | 2603.00194 | null |
| 2026-02-18 | Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation | Zijian Song et.al. | 2603.00110 | null |
| 2026-02-27 | HumanOrbit: 3D Human Reconstruction as 360° Orbit Generation | Keito Suzuki et.al. | 2602.24148 | link |
| 2026-02-27 | SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls | Qianxun Xu et.al. | 2602.23956 | link |
| 2026-02-26 | The Trinity of Consistency as a Defining Principle for General World Models | Jingxuan Wei et.al. | 2602.23152 | link |
| 2026-02-26 | Solaris: Building a Multiplayer Video World Model in Minecraft | Georgy Savva et.al. | 2602.22208 | null |
| 2026-02-25 | Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context | JiaKui Hu et.al. | 2602.21929 | null |
| 2026-02-24 | Human Video Generation from a Single Image with 3D Pose and View Control | Tiantian Wang et.al. | 2602.21188 | null |
| 2026-03-01 | VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models | Bowen Zheng et.al. | 2602.20999 | null |
| 2026-02-24 | GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generatio | Hao Zhang et.al. | 2602.20673 | null |
| 2026-02-24 | PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models | Wonyong Seo et.al. | 2602.20583 | link |
| 2026-02-23 | NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning | Jiahui Fu et.al. | 2602.20119 | null |
| 2026-02-23 | Pixel2Phys: Distilling Governing Laws from Visual Dynamics | Ruikun Li et.al. | 2602.19516 | link |
| 2026-02-22 | UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models | Gang Xu et.al. | 2602.19202 | null |
| 2026-02-22 | Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling | Qi Sun et.al. | 2602.19089 | link |
| 2026-02-21 | RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning | Seungku Kim et.al. | 2602.18742 | null |
| 2026-02-20 | Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control | Linxi Xie et.al. | 2602.18422 | link |
| 2026-02-20 | Predict to Skip: Linear Multistep Feature Forecasting for Efficient Diffusion Transformers | Hanshuai Cui et.al. | 2602.18093 | null |
| 2026-02-18 | CHAI: CacHe Attention Inference for text2video | Joel Mathew Cherian et.al. | 2602.16132 | null |
| 2026-02-17 | World Action Models are Zero-shot Policies | Seonghyeon Ye et.al. | 2602.15922 | null |
| 2026-02-17 | VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation | Hui Ren et.al. | 2602.15819 | link |
| 2026-02-17 | Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation | Jia Li et.al. | 2602.14027 | null |
| 2026-02-14 | High-Fidelity Causal Video Diffusion Models for Real-Time Ultra-Low-Bitrate Semantic Communication | Cem Eteke et.al. | 2602.13837 | null |
| 2026-02-14 | EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation | Rang Meng et.al. | 2602.13669 | null |
| 2026-02-14 | DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation | Haoyu Zhao et.al. | 2602.13637 | null |
| 2026-02-13 | SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning | Jintao Zhang et.al. | 2602.13515 | null |
| 2026-02-13 | SLA2: Sparse-Linear Attention with Learnable Routing and QAT | Jintao Zhang et.al. | 2602.12675 | null |
| 2026-02-13 | Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL | Xin Liu et.al. | 2602.12636 | null |
| 2026-02-12 | OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model | Maomao Li et.al. | 2602.12304 | null |
| 2026-02-12 | MonarchRT: Efficient Attention for Real-Time Video Generation | Krish Agarwal et.al. | 2602.12271 | null |
| 2026-02-15 | VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model | Yanjiang Guo et.al. | 2602.12063 | null |
| 2026-02-12 | Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation | Lingyong Yan et.al. | 2602.11790 | null |
| 2026-02-12 | LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts | Chen Zhao et.al. | 2602.11564 | null |
| 2026-02-11 | SurfPhase: 3D Interfacial Dynamics in Two-Phase Flows from Sparse Videos | Yue Gao et.al. | 2602.11154 | null |
| 2026-02-11 | Flow caching for autoregressive video generation | Yuexiao Ma et.al. | 2602.10825 | null |
| 2026-02-11 | Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation | Songen Gu et.al. | 2602.10717 | null |
| 2026-02-10 | ArtisanGS: Interactive Tools for Gaussian Splat Selection with AI and Human in the Loop | Clement Fuji Tsang et.al. | 2602.10173 | null |
| 2026-02-10 | ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation | Mingyang Wu et.al. | 2602.10113 | null |
| 2026-02-10 | DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos | Juncheng Mu et.al. | 2602.10105 | link |
| 2026-02-10 | VideoWorld 2: Learning Transferable Knowledge from Real-world Videos | Zhongwei Ren et.al. | 2602.10102 | null |
| 2026-02-11 | Monocular Normal Estimation via Shading Sequence Estimation | Zongrui Li et.al. | 2602.09929 | link |
| 2026-02-09 | Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing | Hao Yang et.al. | 2602.08820 | null |
| 2026-02-10 | ALIVE: Animate Your World with Lifelike Audio-Video Generation | Ying Guo et.al. | 2602.08682 | null |
| 2026-02-09 | PISCO: Precise Video Instance Insertion with Sparse Control | Xiangbo Gao et.al. | 2602.08277 | link |
| 2026-02-04 | Reliable and Responsible Foundation Models: A Comprehensive Survey | Xinyu Yang et.al. | 2602.08145 | null |
| 2026-02-08 | ReRoPE: Repurposing RoPE for Relative Camera Control | Chunyang Li et.al. | 2602.08068 | null |
| 2026-02-08 | Geometry-Aware Rotary Position Embedding for Consistent Video World Model | Chendong Xiang et.al. | 2602.07854 | link |
| 2026-02-08 | Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion | Haodong Li et.al. | 2602.07775 | link |
| 2026-02-07 | IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation | Zhufeng Xu et.al. | 2602.07498 | null |
| 2026-02-06 | VideoNeuMat: Neural Material Extraction from Generative Video Models | Bowen Xue et.al. | 2602.07272 | null |
| 2026-02-04 | Interpreting Physics in Video World Models | Sonia Joseph et.al. | 2602.07050 | link |
| 2026-02-06 | CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation | Kaiyi Huang et.al. | 2602.06959 | null |
| 2026-02-05 | LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation | Mirlan Karimov et.al. | 2602.05966 | link |
| 2026-02-05 | Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation | Hai Zhang et.al. | 2602.05827 | null |
| 2026-02-05 | GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling | Shivanshu Shekhar et.al. | 2602.05202 | null |
| 2026-02-04 | Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention | Chengtao Lv et.al. | 2602.04789 | link |
| 2026-02-04 | Adaptive 1D Video Diffusion Autoencoder | Yao Teng et.al. | 2602.04220 | null |
| 2026-02-03 | WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling | Michael Aich et.al. | 2602.03924 | null |
| 2026-02-03 | BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks | Yixiang Chen et.al. | 2602.03793 | null |
| 2026-02-03 | Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization | Haocheng Xi et.al. | 2602.02958 | null |
| 2026-02-06 | Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation | Hongzhou Zhu et.al. | 2602.02214 | link |
| 2026-02-02 | FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space | FSVideo Team et.al. | 2602.02092 | null |
| 2026-02-02 | Grounding Generated Videos in Feasible Plans via World Models | Christos Ziakas et.al. | 2602.01960 | null |
| 2026-02-02 | Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention | Dvir Samuel et.al. | 2602.01801 | null |
| 2026-02-02 | FastPhysGS: Accelerating Physics-based Dynamic 3DGS Simulation via Interior Completion and Adaptive Optimization | Yikun Ma et.al. | 2602.01723 | null |
| 2026-02-02 | Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation? | Susan Liang et.al. | 2602.01623 | null |
| 2026-02-01 | MTC-VAE: Multi-Level Temporal Compression with Content Awareness | Yubo Dong et.al. | 2602.01340 | null |
| 2026-01-29 | Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields | Shiqian Li et.al. | 2602.00148 | link |
| 2026-01-30 | VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation | Hongyang Du et.al. | 2601.23286 | null |
| 2026-01-29 | JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion | Anthony Chen et.al. | 2601.22143 | link |
| 2026-01-29 | EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers | John Flynn et.al. | 2601.22127 | null |
| 2026-01-29 | Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models | Cong Cao et.al. | 2601.21922 | null |
| 2026-02-02 | MPF-Net: Exposing High-Fidelity AI-Generated Video Forgeries via Hierarchical Manifold Deviation and Micro-Temporal Fluctuations | Xinan He et.al. | 2601.21408 | null |
| 2026-01-28 | Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning | Chengzu Li et.al. | 2601.21037 | null |
| 2026-01-28 | FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion Models | Hongyu Zhou et.al. | 2601.20857 | link |
| 2026-01-28 | FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models | Haonan Zhong et.al. | 2601.20791 | null |
| 2026-01-28 | Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V | Meiqi Wu et.al. | 2601.20504 | null |
| 2026-01-28 | Efficient Autoregressive Video Diffusion with Dummy Head | Hang Guo et.al. | 2601.20499 | null |
| 2026-01-28 | Artifact-Aware Evaluation for High-Quality Video Generation | Chen Zhu et.al. | 2601.20297 | null |
| 2026-01-27 | VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics | Zhiyu Yin et.al. | 2601.19236 | null |
| 2026-01-26 | FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Geometry-Complete 4D Reconstruction | Wei Cao et.al. | 2601.18993 | null |
| 2026-01-26 | Are Video Generation Models Geographically Fair? An Attraction-Centric Evaluation of Global Visual Knowledge | Xiao Liu et.al. | 2601.18698 | null |
| 2026-01-29 | SkyReels-V3 Technique Report | Debang Li et.al. | 2601.17323 | null |
| 2026-01-22 | A Mechanistic View on Video Generation as World Models: State and Dynamics | Luozhou Wang et.al. | 2601.17067 | link |
| 2026-01-22 | Memory-V2V: Augmenting Video-to-Video Diffusion Models with Memory | Dohun Lee et.al. | 2601.16296 | null |
| 2026-01-29 | GR3EN: Generative Relighting for 3D Environments | Xiaoyan Xing et.al. | 2601.16272 | null |
| 2026-01-22 | CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback | Wenhang Ge et.al. | 2601.16214 | null |
| 2026-01-22 | Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning | Moo Jin Kim et.al. | 2601.16163 | null |
| 2026-01-22 | PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models | Chak-Wing Mak et.al. | 2601.16007 | null |
| 2026-01-21 | Walk through Paintings: Egocentric World Models from Internet Priors | Anurag Bagchi et.al. | 2601.15284 | null |
| 2026-01-21 | Rethinking Video Generation Model for the Embodied World | Yufan Deng et.al. | 2601.15282 | null |
| 2026-01-21 | ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generation | Hanlei Guo et.al. | 2601.15221 | null |
| 2026-01-20 | VideoMaMa: Mask-Guided Video Matting via Generative Prior | Sangbeom Lim et.al. | 2601.14255 | null |
| 2026-01-21 | Human detectors are surprisingly powerful reward models | Kumar Ashutosh et.al. | 2601.14037 | null |
| 2026-01-19 | Moaw: Unleashing Motion Awareness for Video Diffusion Models | Tianqi Zhang et.al. | 2601.12761 | null |
| 2026-01-16 | PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models | Qiyuan Zhang et.al. | 2601.11087 | null |
| 2026-01-15 | CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos | Chengfeng Zhao et.al. | 2601.10632 | link |
| 2026-01-15 | Inference-time Physics Alignment of Video Generative Models with Latent World Models | Jianhao Yuan et.al. | 2601.10553 | null |
| 2026-01-15 | Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation | Dong-Yu Chen et.al. | 2601.10214 | null |
| 2026-01-15 | CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation | Chengzhuo Tong et.al. | 2601.10061 | null |
| 2026-01-14 | Transition Matching Distillation for Fast Video Generation | Weili Nie et.al. | 2601.09881 | null |
| 2026-01-14 | Efficient Camera-Controlled Video Generation of Static Scenes via Sparse Diffusion and 3D Rendering | Jieying Chen et.al. | 2601.09697 | null |
| 2026-01-14 | MAD: Motion Appearance Decoupling for efficient Driving World Models | Ahmad Rahimi et.al. | 2601.09452 | link |
| 2026-01-14 | PhyRPR: Training-Free Physics-Constrained Video Generation | Yibo Zhao et.al. | 2601.09255 | null |
| 2026-01-14 | Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers | Jonas Römer et.al. | 2601.09040 | link |
| 2026-01-13 | Motion Attribution for Video Generation | Xindi Wu et.al. | 2601.08828 | null |
| 2026-01-12 | Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions | Zhiting Mei et.al. | 2601.07823 | null |
| 2026-01-12 | Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models | Yuanyang Yin et.al. | 2601.07287 | null |
| 2026-01-09 | Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals | Nate Gillman et.al. | 2601.05848 | link |
| 2026-01-09 | Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation | Jin Wang et.al. | 2601.05722 | null |
| 2026-01-08 | VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control | Sixiao Zheng et.al. | 2601.05138 | link |
| 2026-01-07 | ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers | Mohsen Ghafoorian et.al. | 2601.04342 | null |
| 2026-01-07 | Choreographing a World of Dynamic Objects | Yanzhe Lyu et.al. | 2601.04194 | link |
| 2026-01-07 | Diffusion-DRF: Differentiable Reward Flow for Video Diffusion Fine-Tuning | Yifan Wang et.al. | 2601.04153 | null |
| 2026-01-07 | Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction | Jiaxin Huang et.al. | 2601.04090 | link |
| 2026-01-08 | Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models | Zitong Huang et.al. | 2601.04068 | null |
| 2026-01-07 | PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance | Siddarth Nilol Kundur Satish et.al. | 2601.03665 | null |
| 2026-01-06 | LTX-2: Efficient Joint Audio-Visual Foundation Model | Yoav HaCohen et.al. | 2601.03233 | null |
| 2026-01-06 | DreamStyle: A Unified Framework for Video Stylization | Mengtian Li et.al. | 2601.02785 | link |
| 2026-01-06 | DreamLoop: Controllable Cinemagraph Generation from a Single Photograph | Aniruddha Mahapatra et.al. | 2601.02646 | null |
| 2026-01-05 | SingingBot: An Avatar-Driven System for Robotic Face Singing Performance | Zhuoxiong Xu et.al. | 2601.02125 | null |
| 2026-01-04 | DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving | Yang Zhou et.al. | 2601.01528 | link |
| 2026-01-02 | Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians | Melonie de Almeida et.al. | 2601.00678 | null |
| 2026-01-01 | MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation | Miaowei Wang et.al. | 2601.00504 | null |
| 2025-12-31 | TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model | Yabo Chen et.al. | 2601.00051 | null |
| 2025-12-31 | SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time | Zhening Huang et.al. | 2512.25075 | link |
| 2025-12-31 | Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow | Karthik Dharmarajan et.al. | 2512.24766 | link |
| 2025-12-30 | AI-Driven Evaluation of Surgical Skill via Action Recognition | Yan Meng et.al. | 2512.24411 | null |
| 2025-12-30 | Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes | Shuyun Wang et.al. | 2512.24227 | link |
| 2025-12-30 | DriveExplorer: Images-Only Decoupled 4D Reconstruction with Progressive Restoration for Driving View Extrapolation | Yuang Jia et.al. | 2512.23983 | null |
| 2025-12-30 | T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models | Changzhen Li et.al. | 2512.23953 | null |
| 2025-12-29 | Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation | Shaocong Xu et.al. | 2512.23705 | link |
| 2025-12-29 | Bridging Your Imagination with Audio-Video Generation via a Unified Director | Jiaxu Zhang et.al. | 2512.23222 | null |
| 2025-12-27 | Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion | Yuming Gu et.al. | 2512.22626 | null |
| 2025-12-25 | GeCo: A Differentiable Geometric Consistency Metric for Video Generation | Leslie Gu et.al. | 2512.22274 | link |
| 2025-12-26 | StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars | Zhiyao Sun et.al. | 2512.22065 | link |
| 2025-12-31 | Inference-based GAN Video Generation | Jingbo Yang et.al. | 2512.21776 | null |
| 2025-12-30 | SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wild | Xindi Zhang et.al. | 2512.21736 | null |
| 2025-12-29 | Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation | Steven Xiao et.al. | 2512.21734 | link |
| 2025-12-25 | SVBench: Evaluation of Video Generation Models on Social Reasoning | Wenshuo Peng et.al. | 2512.21507 | link |
| 2025-12-24 | Surgical Scene Segmentation using a Spike-Driven Video Transformer with Real-Time Potential | Shihao Zou et.al. | 2512.21284 | null |
| 2025-12-24 | ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision | Weiqi Li et.al. | 2512.21268 | null |
| 2025-12-25 | DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation | Jiawei Liu et.al. | 2512.21252 | null |
| 2025-12-25 | SemanticGen: Video Generation in Semantic Space | Jianhong Bai et.al. | 2512.20619 | null |
| 2025-12-23 | Learning Skills from Action-Free Videos | Hung-Chieh Fang et.al. | 2512.20052 | null |
| 2025-12-23 | How Much 3D Do Video Foundation Models Encode? | Zixuan Huang et.al. | 2512.19949 | link |
| 2025-12-29 | Learning to Refocus with Video Diffusion Models | SaiKiran Tedla et.al. | 2512.19823 | link |
| 2025-12-22 | Generating the Past, Present and Future from a Motion-Blurred Image | SaiKiran Tedla et.al. | 2512.19817 | null |
| 2025-12-22 | Over++: Generative Video Compositing for Layer Interaction Effects | Luchao Qi et.al. | 2512.19661 | null |
| 2025-12-22 | StoryMem: Multi-shot Long Video Storytelling with Memory | Kaiwen Zhang et.al. | 2512.19539 | link |
| 2025-12-22 | Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface | Yujie Zhao et.al. | 2512.19402 | link |
| 2025-12-21 | EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer | Yuxiao Yang et.al. | 2512.18814 | link |
| 2025-12-21 | A Study of Finetuning Video Transformers for Multi-view Geometry Tasks | Huimin Wu et.al. | 2512.18684 | null |
| 2025-12-21 | PTTA: A Pure Text-to-Animation Framework for High-Quality Creation | Ruiqi Chen et.al. | 2512.18614 | null |
| 2025-12-19 | Map2Video: Street View Imagery Driven AI Video Generation | Hye-Young Jo et.al. | 2512.17883 | link |
| 2025-12-19 | Vidarc: Embodied Video Diffusion Model for Closed-loop Control | Yao Feng et.al. | 2512.17661 | null |
| 2025-12-19 | InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion | Hoiyeong Jin et.al. | 2512.17504 | link |
| 2025-12-19 | Mitty: Diffusion-based Human-to-Robot Video Generation | Yiren Song et.al. | 2512.17253 | link |
| 2025-12-18 | Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation | Min-Jung Kim et.al. | 2512.17040 | null |
| 2025-12-17 | Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories | Chayan Jain et.al. | 2512.16954 | null |
| 2025-12-18 | Instant Expressive Gaussian Head Avatar via 3D-Aware Expression Distillation | Kaiwen Jiang et.al. | 2512.16893 | null |
| 2025-12-18 | Factorized Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models | Mariam Hassan et.al. | 2512.16371 | link |
| 2025-12-18 | TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times | Jintao Zhang et.al. | 2512.16093 | null |
| 2025-12-17 | CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion | Liudi Yang et.al. | 2512.16023 | null |
| 2025-12-17 | Spatia: Video Generation with Updatable Spatial Memory | Jinjing Zhao et.al. | 2512.15716 | null |
| 2025-12-17 | End-to-End Training for Autoregressive Video Diffusion via Self-Resampling | Yuwei Guo et.al. | 2512.15702 | null |
| 2025-12-17 | GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models | Bozhou Li et.al. | 2512.15560 | null |
| 2025-12-17 | DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations | Yuxiang Shi et.al. | 2512.15524 | null |
| 2025-12-16 | MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives | Sihui Ji et.al. | 2512.14699 | link |
| 2025-12-16 | WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling | Wenqiang Sun et.al. | 2512.14614 | null |
| 2025-12-16 | SS4D: Native 4D Generative Model via Structured Spacetime Latents | Zhibing Li et.al. | 2512.14284 | link |
| 2025-12-16 | DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos | Yang Bai et.al. | 2512.14217 | null |
| 2025-12-16 | AnimaMimic: Imitating 3D Animation from Video Priors | Tianyi Xie et.al. | 2512.14133 | null |
| 2025-12-16 | AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation | Sisi Dai et.al. | 2512.14095 | null |
| 2025-12-15 | DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders | Susung Hong et.al. | 2512.13690 | null |
| 2025-12-15 | KlingAvatar 2.0 Technical Report | Kling Team et.al. | 2512.13313 | null |
| 2025-12-18 | Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? | Jiaqi Wang et.al. | 2512.13281 | null |
| 2025-12-15 | STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits | Foivos Paraperas Papantoniou et.al. | 2512.13247 | link |
| 2025-12-15 | Motus: A Unified Latent Action World Model | Hongzhe Bi et.al. | 2512.13030 | link |
| 2025-12-15 | SneakPeek: Future-Guided Instructional Streaming Video Generation | Cheeun Hong et.al. | 2512.13019 | null |
| 2025-12-14 | GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation | Zhenya Yang et.al. | 2512.12751 | link |
| 2025-12-14 | Supervised Contrastive Frame Aggregation for Video Representation Learning | Shaif Chowdhury et.al. | 2512.12549 | null |
| 2025-12-14 | Animus3D: Text-driven 3D Animation via Motion Score Distillation | Qi Sun et.al. | 2512.12534 | null |
| 2025-12-14 | Generative Spatiotemporal Data Augmentation | Jinfan Zhou et.al. | 2512.12508 | null |
| 2025-12-13 | V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping | Hyunkoo Lee et.al. | 2512.12375 | null |
| 2025-12-12 | SPDMark: Selective Parameter Displacement for Robust Video Watermarking | Samar Fares et.al. | 2512.12090 | null |
| 2025-12-12 | BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models | Ryan Po et.al. | 2512.12080 | null |
| 2025-12-12 | V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties | Ye Fang et.al. | 2512.11799 | link |
| 2025-12-12 | AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis | Junjie Ye et.al. | 2512.11797 | link |
| 2025-12-12 | Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation | Yang Fei et.al. | 2512.11792 | null |
| 2025-12-12 | FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion | Xiangyang Luo et.al. | 2512.11274 | null |
| 2025-12-15 | AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path | Zhengyang Yu et.al. | 2512.11203 | null |
| 2025-12-11 | AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation | Sharath Girish et.al. | 2512.10943 | null |
| 2025-12-10 | UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving | Hao Lu et.al. | 2512.09864 | null |
| 2025-12-10 | VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification | Wanyue Zhang et.al. | 2512.09646 | null |
| 2025-12-10 | DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping | Yanan Wang et.al. | 2512.09417 | null |
| 2025-12-10 | H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos | Hai Ci et.al. | 2512.09406 | null |
| 2025-12-10 | VABench: A Comprehensive Benchmark for Audio-Video Generation | Daili Hua et.al. | 2512.09299 | null |
| 2025-12-09 | Astra: General Interactive World Model with Autoregressive Denoising | Yixuan Zhu et.al. | 2512.08931 | link |
| 2025-12-09 | Self-Evolving 3D Scene Generation from a Single Image | Kaizhi Zheng et.al. | 2512.08905 | null |
| 2025-12-09 | Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance | Ruihang Chu et.al. | 2512.08765 | link |
| 2025-12-09 | EgoX: Egocentric Video Generation from a Single Exocentric Video | Taewoong Kang et.al. | 2512.08269 | link |
| 2025-12-09 | Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model | Wenjiang Xu et.al. | 2512.08188 | link |
| 2025-12-08 | UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation | Jiehui Huang et.al. | 2512.07831 | link |
| 2025-12-09 | ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation | Fan Yang et.al. | 2512.07720 | null |
| 2025-12-08 | Unified Video Editing with Temporal Reasoner | Xiangpeng Yang et.al. | 2512.07469 | link |
| 2025-12-08 | Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism | Zhiyuan Wu et.al. | 2512.07350 | null |
| 2025-12-07 | VideoVLA: Video Generators Can Be Generalizable Robot Manipulators | Yichao Shen et.al. | 2512.06963 | link |
| 2025-12-07 | Pseudo Anomalies Are All You Need: Diffusion-Based Generation for Weakly-Supervised Video Anomaly Detection | Satoshi Hashimoto et.al. | 2512.06845 | null |
| 2025-12-07 | RunawayEvil: Jailbreaking the Image-to-Video Generative Models | Songping Wang et.al. | 2512.06674 | null |
| 2025-12-07 | MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment | Ruicheng Zhang et.al. | 2512.06628 | null |
| 2025-12-05 | Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation | Su Sun et.al. | 2512.06158 | null |
| 2025-12-05 | AQUA-Net: Adaptive Frequency Fusion and Illumination Aware Network for Underwater Image Enhancement | Munsif Ali et.al. | 2512.05960 | null |
| 2025-12-05 | World Models That Know When They Don't Know: Controllable Video Generation with Calibrated Uncertainty | Zhiting Mei et.al. | 2512.05927 | link |
| 2025-12-05 | SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations | Wenhao Yan et.al. | 2512.05905 | link |
| 2025-12-05 | Bring Your Dreams to Life: Continual Text-to-Video Customization | Jiahua Dong et.al. | 2512.05802 | link |
| 2025-12-05 | USV: Unified Sparsification for Accelerating Video Diffusion Models | Xinjian Wu et.al. | 2512.05754 | null |
| 2025-12-05 | ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior | Weikai Lu et.al. | 2512.05745 | link |
| 2025-12-05 | InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem | Yeobin Hong et.al. | 2512.05672 | link |
| 2025-12-05 | ProPhy: Progressive Physical Alignment for Dynamic World Simulation | Zijun Wang et.al. | 2512.05564 | link |
| 2025-12-05 | VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation | Chinthani Sugandhika et.al. | 2512.05524 | null |
| 2025-12-05 | User Negotiations of Authenticity, Ownership, and Governance on AI-Generated Video Platforms: Evidence from Sora | Bohui Shen et.al. | 2512.05519 | null |
| 2025-12-05 | WaterWave: Bridging Underwater Image Enhancement into Video Streams via Wavelet-based Temporal Consistency Field | Qi Zhu et.al. | 2512.05492 | null |
| 2025-12-05 | Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability | Shizhan Liu et.al. | 2512.05394 | link |
| 2025-12-05 | CATNUS: Coordinate-Aware Thalamic Nuclei Segmentation Using T1-Weighted MRI | Anqi Feng et.al. | 2512.05329 | null |
| 2025-12-04 | IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction | Dmitrii Torbunov et.al. | 2512.05240 | link |
| 2025-12-04 | Invariance Co-training for Robot Visual Generalization | Jonathan Yang et.al. | 2512.05230 | null |
| 2025-12-04 | Light-X: Generative 4D Video Rendering with Camera and Illumination Control | Tianqi Liu et.al. | 2512.05115 | link |
| 2025-12-04 | NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation | Yu Zeng et.al. | 2512.05106 | null |
| 2025-12-08 | TV2TV: A Unified Framework for Interleaved Language and Video Generation | Xiaochuang Han et.al. | 2512.05103 | null |
| 2025-12-04 | From Generated Human Videos to Physically Plausible Robot Trajectories | James Ni et.al. | 2512.05094 | link |
| 2025-12-04 | Object Reconstruction under Occlusion with Generative Priors and Contact-induced Constraints | Minghan Zhu et.al. | 2512.05079 | null |
| 2025-12-04 | BulletTime: Decoupled Control of Time and Camera Pose for Video Generation | Yiming Wang et.al. | 2512.05076 | null |
| 2025-12-04 | Generative Neural Video Compression via Video Diffusion Prior | Qi Mao et.al. | 2512.05016 | null |
| 2025-12-04 | Hybrid-Diffusion Models: Combining Open-loop Routines with Visuomotor Diffusion Policies | Jonne Van Haastregt et.al. | 2512.04960 | null |
| 2025-12-04 | Multi Task Denoiser Training for Solving Linear Inverse Problems | Clément Bled et.al. | 2512.04709 | null |
| 2025-12-04 | Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation | Yunhong Lu et.al. | 2512.04678 | link |
| 2025-12-04 | Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length | Yubo Huang et.al. | 2512.04677 | link |
| 2025-12-04 | SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding | Chang-Hsun Wu et.al. | 2512.04643 | null |
| 2025-12-04 | Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence | Tianyu Yuan et.al. | 2512.04619 | null |
| 2025-12-04 | VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management | Hongbo Jin et.al. | 2512.04540 | null |
| 2025-12-04 | X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale | Pei Yang et.al. | 2512.04537 | null |
| 2025-12-04 | Refaçade: Editing Object with Given Reference Texture | Youze Huang et.al. | 2512.04534 | null |
| 2025-12-04 | PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement | Yu-Wei Zhan et.al. | 2512.04532 | null |
| 2025-12-04 | VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory | Yifei Yu et.al. | 2512.04519 | null |
| 2025-12-04 | EgoLCD: Egocentric Video Generation with Long Context Diffusion | Liuzhou Zhang et.al. | 2512.04515 | link |
| 2025-12-04 | Not All Birds Look The Same: Identity-Preserving Generation For Birds | Aaron Sun et.al. | 2512.04485 | null |
| 2025-12-03 | Stable Signer: Hierarchical Sign Language Generative Model | Sen Fang et.al. | 2512.04048 | null |
| 2025-12-03 | RELIC: Interactive Video World Model with Long-Horizon Memory | Yicong Hong et.al. | 2512.04040 | null |
| 2025-12-03 | PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation | Xiaolong Li et.al. | 2512.04025 | link |
| 2025-12-03 | TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning | Tao Wu et.al. | 2512.03963 | null |
| 2025-12-03 | Classification of User Satisfaction in HRI with Social Signals in the Wild | Michael Schiffmann et.al. | 2512.03945 | null |
| 2025-12-03 | UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework | Youxin Pang et.al. | 2512.03918 | null |
| 2025-12-03 | Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence | Shuai Yang et.al. | 2512.03905 | link |
| 2025-12-03 | A Robust Camera-based Method for Breath Rate Measurement | Alexey Protopopov et.al. | 2512.03827 | null |
| 2025-12-03 | ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos | Qi'ao Xu et.al. | 2512.03666 | link |
| 2025-12-03 | The promising potential of vision language models for the generation of textual weather forecasts | Edward C. C. Steele et.al. | 2512.03623 | null |
| 2025-12-03 | ReCamDriving: LiDAR-Free Camera-Controlled Novel Trajectory Video Generation | Yaokun Li et.al. | 2512.03621 | link |
| 2025-12-03 | LAMP: Language-Assisted Motion Planning for Controllable Video Generation | Muhammed Burak Kizil et.al. | 2512.03619 | null |
| 2025-12-03 | Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding | Haoran Zhou et.al. | 2512.03601 | null |
| 2025-12-03 | Beyond Boundary Frames: Audio-Visual Semantic Guidance for Context-Aware Video Interpolation | Yuchen Deng et.al. | 2512.03590 | null |
| 2025-12-03 | Dynamic Optical Test for Bot Identification (DOT-BI): A simple check to identify bots in surveys and online processes | Malte Bleeker et.al. | 2512.03580 | link |
| 2025-12-03 | Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching | Wei Chee Yew et.al. | 2512.03553 | null |
| 2025-12-03 | Real-Time Control and Automation Framework for Acousto-Holographic Microscopy | Hasan Berkay Abdioğlu et.al. | 2512.03539 | null |
| 2025-12-03 | FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation | Yiyi Cai et.al. | 2512.03520 | link |
| 2025-12-03 | Towards Object-centric Understanding for Instructional Videos | Wenliang Guo et.al. | 2512.03479 | null |
| 2025-12-03 | GeoVideo: Introducing Geometric Regularization into Video Generation Model | Yunpeng Bai et.al. | 2512.03453 | link |
| 2025-12-02 | Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling | Yueru Jia et.al. | 2512.03044 | link |
| 2025-12-02 | OneThinker: All-in-one Reasoning Model for Image and Video | Kaituo Feng et.al. | 2512.03043 | link |
| 2025-12-02 | MultiShotMaster: A Controllable Multi-Shot Video Generation Framework | Qinghe Wang et.al. | 2512.03041 | link |
| 2025-12-02 | Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation | Zeqi Xiao et.al. | 2512.03040 | null |
| 2025-12-02 | ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation | Mengchen Zhang et.al. | 2512.03036 | link |
| 2025-12-02 | MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation | Youxin Pang et.al. | 2512.03034 | null |
| 2025-12-02 | SMP: Reusable Score-Matching Motion Priors for Physics-Based Character Control | Yuxuan Mu et.al. | 2512.03028 | null |
| 2025-12-02 | Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks | Matthew Dutson et.al. | 2512.03014 | link |
| 2025-12-02 | In-Context Sync-LoRA for Portrait Video Editing | Sagi Polaczek et.al. | 2512.03013 | null |
| 2025-12-02 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling | Kairun Wen et.al. | 2512.03000 | link |
| 2025-12-02 | Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench | Lanxiang Hu et.al. | 2512.02942 | null |
| 2025-12-02 | Maintaining SUV Accuracy in Low-Count PET with PETfectior: A Deep Learning Denoising Solution | Yamila Rotstein Habarnau et.al. | 2512.02917 | null |
| 2025-12-02 | Taming Camera-Controlled Video Generation with Verifiable Geometry Reward | Zhaoqing Wang et.al. | 2512.02870 | null |
| 2025-12-02 | Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video? | Manuel Benavent-Lledo et.al. | 2512.02846 | link |
| 2025-12-02 | ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning | Yifan Li et.al. | 2512.02835 | link |
| 2025-12-02 | From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity | Haoming Liu et.al. | 2512.02826 | null |
| 2025-12-02 | Learning Science and the Illusion of Understanding: Exploring the Effects of Integrating Learning Tasks after Explainer Videos | Madeleine Hörnlein et.al. | 2512.02824 | null |
| 2025-12-02 | FiMMIA: scaling semantic perturbation-based membership inference across modalities | Anton Emelyanov et.al. | 2512.02786 | link |
| 2025-12-02 | Reasoning-Aware Multimodal Fusion for Hateful Video Detection | Shuonan Yang et.al. | 2512.02743 | link |
| 2025-12-02 | RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning | Yuhong Zhang et.al. | 2512.02729 | null |
| 2025-12-02 | RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence | Xuming He et.al. | 2512.02622 | null |
| 2025-12-02 | Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision | Chenshuang Zhang et.al. | 2512.02339 | null |
| 2025-12-01 | Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now | Varun Varma Thozhiyoor et.al. | 2512.02016 | null |
| 2025-12-01 | Generative Video Motion Editing with 3D Point Tracks | Yao-Chih Lee et.al. | 2512.02015 | null |
| 2025-12-01 | TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models | Zhiheng Liu et.al. | 2512.02014 | null |
| 2025-12-01 | Learning Dexterous Manipulation Skills from Imperfect Simulations | Elvis Hsieh et.al. | 2512.02011 | link |
| 2025-12-01 | Learning Visual Affordance from Audio | Lidong Lu et.al. | 2512.02005 | null |
| 2025-12-01 | PAI-Bench: A Comprehensive Benchmark For Physical AI | Fengzhe Zhou et.al. | 2512.01989 | link |
| 2025-12-01 | SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation | Zisu Li et.al. | 2512.01960 | null |
| 2025-12-01 | GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment | Haoyang He et.al. | 2512.01952 | link |
| 2025-12-01 | Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models | Zhongyu Yang et.al. | 2512.01949 | null |
| 2025-12-01 | TransientTrack: Advanced Multi-Object Tracking and Classification of Cancer Cells with Transient Fluorescent Signals | Florian Bürger et.al. | 2512.01885 | null |
| 2025-12-01 | COACH: Collaborative Agents for Contextual Highlighting - A Multi-Agent Framework for Sports Video Analysis | Tsz-To Wong et.al. | 2512.01853 | null |
| 2025-12-01 | JPEGs Just Got Snipped: Croppable Signatures Against Deepfake Images | Pericle Perazzo et.al. | 2512.01845 | null |
| 2025-12-01 | PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models | Zeqing Wang et.al. | 2512.01843 | link |
| 2025-12-01 | Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling | Meng Cao et.al. | 2512.01821 | link |
| 2025-12-02 | Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos | Xavier Thomas et.al. | 2512.01803 | null |
| 2025-12-01 | Evaluating SAM2 for Video Semantic Segmentation | Syed Hesham Syed Ariff et.al. | 2512.01774 | null |
| 2025-12-01 | VideoScoop: A Non-Traditional Domain-Independent Framework For Video Analysis | Hafsa Billah et.al. | 2512.01769 | null |
| 2025-12-01 | StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos | Daeun Lee et.al. | 2512.01707 | link |
| 2025-12-01 | DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models | Patrick Kwon et.al. | 2512.01686 | link |
| 2025-12-01 | Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation | Haodong Yan et.al. | 2512.01677 | null |
| 2025-12-01 | ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling | Qisen Wang et.al. | 2512.01481 | null |
| 2025-11-28 | Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models | Muhammad Maaz et.al. | 2511.23478 | link |
| 2025-11-28 | Video-CoM: Interactive Video Reasoning via Chain of Manipulations | Hanoona Rasheed et.al. | 2511.23477 | link |
| 2025-11-28 | AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement | Zhizhou Zhong et.al. | 2511.23475 | link |
| 2025-11-28 | Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model | Junshu Tang et.al. | 2511.23429 | null |
| 2025-11-28 | DisMo: Disentangled Motion Representations for Open-World Motion Transfer | Thomas Ressler-Antal et.al. | 2511.23428 | link |
| 2025-11-28 | Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach | Haruki Sakajo et.al. | 2511.23311 | null |
| 2025-11-28 | Vision Bridge Transformer at Scale | Zhenxiong Tan et.al. | 2511.23199 | link |
| 2025-11-28 | GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation | Yuhao Wan et.al. | 2511.23191 | null |
| 2025-11-28 | Fast Multi-view Consistent 3D Editing with Video Priors | Liyi Chen et.al. | 2511.23172 | link |
| 2025-11-28 | InstanceV: Instance-Level Video Generation | Yuheng Chen et.al. | 2511.23146 | null |
| 2025-11-28 | DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation | Hongfei Zhang et.al. | 2511.23127 | link |
| 2025-11-28 | LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models | Zuolei Li et.al. | 2511.23034 | null |
| 2025-11-28 | McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning | Qiushi Yang et.al. | 2511.22974 | link |
| 2025-11-28 | BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation | Zeyu Zhang et.al. | 2511.22973 | null |
| 2025-11-28 | RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video | Haiyang Mei et.al. | 2511.22950 | null |
| 2025-11-28 | One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfe | Shijun Shi et.al. | 2511.22940 | null |
| 2025-11-28 | TARFVAE: Efficient One-Step Generative Time Series Forecasting via TARFLOW based VAE | Jiawen Wei et.al. | 2511.22853 | null |
| 2025-11-28 | Captain Safari: A World Engine | Yu-Cheng Chou et.al. | 2511.22815 | null |
| 2025-11-27 | ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering | Alberto Compagnoni et.al. | 2511.22715 | link |
| 2025-11-27 | Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration | Mengyu Yang et.al. | 2511.22533 | link |
| 2025-11-27 | AI killed the video star. Audio-driven diffusion model for expressive talking head generation | Baptiste Chopin et.al. | 2511.22488 | null |
| 2025-11-27 | Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation | François Provost et.al. | 2511.22467 | null |
| 2025-11-27 | Beyond Real versus Fake Towards Intent-Aware Video Analysis | Saurabh Atreya et.al. | 2511.22455 | null |
| 2025-11-27 | Prompt-based Consistent Video Colorization | Silvia Dani et.al. | 2511.22330 | null |
| 2025-11-27 | Match-and-Fuse: Consistent Generation from Unstructured Image Sets | Kate Feingold et.al. | 2511.22287 | link |
| 2025-11-27 | DriveVGGT: Visual Geometry Transformer for Autonomous Driving | Xiaosong Jia et.al. | 2511.22264 | null |
| 2025-11-27 | VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task | Yuyue Wang et.al. | 2511.22229 | null |
| 2025-11-27 | 3D-Consistent Multi-View Editing by Diffusion Guidance | Josef Bengtson et.al. | 2511.22228 | link |
| 2025-11-27 | IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer | Bo Chen et.al. | 2511.22167 | link |
| 2025-11-27 | EASL: Multi-Emotion Guided Semantic Disentanglement for Expressive Sign Language Generation | Yanchao Zhao et.al. | 2511.22135 | null |
| 2025-11-27 | GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models | Bin Wang et.al. | 2511.22125 | null |
| 2025-11-27 | WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation | Quanjian Song et.al. | 2511.22098 | link |
| 2025-11-27 | GACELLE: GPU-accelerated tools for model parameter estimation and image reconstruction | Kwok-Shing Chan et.al. | 2511.22094 | null |
| 2025-11-27 | AutoRec: Accelerating Loss Recovery for Live Streaming in a Multi-Supplier Market | Tong Li et.al. | 2511.22046 | null |
| 2025-11-26 | Digital Elevation Model Estimation from RGB Satellite Imagery using Generative Deep Learning | Alif Ilham Madani et.al. | 2511.21985 | null |
| 2025-11-26 | Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data | Satrajit Chakrabarty et.al. | 2511.21926 | null |
| 2025-11-26 | Multi-Modal Machine Learning for Early Trust Prediction in Human-AI Interaction Using Face Image and GSR Bio Signals | Hamid Shamszare et.al. | 2511.21908 | null |
| 2025-11-26 | TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos | Seungjae Lee et.al. | 2511.21690 | link |
| 2025-11-26 | Entropy Coding for Non-Rectangular Transform Blocks using Partitioned DCT Dictionaries for AV1 | Priyanka Das et.al. | 2511.21609 | null |
| 2025-11-26 | MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training | Haotian Xue et.al. | 2511.21592 | null |
| 2025-11-26 | Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy | Teng Hu et.al. | 2511.21579 | null |
| 2025-11-26 | Video Generation Models Are Good Latent Reward Models | Xiaoyue Mi et.al. | 2511.21541 | null |
| 2025-11-26 | MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices | Shuai Zhang et.al. | 2511.21475 | link |
| 2025-11-26 | Making sense of quantum teleportation: An intervention study on students' conceptions using a diagrammatic approach | Sebastian Kilde-Westberg et.al. | 2511.21443 | null |
| 2025-11-26 | Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning | Xin Gu et.al. | 2511.21375 | null |
| 2025-11-26 | AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs | Shuhan Xia et.al. | 2511.21251 | null |
| 2025-11-26 | AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control | Xinyue Guo et.al. | 2511.21146 | null |
| 2025-11-26 | TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models | Jiaming He et.al. | 2511.21145 | null |
| 2025-11-26 | Referring Video Object Segmentation with Cross-Modality Proxy Queries | Baoli Sun et.al. | 2511.21139 | link |
| 2025-11-26 | Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning | Changlin Li et.al. | 2511.21136 | null |
| 2025-11-26 | SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation | Ziyi Chen et.al. | 2511.21135 | link |
| 2025-11-26 | CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion | Dianbing Xi et.al. | 2511.21129 | null |
| 2025-11-26 | Dataset Poisoning Attacks on Behavioral Cloning Policies | Akansha Kalra et.al. | 2511.20992 | null |
| 2025-11-26 | TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs | Md Adnan Arefeen et.al. | 2511.20965 | null |
| 2025-11-25 | V |
Jiancheng Pan et.al. | 2511.20886 | null |
| 2025-11-25 | Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries | Sree Bhattacharyya et.al. | 2511.20854 | link |
| 2025-11-25 | Layer-Aware Video Composition via Split-then-Merge | Ozgur Kara et.al. | 2511.20809 | null |
| 2025-11-25 | Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout | Hidir Yesiltepe et.al. | 2511.20649 | null |
| 2025-11-25 | Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization | Tahira Kazimi et.al. | 2511.20647 | null |
| 2025-11-25 | MotionV2V: Editing Motion in a Video | Ryan Burgert et.al. | 2511.20640 | null |
| 2025-11-25 | iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation | Zhoujie Fu et.al. | 2511.20635 | link |
| 2025-11-25 | MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models | Chieh-Yun Chen et.al. | 2511.20629 | link |
| 2025-11-25 | ShapeGen: Towards High-Quality 3D Shape Synthesis | Yangguang Li et.al. | 2511.20624 | null |
| 2025-11-25 | E2E-GRec: An End-to-End Joint Training Framework for Graph Neural Networks and Recommender Systems | Rui Xue et.al. | 2511.20564 | null |
| 2025-11-25 | A Reason-then-Describe Instruction Interpreter for Controllable Video Generation | Shengqiong Wu et.al. | 2511.20563 | link |
| 2025-11-25 | PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding | Haoze Zhang et.al. | 2511.20562 | null |
| 2025-11-25 | Dance Style Classification using Laban-Inspired and Frequency-Domain Motion Features | Ben Hamscher et.al. | 2511.20469 | link |
| 2025-11-25 | STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flow | Jiatao Gu et.al. | 2511.20462 | null |
| 2025-11-25 | Block Cascading: Training Free Acceleration of Block-Causal Video Models | Hmrishav Bandyopadhyay et.al. | 2511.20426 | null |
| 2025-11-25 | TReFT: Taming Rectified Flow Models For One-Step Image Translation | Shengqian Li et.al. | 2511.20307 | null |
| 2025-11-25 | Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations | Chao Wang et.al. | 2511.20295 | link |
| 2025-11-25 | Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement | Yang Liu et.al. | 2511.20280 | null |
| 2025-11-25 | VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs | Tianxiang Jiang et.al. | 2511.20272 | link |
| 2025-11-25 | Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation | Daniel Kienzle et.al. | 2511.20250 | link |
| 2025-11-25 | GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering | Dionysia Danai Brilli et.al. | 2511.20201 | null |
| 2025-11-25 | SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA | Haibin He et.al. | 2511.20190 | link |
| 2025-11-25 | Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis | Mohammad Mahdi et.al. | 2511.20186 | null |
| 2025-11-24 | VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection | Qiang Wang et.al. | 2511.19436 | link |
| 2025-11-24 | Are Image-to-Video Models Good Zero-Shot Image Editors? | Zechuan Zhang et.al. | 2511.19435 | null |
| 2025-11-24 | In-Video Instructions: Visual Signals as Generative Control | Gongfan Fang et.al. | 2511.19401 | link |
| 2025-11-24 | Growing with the Generator: Self-paced GRPO for Video Generation | Rui Li et.al. | 2511.19356 | null |
| 2025-11-24 | MonoMSK: Monocular 3D Musculoskeletal Dynamics Estimation | Farnoosh Koleini et.al. | 2511.19326 | null |
| 2025-11-24 | SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation | Jiaming Zhang et.al. | 2511.19320 | link |
| 2025-11-24 | SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis | Lingwei Dang et.al. | 2511.19319 | link |
| 2025-11-24 | LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models | Shuai Wang et.al. | 2511.19261 | null |
| 2025-11-24 | IDSplat: Instance-Decomposed 3D Gaussian Splatting for Driving Scenes | Carl Lindström et.al. | 2511.19235 | link |
| 2025-11-24 | Learning Plug-and-play Memory for Guiding Video Diffusion Models | Selena Song et.al. | 2511.19229 | link |
| 2025-11-24 | AvatarBrush: Monocular Reconstruction of Gaussian Avatars with Intuitive Local Editing | Mengtian Li et.al. | 2511.19189 | null |
| 2025-11-24 | RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning | Deyi Ji et.al. | 2511.19168 | null |
| 2025-11-24 | HABIT: Human Action Benchmark for Interactive Traffic in CARLA | Mohan Ramesh et.al. | 2511.19109 | null |
| 2025-11-24 | Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation | Ruojun Xu et.al. | 2511.19049 | null |
| 2025-11-24 | View-Consistent Diffusion Representations for 3D-Consistent Video Generation | Duolikun Danier et.al. | 2511.18991 | null |
| 2025-11-24 | Eevee: Towards Close-up High-resolution Video-based Virtual Try-on | Jianhao Zeng et.al. | 2511.18957 | link |
| 2025-11-24 | One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control | Zhenxing Mi et.al. | 2511.18922 | null |
| 2025-11-24 | EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models | Wenhao Xu et.al. | 2511.18920 | null |
| 2025-11-24 | Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation | Ruiying Liu et.al. | 2511.18919 | null |
| 2025-11-24 | MagicWorld: Interactive Geometry-driven Video World Exploration | Guangyuan Li et.al. | 2511.18886 | null |
| 2025-11-21 | EvDiff: High Quality Video with an Event Camera | Weilun Li et.al. | 2511.17492 | null |
| 2025-11-21 | Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination | Yolo Yunlong Tang et.al. | 2511.17490 | null |
| 2025-11-21 | Counterfactual World Models via Digital Twin-conditioned Video Diffusion | Yiqing Shen et.al. | 2511.17481 | null |
| 2025-11-21 | Planning with Sketch-Guided Verification for Physics-Aware Video Generation | Yidong Huang et.al. | 2511.17450 | link |
| 2025-11-21 | Learning Latent Transmission and Glare Maps for Lens Veiling Glare Removal | Xiaolong Qian et.al. | 2511.17353 | link |
| 2025-11-21 | Loomis Painter: Reconstructing the Painting Process | Markus Pobitzer et.al. | 2511.17344 | link |
| 2025-11-21 | Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM | Chiori Hori et.al. | 2511.17335 | null |
| 2025-11-21 | FORWARD: Dataset of a forwarder operating in rough terrain | Mikael Lundbäck et.al. | 2511.17318 | null |
| 2025-11-21 | PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention | Yipeng Chen et.al. | 2511.17185 | null |
| 2025-11-21 | Investigating self-supervised representations for audio-visual deepfake detection | Dragos-Alexandru Boldisor et.al. | 2511.17181 | null |
| 2025-11-21 | OmniLens++: Blind Lens Aberration Correction via Large LensLib Pre-Training and Latent PSF Representation | Qi Jiang et.al. | 2511.17126 | null |
| 2025-11-21 | Sparse Reasoning is Enough: Biological-Inspired Framework for Video Anomaly Detection with Large Pre-trained Models | He Huang et.al. | 2511.17094 | null |
| 2025-11-21 | H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation | Yijie Zhu et.al. | 2511.17079 | null |
| 2025-11-21 | Feature Partitioning and Semantic Equalization for Intrinsic Robustness in Semantic Communication under Packet Loss | Xiao Yang et.al. | 2511.16983 | null |
| 2025-11-21 | MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis | Di Luo et.al. | 2511.16957 | null |
| 2025-11-21 | Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models | Dailan He et.al. | 2511.16955 | null |
| 2025-11-21 | Point-Supervised Facial Expression Spotting with Gaussian-Based Instance-Adaptive Intensity Modeling | Yicheng Deng et.al. | 2511.16952 | null |
| 2025-11-21 | FingerCap: Fine-grained Finger-level Hand Motion Captioning | Xin Shen et.al. | 2511.16951 | null |
| 2025-11-21 | Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features | Jingyi Xu et.al. | 2511.16928 | null |
| 2025-11-21 | R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios | Lu Zhu et.al. | 2511.16901 | link |
| 2025-11-20 | Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO | Junhao Cheng et.al. | 2511.16669 | link |
| 2025-11-20 | V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models | Yang Luo et.al. | 2511.16668 | null |
| 2025-11-20 | SAM2S: Segment Anything in Surgical Videos via Semantic Long-term Tracking | Haofeng Liu et.al. | 2511.16618 | link |
| 2025-11-20 | TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding | Boshen Xu et.al. | 2511.16595 | link |
| 2025-11-20 | An analytical and experimental study of the energy transition discourse on YouTube | Aleix Bassolas et.al. | 2511.16497 | null |
| 2025-11-20 | Flow and Depth Assisted Video Prediction with Latent Transformer | Eliyas Suleyman et.al. | 2511.16484 | null |
| 2025-11-20 | Dynamic Multiple-Parameter Joint Time-Vertex Fractional Fourier Transform and its Intelligent Filtering Methods | Manjun Cui et.al. | 2511.16277 | null |
| 2025-11-20 | PIPHEN: Physical Interaction Prediction with Hamiltonian Energy Networks | Kewei Chen et.al. | 2511.16200 | null |
| 2025-11-20 | FOOTPASS: A Multi-Modal Multi-Agent Tactical Context Dataset for Play-by-Play Action Spotting in Soccer Broadcast Videos | Jeremie Ochin et.al. | 2511.16183 | null |
| 2025-11-20 | Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight | Yi Yang et.al. | 2511.16175 | link |
| 2025-11-20 | Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning | Yibin Huang et.al. | 2511.16160 | link |
| 2025-11-20 | MagBotSim: Physics-Based Simulation and Reinforcement Learning Environments for Magnetic Robotics | Lara Bergmann et.al. | 2511.16158 | null |
| 2025-11-20 | Degradation-Aware Hierarchical Termination for Blind Quality Enhancement of Compressed Video | Li Yu et.al. | 2511.16137 | null |
| 2025-11-20 | Decoupling Complexity from Scale in Latent Diffusion Model | Tianxiong Zhong et.al. | 2511.16117 | null |
| 2025-11-20 | VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning | Zishan Xu et.al. | 2511.16077 | null |
| 2025-11-20 | Panel-by-Panel Souls: A Performative Workflow for Expressive Faces in AI-Assisted Manga Creation | Qing Zhang et.al. | 2511.16038 | null |
| 2025-11-20 | Physically Realistic Sequence-Level Adversarial Clothing for Robust Human-Detection Evasion | Dingkun Zhou et.al. | 2511.16020 | null |
| 2025-11-20 | Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click | Raphael Ruschel et.al. | 2511.15948 | null |
| 2025-11-20 | Automated Interpretable 2D Video Extraction from 3D Echocardiography | Milos Vukadinovic et.al. | 2511.15946 | null |
| 2025-11-19 | RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification | Meilong Xu et.al. | 2511.15923 | null |
| 2025-11-19 | First Frame Is the Place to Go for Video Content Customization | Jingxi Chen et.al. | 2511.15700 | link |
| 2025-11-19 | Joint Semantic-Channel Coding and Modulation for Token Communications | Jingkai Ying et.al. | 2511.15699 | null |
| 2025-11-19 | The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification | Dante Francisco Wasmuht et.al. | 2511.15622 | null |
| 2025-11-19 | Multimodal Evaluation of Russian-language Architectures | Artem Chervyakov et.al. | 2511.15552 | link |
| 2025-11-19 | Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners | Xabier Lekunberri et.al. | 2511.15468 | null |
| 2025-11-19 | ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation | Simon Boeder et.al. | 2511.15396 | null |
| 2025-11-19 | A Multimodal Transformer Approach for UAV Detection and Aerial Object Recognition Using Radar, Audio, and Video Data | Mauro Larrat et.al. | 2511.15312 | null |
| 2025-11-19 | PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback | Sirui Chen et.al. | 2511.15253 | null |
| 2025-11-19 | Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation | Firdavs Nasriddinov et.al. | 2511.15159 | null |
| 2025-11-19 | MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm | Vineeth Sai Narajala et.al. | 2511.15097 | null |
| 2025-11-19 | Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks | Cheng Yang et.al. | 2511.15065 | null |
| 2025-11-19 | Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation | Vladimir Arkhipkin et.al. | 2511.14993 | null |
| 2025-11-18 | Reconstruction of three-dimensional shapes of normal and disease-related erythrocytes from partial observations using multi-fidelity neural networks | Haizhou Wen et.al. | 2511.14962 | null |
| 2025-11-18 | GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis | Antonio Ruiz et.al. | 2511.14884 | null |
| 2025-11-18 | Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising | Yifan Wang et.al. | 2511.14719 | link |
| 2025-11-18 | FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation | Yunfeng Wu et.al. | 2511.14712 | link |
| 2025-11-18 | NERD: Network-Regularized Diffusion Sampling For 3D Computed Tomography | Shijun Liang et.al. | 2511.14680 | null |
| 2025-11-18 | ForensicFlow: A Tri-Modal Adaptive Network for Robust Deepfake Detection | Mohammad Romani et.al. | 2511.14554 | link |
| 2025-11-18 | DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation | Xiangchen Yin et.al. | 2511.14530 | null |
| 2025-11-18 | ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries | Junfu Pu et.al. | 2511.14349 | null |
| 2025-11-18 | Dental3R: Geometry-Aware Pairing for Intraoral 3D Reconstruction from Sparse-View Photographs | Yiyi Miao et.al. | 2511.14315 | null |
| 2025-11-18 | Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning | Rui Liu et.al. | 2511.14249 | link |
| 2025-11-18 | TailCue: Exploring Animal-inspired Robotic Tail for Automated Vehicles Interaction | Yuan Li et.al. | 2511.14242 | null |
| 2025-11-18 | InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion Prior | Weimin Bai et.al. | 2511.14208 | link |
| 2025-11-18 | Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion | Zhuo Li et.al. | 2511.14178 | null |
| 2025-11-18 | Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning Framework with Vision-Language Models | Hao Zhen et.al. | 2511.14120 | null |
| 2025-11-18 | Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services | Liuyi Jin et.al. | 2511.14119 | null |
| 2025-11-18 | A Patient-Independent Neonatal Seizure Prediction Model Using Reduced Montage EEG and ECG | Sithmini Ranasingha et.al. | 2511.14110 | null |
| 2025-11-18 | Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations | Yiqing Shen et.al. | 2511.14100 | null |
| 2025-11-17 | Learning Skill-Attributes for Transferable Assessment in Video | Kumar Ashutosh et.al. | 2511.13993 | link |
| 2025-11-17 | PoCGM: Poisson-Conditioned Generative Model for Sparse-View CT Reconstruction | Changsheng Fang et.al. | 2511.13967 | null |
| 2025-11-17 | SAE-MCVT: A Real-Time and Scalable Multi-Camera Vehicle Tracking Framework Powered by Edge Computing | Yuqiang Lin et.al. | 2511.13904 | null |
| 2025-11-17 | Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors | Mert Onur Cakiroglu et.al. | 2511.13897 | null |
| 2025-11-17 | Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark | Xinxin Liu et.al. | 2511.13853 | null |
| 2025-11-17 | Segment Anything Across Shots: A Method and Benchmark | Hengrui Hu et.al. | 2511.13715 | link |
| 2025-11-17 | UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity | Junwei Yu et.al. | 2511.13714 | link |
| 2025-11-17 | OpenRoboCare: A Multimodal Multi-Task Expert Demonstration Dataset for Robot Caregiving | Xiaoyu Liang et.al. | 2511.13707 | null |
| 2025-11-17 | TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models | Harold Haodong Chen et.al. | 2511.13704 | link |
| 2025-11-17 | Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting | Jiangnan Ye et.al. | 2511.13684 | null |
| 2025-11-17 | CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding | Shrenik Patel et.al. | 2511.13644 | null |
| 2025-11-17 | Smooth Total variation Regularization for Interference Detection and Elimination (STRIDE) for MRI | Alexander Mertens et.al. | 2511.13628 | null |
| 2025-11-17 | Computer Vision based group activity detection and action spotting | Narthana Sivalingam et.al. | 2511.13315 | null |
| 2025-11-17 | PyPeT: A Python Perfusion Tool for Automated Quantitative Brain CT and MR Perfusion Analysis | Marijn Borghouts et.al. | 2511.13310 | null |
| 2025-11-17 | CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving | Enhui Ma et.al. | 2511.13297 | null |
| 2025-11-17 | Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models | Noam Tsfaty et.al. | 2511.13276 | null |
| 2025-11-17 | FoleyBench: A Benchmark For Video-to-Audio Models | Satvik Dixit et.al. | 2511.13219 | null |
| 2025-11-17 | End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer | Yonghui Yu et.al. | 2511.13208 | link |
| 2025-11-17 | RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection | Junhee Lee et.al. | 2511.13204 | null |
| 2025-11-17 | Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification | Rifen Lin et.al. | 2511.13150 | link |
| 2025-11-17 | VEIL: Jailbreaking Text-to-Video Models via Visual Exploitation from Implicit Language | Zonghao Ying et.al. | 2511.13127 | link |
| 2025-11-17 | CloseUpShot: Close-up Novel View Synthesis from Sparse-views via Point-conditioned Diffusion Model | Yuqi Zhang et.al. | 2511.13121 | null |
| 2025-11-17 | Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining | Zhaocheng Yu et.al. | 2511.13113 | null |
| 2025-11-17 | ViSS-R1: Self-Supervised Reinforcement Video Reasoning | Bo Fang et.al. | 2511.13054 | link |
| 2025-11-17 | Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention | Taiye Chen et.al. | 2511.12940 | null |
| 2025-11-14 | Scalable Policy Evaluation with Video World Models | Wei-Cheng Tseng et.al. | 2511.11520 | link |
| 2025-11-14 | Disentangling Emotional Bases and Transient Fluctuations: A Low-Rank Sparse Decomposition Approach for Video Affective Analysis | Feng-Qi Cui et.al. | 2511.11406 | null |
| 2025-11-14 | YCB-Ev SD: Synthetic event-vision dataset for 6DoF object pose estimation | Pavel Rojtberg et.al. | 2511.11344 | null |
| 2025-11-14 | RealisticDreamer: Guidance Score Distillation for Few-shot Gaussian Splatting | Ruocheng Wu et.al. | 2511.11213 | null |
| 2025-11-14 | VIDEOP2R: Video Understanding from Perception to Reasoning | Yifan Jiang et.al. | 2511.11113 | null |
| 2025-11-14 | A Space-Time Transformer for Precipitation Forecasting | Levi Harris et.al. | 2511.11090 | link |
| 2025-11-14 | LiteAttention: A Temporal Sparse Attention for Diffusion Transformers | Dor Shmilovich et.al. | 2511.11062 | null |
| 2025-11-14 | EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation | Zongyang Qiu et.al. | 2511.11002 | null |
| 2025-11-14 | Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition | Gunho Jung et.al. | 2511.10958 | null |
| 2025-11-14 | Language-Guided Graph Representation Learning for Video Summarization | Wenrui Li et.al. | 2511.10953 | null |
| 2025-11-14 | DINOv3 as a Frozen Encoder for CRPS-Oriented Probabilistic Rainfall Nowcasting | Luciano Araujo Dourado Filho et.al. | 2511.10894 | link |
| 2025-11-14 | Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling | Seoik Jung et.al. | 2511.10866 | null |
| 2025-11-13 | From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring | Syed Mumtahin Mahmud et.al. | 2511.10806 | null |
| 2025-11-13 | Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals | Shruti Singh Baghel et.al. | 2511.10615 | link |
| 2025-11-13 | Dynamic Avatar-Scene Rendering from Human-centric Context | Wenqing Wang et.al. | 2511.10539 | link |
| 2025-11-14 | RodEpil: A Video Dataset of Laboratory Rodents for Seizure Detection and Benchmark Evaluation | Daniele Perlo et.al. | 2511.10431 | null |
| 2025-11-13 | TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding | Jinxuan Li et.al. | 2511.10241 | null |
| 2025-11-13 | Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization | Ashutosh Anshul et.al. | 2511.10212 | link |
| 2025-11-13 | Using an instrumented hammer during Summers osteotomy: an animal model | Yasuhiro Homma et.al. | 2511.10126 | null |
| 2025-11-13 | An Instrumented Hammer to Detect the Bone Transitions During an High Tibial Osteotomy: An Animal Study | Bas-Dit-Nugues Manon et.al. | 2511.10121 | null |
| 2025-11-13 | SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition | Qilang Ye et.al. | 2511.10091 | link |
| 2025-11-13 | When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion? | Qilang Ye et.al. | 2511.10059 | link |
| 2025-11-13 | Reinforcing Trustworthiness in Multimodal Emotional Support Systems | Huy M. Le et.al. | 2511.10011 | null |
| 2025-11-13 | Learning phase diversity for solving ill-posed inverse problems in imaging | Jasleen Birdi et.al. | 2511.09952 | null |
| 2025-11-12 | Density Estimation and Crowd Counting | Balachandra Devarangadi Sunil et.al. | 2511.09723 | link |
| 2025-11-12 | PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild | Felix B. Mueller et.al. | 2511.09675 | null |
| 2025-11-12 | TempRetinex: Retinex-based Unsupervised Enhancement for Low-light Video Under Diverse Lighting Conditions | Yini Li et.al. | 2511.09609 | null |
| 2025-11-12 | Bridging the Data Gap: Spatially Conditioned Diffusion Model for Anomaly Generation in Photovoltaic Electroluminescence Images | Shiva Hanifi et.al. | 2511.09604 | null |
| 2025-11-12 | SPIDER: Scalable Physics-Informed Dexterous Retargeting | Chaoyi Pan et.al. | 2511.09484 | link |
| 2025-11-12 | Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models | Ying Peng et.al. | 2511.09469 | null |
| 2025-11-12 | Hand Held Multi-Object Tracking Dataset in American Football | Rintaro Otsubo et.al. | 2511.09455 | null |
| 2025-11-12 | MCAD: Multimodal Context-Aware Audio Description Generation For Soccer | Lipisha Chaudhary et.al. | 2511.09448 | null |
| 2025-11-12 | Augment to Augment: Diverse Augmentations Enable Competitive Ultra-Low-Field MRI Enhancement | Felix F Zimmermann et.al. | 2511.09366 | link |
| 2025-11-10 | Robot Learning from a Physical World Model | Jiageng Mao et.al. | 2511.07416 | link |
| 2025-11-10 | StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation | Tianrui Feng et.al. | 2511.07399 | null |
| 2025-11-10 | ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search | Zhenjie Liu et.al. | 2511.06833 | null |
| 2025-11-09 | GenAI vs. Human Creators: Procurement Mechanism Design in Two-/Three-Layer Markets | Rui Ai et.al. | 2511.06559 | null |
| 2025-11-08 | Neodragon: Mobile Video Generation using Diffusion Transformer | Animesh Karnewar et.al. | 2511.06055 | link |
| 2025-11-06 | Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm | Jingqi Tong et.al. | 2511.04570 | null |
| 2025-11-06 | RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation | Xiangjun Zhang et.al. | 2511.04317 | null |
| 2025-11-07 | PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection | Peiyao Wang et.al. | 2511.03997 | null |
| 2025-11-05 | Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising | Shuangquan Lyu et.al. | 2511.03272 | null |
| 2025-11-03 | How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment | Zhen Chen et.al. | 2511.01775 | null |
| 2025-11-03 | Towards One-step Causal Video Generation via Adversarial Self-Distillation | Yongqi Yang et.al. | 2511.01419 | link |
| 2025-11-02 | Anatomically Constrained Transformers for Echocardiogram Analysis | Alexander Thorley et.al. | 2511.01109 | null |
| 2025-11-04 | ID-Composer: Multi-Subject Video Synthesis with Hierarchical Identity Preservation | Panwang Pan et.al. | 2511.00511 | null |
| 2025-11-01 | Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models | Panwang Pan et.al. | 2511.00503 | link |
| 2025-10-31 | Object-Aware 4D Human Motion Generation | Shurui Gui et.al. | 2511.00248 | null |
| 2025-10-31 | Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals | Xiangyu Fan et.al. | 2510.27684 | null |
| 2025-10-31 | DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model | Yucheng Xing et.al. | 2510.27169 | null |
| 2025-10-30 | Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark | Ziyu Guo et.al. | 2510.26802 | null |
| 2025-10-30 | Co-Evolving Latent Action World Models | Yucen Wang et.al. | 2510.26433 | null |
| 2025-10-29 | 4-Doodle: Text to 3D Sketches that Move! | Hao Chen et.al. | 2510.25319 | null |
| 2025-10-28 | VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos | Qiucheng Wu et.al. | 2510.24904 | null |
| 2025-11-05 | Generative View Stitching | Chonghyuk Song et.al. | 2510.24718 | link |
| 2025-11-03 | Rethinking Visual Intelligence: Insights from Video Pretraining | Pablo Acuaviva et.al. | 2510.24448 | link |
| 2025-10-27 | Yesnt: Are Diffusion Relighting Models Ready for Capture Stage Compositing? A Hybrid Alternative to Bridge the Gap | Elisabeth Jüttner et.al. | 2510.23494 | null |
| 2025-10-28 | LongCat-Video Technical Report | Meituan LongCat Team et.al. | 2510.22200 | null |
| 2025-10-22 | Improving the Physics of Video Generation with VJEPA-2 Reward Signal | Jianhao Yuan et.al. | 2510.21840 | null |
| 2025-10-24 | Epipolar Geometry Improves Video Generation Models | Orest Kupyn et.al. | 2510.21615 | link |
| 2025-10-27 | Video-As-Prompt: Unified Semantic Control for Video Generation | Yuxuan Bian et.al. | 2510.20888 | link |
| 2025-10-23 | AutoScape: Geometry-Consistent Long-Horizon Scene Generation | Jiacheng Chen et.al. | 2510.20726 | null |
| 2025-10-23 | Evaluating Video Models as Simulators of Multi-Person Pedestrian Trajectories | Aaron Appelle et.al. | 2510.20182 | null |
| 2025-10-23 | Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning | Takehiro Aoshima et.al. | 2510.19193 | null |
| 2025-10-21 | MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models | Aritra Bhowmik et.al. | 2510.19022 | null |
| 2025-10-21 | UltraGen: High-Resolution Video Generation with Hierarchical Attention | Teng Hu et.al. | 2510.18775 | null |
| 2025-10-21 | MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation | Weinan Jia et.al. | 2510.18692 | null |
| 2025-10-21 | Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model | Zhenxing Zhang et.al. | 2510.18573 | link |
| 2025-10-22 | MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models | Yongshun Zhang et.al. | 2510.17519 | link |
| 2025-10-20 | From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models | Zefan Cai et.al. | 2510.17247 | null |
| 2025-10-16 | RealDPO: Real or Not Real, that is the Preference | Guo Cheng et.al. | 2510.14955 | null |
| 2025-10-16 | DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation | Yu Zhou et.al. | 2510.14949 | null |
| 2025-10-22 | ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints | Meiqi Wu et.al. | 2510.14847 | link |
| 2025-10-16 | In-Context Learning with Unpaired Clips for Instruction-based Video Editing | Xinyao Liao et.al. | 2510.14648 | link |
| 2025-10-16 | Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures | Yuancheng Xu et.al. | 2510.14179 | link |
| 2025-10-15 | PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning | Sihui Ji et.al. | 2510.13809 | link |
| 2025-10-15 | FlashWorld: High-quality 3D Scene Generation within Seconds | Xinyang Li et.al. | 2510.13678 | link |
| 2025-10-14 | MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars | Felix Taubner et.al. | 2510.12785 | link |
| 2025-10-14 | G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior | Junfeng Ni et.al. | 2510.12099 | link |
| 2025-10-13 | Point Prompting: Counterfactual Tracking with Video Diffusion Models | Ayush Shrivastava et.al. | 2510.11715 | null |
| 2025-10-14 | LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference | Jianhao Yuan et.al. | 2510.11512 | null |
| 2025-10-12 | AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes | Yu Li et.al. | 2510.10670 | null |
| 2025-10-13 | Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization | Shuo Xing et.al. | 2510.08789 | null |
| 2025-10-09 | NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos | Hongyu Li et.al. | 2510.08568 | null |
| 2025-10-11 | MultiCOIN: Multi-Modal COntrollable Video INbetweening | Maham Tanveer et.al. | 2510.08561 | null |
| 2025-10-09 | VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning | Minghong Cai et.al. | 2510.08555 | link |
| 2025-10-09 | Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency | Kaiwen Zheng et.al. | 2510.08431 | null |
| 2025-10-09 | LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation | Yushi Huang et.al. | 2510.08318 | null |
| 2025-10-09 | SViM3D: Stable Video Material Diffusion for Single Image 3D Generation | Andreas Engelhardt et.al. | 2510.08271 | null |
| 2025-10-09 | UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution | Shian Du et.al. | 2510.08143 | link |
| 2025-10-15 | Real-Time Motion-Controllable Autoregressive Video Diffusion | Kesen Zhao et.al. | 2510.08131 | null |
| 2025-10-16 | CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving | Tianrui Zhang et.al. | 2510.07944 | null |
| 2025-10-09 | Controllable Video Synthesis via Variational Inference | Haoyi Duan et.al. | 2510.07670 | null |
| 2025-10-09 | Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection | Yanjie Pan et.al. | 2510.07654 | null |
| 2025-10-08 | TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility | Saman Motamed et.al. | 2510.07550 | link |
| 2025-10-07 | Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model | Danush Kumar Venkatesh et.al. | 2510.07345 | null |
| 2025-10-08 | WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation | Zezhong Qian et.al. | 2510.07313 | null |
| 2025-10-08 | MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis | Yihao Zhi et.al. | 2510.07190 | link |
| 2025-10-07 | Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models | Jiahao Wang et.al. | 2510.06209 | null |
| 2025-10-06 | VChain: Chain-of-Visual-Thought for Reasoning in Video Generation | Ziqi Huang et.al. | 2510.05094 | link |
| 2025-10-06 | Bridging Text and Video Generation: A Survey | Nilay Kumar et.al. | 2510.04999 | null |
| 2025-10-05 | ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation | Jay Zhangjie Wu et.al. | 2510.04290 | link |
| 2025-10-04 | Generating Human Motion Videos using a Cascaded Text-to-Video Framework | Hyelin Nam et.al. | 2510.03909 | link |
| 2025-10-04 | Towards Robust and Generalizable Continuous Space-Time Video Super-Resolution with Events | Shuoyan Wei et.al. | 2510.03833 | link |
| 2025-10-03 | Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime! | Junbao Zhou et.al. | 2510.03550 | null |
| 2025-10-03 | Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft | Junchao Huang et.al. | 2510.03198 | link |
| 2025-10-02 | How Confident are Video Models? Empowering Video Models to Express their Uncertainty | Zhiting Mei et.al. | 2510.02571 | link |
| 2025-10-02 | Learning to Generate Object Interactions with Physics-Guided Video Diffusion | David Romero et.al. | 2510.02284 | null |
| 2025-10-02 | TempoControl: Temporal Attention Guidance for Text-to-Video Models | Shira Schiber et.al. | 2510.02226 | link |
| 2025-10-03 | UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction | Jin Cao et.al. | 2510.01669 | link |
| 2025-10-01 | IMAGEdit: Let Any Subject Transform | Fei Shen et.al. | 2510.01186 | link |
| 2025-10-01 | Can World Models Benefit VLMs for World Dynamics? | Kevin Zhang et.al. | 2510.00855 | null |
| 2025-10-01 | From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation | Fan Yang et.al. | 2510.00806 | null |
| 2025-10-01 | BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration | Zhaoyang Li et.al. | 2510.00438 | link |
| 2025-09-27 | Object-AVEdit: An Object-level Audio-Visual Editing Model | Youquan Fu et.al. | 2510.00050 | null |
| 2025-09-30 | Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation | Agneet Chatterjee et.al. | 2509.26555 | null |
| 2025-09-30 | MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation | Chenhui Zhu et.al. | 2509.26391 | link |
| 2025-09-30 | PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution | Shian Du et.al. | 2509.26025 | null |
| 2025-09-29 | DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder | Junyu Chen et.al. | 2509.25182 | link |
| 2025-09-29 | Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer | Mohsen Ghafoorian et.al. | 2509.24899 | null |
| 2025-09-30 | Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility | Yutong Hao et.al. | 2509.24702 | null |
| 2025-09-29 | CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers | Kai Liu et.al. | 2509.24416 | null |
| 2025-09-29 | NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis | Yixuan Ren et.al. | 2509.24353 | null |
| 2025-09-28 | ReLumix: Extending Image Relighting to Video via Video Diffusion Models | Lezhong Wang et.al. | 2509.23769 | null |
| 2025-09-28 | VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement | Shulian Zhang et.al. | 2509.23584 | null |
| 2025-09-27 | WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving | Ziyue Zhu et.al. | 2509.23402 | link |
| 2025-10-01 | Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs | Xingyu Fu et.al. | 2509.22646 | null |
| 2025-09-26 | EgoDemoGen: Novel Egocentric Demonstration Generation Enables Viewpoint-Robust Manipulation | Yuan Xu et.al. | 2509.22578 | null |
| 2025-09-26 | EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer | Zhehao Dong et.al. | 2509.22407 | null |
| 2025-09-26 | From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment | Ke Ye et.al. | 2509.22205 | null |
| 2025-09-29 | MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training | Haoyun Li et.al. | 2509.22199 | null |
| 2025-09-26 | Drag4D: Align Your Motion with Text-Driven 3D Scene Generation | Minjun Kang et.al. | 2509.21888 | null |
| 2025-09-29 | DiTraj: training-free trajectory control for video diffusion transformer | Cheng Lei et.al. | 2509.21839 | link |
| 2025-09-26 | UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models | Lan Chen et.al. | 2509.21760 | link |
| 2025-09-29 | ControlHair: Physically-based Video Diffusion for Controllable Dynamic Hair Rendering | Weikai Lin et.al. | 2509.21541 | null |
| 2025-09-25 | NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics | Yu Yuan et.al. | 2509.21309 | link |
| 2025-09-24 | PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation | Chen Wang et.al. | 2509.20358 | link |
| 2025-09-24 | 4D Driving Scene Generation With Stereo Forcing | Hao Lu et.al. | 2509.20251 | null |
| 2025-09-24 | Anatomically Constrained Transformers for Cardiac Amyloidosis Classification | Alexander Thorley et.al. | 2509.19691 | null |
| 2025-09-24 | From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition | Ling Lo et.al. | 2509.19690 | null |
| 2025-09-23 | Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation | Sherwin Bahmani et.al. | 2509.19296 | link |
| 2025-09-23 | Flow marching for a generative PDE foundation model | Zituo Chen et.al. | 2509.18611 | null |
| 2025-09-22 | VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models | Geonung Kim et.al. | 2509.17985 | link |
| 2025-09-20 | FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers | Sankeerth Durvasula et.al. | 2509.16518 | null |
| 2025-09-20 | RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation | Tianyi Yan et.al. | 2509.16500 | null |
| 2025-09-27 | WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance | Chenxi Song et.al. | 2509.15130 | link |
| 2025-09-18 | BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching | Hanshuai Cui et.al. | 2509.13789 | link |
| 2025-09-17 | TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving | Jiawei Wang et.al. | 2509.13164 | null |
| 2025-09-15 | MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment | Yanyun Pu et.al. | 2509.11589 | null |
| 2025-09-14 | PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation | Zeyu Dong et.al. | 2509.11092 | null |
| 2025-09-12 | T2Bs: Text-to-Character Blendshapes via Video Generation | Jiahao Luo et.al. | 2509.10678 | null |
| 2025-09-11 | Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders | Dohun Lee et.al. | 2509.09547 | null |
| 2025-09-09 | LINR Bridge: Vector Graphic Animation via Neural Implicits and Video Diffusion Priors | Wenshuo Gao et.al. | 2509.07484 | null |
| 2025-09-09 | ANYPORTAL: Zero-Shot Consistent Video Background Replacement | Wenshuo Gao et.al. | 2509.07472 | null |
| 2025-09-08 | From Rigging to Waving: 3D-Guided Diffusion for Natural Animation of Hand-Drawn Characters | Jie Zhou et.al. | 2509.06573 | link |
| 2025-09-16 | BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models | Yuming Li et.al. | 2509.06040 | link |
| 2025-08-29 | ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory | Ying Li et.al. | 2509.05314 | null |
| 2025-09-04 | Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview | Jun-Kun Chen et.al. | 2509.04450 | null |
| 2025-09-01 | Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement | Jiayi Gao et.al. | 2509.01362 | null |
| 2025-08-31 | Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion | Xueyang Kang et.al. | 2509.00843 | link |
| 2025-08-30 | DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective | Yushuo Chen et.al. | 2509.00403 | null |
| 2025-08-29 | Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering | Nattapong Kurpukdee et.al. | 2508.21773 | null |
| 2025-08-27 | ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion | Xurui Peng et.al. | 2508.21091 | null |
| 2025-08-28 | POSE: Phased One-Step Adversarial Equilibrium for Video Diffusion Models | Jiaxiang Cheng et.al. | 2508.21019 | null |
| 2025-08-26 | LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding | Julian Ost et.al. | 2508.19204 | null |
| 2025-08-26 | ROSE: Remove Objects with Side Effects in Videos | Chenxuan Miao et.al. | 2508.18633 | link |
| 2025-08-25 | ObjFiller-3D: Consistent Multi-view 3D Inpainting via Video Diffusion Models | Haitang Feng et.al. | 2508.18271 | link |
| 2025-08-24 | A Synthetic Dataset for Manometry Recognition in Robotic Applications | Pedro Antonio Rabelo Saraiva et.al. | 2508.17468 | null |
| 2025-08-24 | Multi-Level LVLM Guidance for Untrimmed Video Action Recognition | Liyang Peng et.al. | 2508.17442 | null |
| 2025-08-24 | MoCo: Motion-Consistent Human Video Generation via Structure-Appearance Decoupling | Haoyu Wang et.al. | 2508.17404 | null |
| 2025-08-26 | Waver: Wave Your Way to Lifelike Video Generation | Yifu Zhang et.al. | 2508.15761 | null |
| 2025-08-27 | VideoEraser: Concept Erasure in Text-to-Video Diffusion Models | Naen Xu et.al. | 2508.15314 | link |
| 2025-08-25 | MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion | Xuyang Chen et.al. | 2508.15169 | null |
| 2025-08-20 | Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving | Leila Cheshmi et.al. | 2508.14729 | null |
| 2025-08-19 | Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing | Feng-Lin Liu et.al. | 2508.13797 | null |
| 2025-08-19 | Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model | Ruixin Zhang et.al. | 2508.13584 | link |
| 2025-08-18 | GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis | Sirshapan Mitra et.al. | 2508.13300 | link |
| 2025-08-18 | 4DNeX: Feed-Forward 4D Generative Modeling Made Easy | Zhaoxi Chen et.al. | 2508.13154 | link |
| 2025-08-18 | Precise Action-to-Video Generation Through Visual Action Prompts | Yuang Wang et.al. | 2508.13104 | null |
| 2025-08-18 | Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models | Jianshu Zeng et.al. | 2508.12945 | null |
| 2025-08-18 | E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model | Ronghao Lin et.al. | 2508.12854 | link |
| 2025-08-15 | CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models | Xiaoxue Wu et.al. | 2508.11484 | link |
| 2025-08-15 | Versatile Video Tokenization with Generative 2D Gaussian Splatting | Zhenghao Chen et.al. | 2508.11183 | null |
| 2025-08-14 | GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning | Kelin Yu et.al. | 2508.11049 | null |
| 2025-08-15 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation | Harold Haodong Chen et.al. | 2508.10858 | null |
| 2025-08-26 | Physical Autoregressive Model for Robotic Manipulation without Action Pretraining | Zijian Song et.al. | 2508.09822 | null |
| 2025-08-13 | GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors | Xingyilang Yin et.al. | 2508.09667 | link |
| 2025-08-21 | Preacher: Paper-to-Video Agentic System | Jingwei Liu et.al. | 2508.09632 | link |
| 2025-08-14 | From Large Angles to Consistent Faces: Identity-Preserving Video Generation via Mixture of Facial Experts | Yuji Wang et.al. | 2508.09476 | link |
| 2025-08-12 | X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents | Guoxian Song et.al. | 2508.09383 | null |
| 2025-08-12 | Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices | Ya Zou et.al. | 2508.09136 | link |
| 2025-08-12 | TaoCache: Structure-Maintained Video Generation Acceleration | Zhentao Fan et.al. | 2508.08978 | null |
| 2025-08-14 | Yan: Foundational Interactive Video Generation | Deheng Ye et.al. | 2508.08601 | null |
| 2025-08-11 | Matrix-3D: Omnidirectional Explorable 3D World Generation | Zhongqi Yang et.al. | 2508.08086 | null |
| 2025-08-11 | S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix | Peng Dai et.al. | 2508.08048 | null |
| 2025-08-12 | Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation | Fangyuan Mao et.al. | 2508.07981 | null |
| 2025-08-11 | Generative Video Matting | Yongtao Ge et.al. | 2508.07905 | link |
| 2025-08-12 | Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation | Bowen Xue et.al. | 2508.07901 | null |
| 2025-08-11 | Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation | Xiaoyan Liu et.al. | 2508.07769 | null |
| 2025-08-11 | Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation | Minghao Yin et.al. | 2508.07557 | null |
| 2025-08-10 | SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models | Ruolin Yang et.al. | 2508.07149 | null |
| 2025-08-09 | eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos | Xuecheng Wu et.al. | 2508.06902 | null |
| 2025-08-08 | Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video | Jixuan He et.al. | 2508.06715 | null |
| 2025-08-08 | FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation | Wenbin Teng et.al. | 2508.06392 | link |
| 2025-08-08 | SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment | Yanxiao Sun et.al. | 2508.06082 | null |
| 2025-08-07 | Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation | Yue Liao et.al. | 2508.05635 | null |
| 2025-08-06 | 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation | Shuzhou Yang et.al. | 2508.04467 | null |
| 2025-08-07 | S |
Weilun Feng et.al. | 2508.04016 | null |
| 2025-08-05 | VideoGuard: Protecting Video Content from Unauthorized Editing | Junjie Cao et.al. | 2508.03480 | null |
| 2025-08-06 | Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation | Xunzhi Xiang et.al. | 2508.03334 | link |
| 2025-08-05 | V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models | Jisoo Kim et.al. | 2508.03254 | null |
| 2025-08-13 | MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention | Qi Xie et.al. | 2508.03034 | null |
| 2025-08-04 | DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework | Tongchun Zuo et.al. | 2508.02807 | link |
| 2025-08-04 | VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling | Yuru Xiao et.al. | 2508.02129 | null |
| 2025-08-03 | Versatile Transition Generation with Image-to-Video Diffusion | Zuhao Yang et.al. | 2508.01698 | null |
| 2025-08-01 | Video Generators are Robot Policies | Junbang Liang et.al. | 2508.00795 | null |
| 2025-08-01 | Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency | Xi Xue et.al. | 2508.00397 | null |
| 2025-08-01 | GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection | Suhang Cai et.al. | 2508.00312 | null |
| 2025-08-01 | TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models | Christian Simon et.al. | 2508.00289 | null |
| 2025-07-31 | World Consistency Score: A Unified Metric for Video Generation Quality | Akshat Rakheja et.al. | 2508.00144 | null |
| 2025-07-30 | GVD: Guiding Video Diffusion Model for Scalable Video Distillation | Kunyang Li et.al. | 2507.22360 | null |
| 2025-07-27 | AnimeColor: Reference-based Animation Colorization with Diffusion Transformers | Yuhong Zhang et.al. | 2507.20158 | link |
| 2025-08-01 | HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly | Chang Liu et.al. | 2507.19924 | link |
| 2025-07-26 | TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection | Suhwan Cho et.al. | 2507.19789 | link |
| 2025-07-25 | RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution | Weisong Zhao et.al. | 2507.19138 | null |
| 2025-07-22 | Controllable Video Generation: A Survey | Yue Ma et.al. | 2507.16869 | link |
| 2025-07-22 | Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model | Mingtao Guo et.al. | 2507.16341 | link |
| 2025-07-22 | PUSA V1.0: Surpassing Wan-I2V with $500 Training Cost by Vectorized Timestep Adaptation | Yaofang Liu et.al. | 2507.16116 | null |
| 2025-07-21 | Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars | Marcel C. Bühler et.al. | 2507.15979 | null |
| 2025-07-21 | Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models | Enes Sanli et.al. | 2507.15824 | null |
| 2025-07-21 | TokensGen: Harnessing Condensed Tokens for Long Video Generation | Wenqi Ouyang et.al. | 2507.15728 | link |
| 2025-07-21 | CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers | Jiaqi Han et.al. | 2507.15260 | link |
| 2025-07-20 | StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation | Shuyuan Tu et.al. | 2507.15064 | null |
| 2025-07-17 | "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models | Jing Gu et.al. | 2507.13428 | null |
| 2025-07-27 | Vidar: Embodied Video Diffusion Model for Generalist Bimanual Manipulation | Yao Feng et.al. | 2507.12898 | null |
| 2025-07-16 | Reconstruct, Inpaint, Finetune: Dynamic Novel-view Synthesis from Monocular Videos | Kaihua Chen et.al. | 2507.12646 | link |
| 2025-07-29 | NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation Models | X. Feng et.al. | 2507.11245 | link |
| 2025-07-10 | Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling | Haoyu Wu et.al. | 2507.07982 | link |
| 2025-07-30 | Scaling RL to Long Videos | Yukang Chen et.al. | 2507.07966 | null |
| 2025-07-09 | A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality | Mohamed Elmoghany et.al. | 2507.07202 | null |
| 2025-07-09 | Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation | Tao Feng et.al. | 2507.06830 | null |
| 2025-07-14 | Democratizing High-Fidelity Co-Speech Gesture Video Generation | Xu Yang et.al. | 2507.06812 | link |
| 2025-07-24 | Bridging Sequential Deep Operator Network and Video Diffusion: Residual Refinement of Spatio-Temporal PDE Solutions | Jaewan Park et.al. | 2507.06133 | null |
| 2025-07-08 | DreamArt: Generating Interactable Articulated Objects from a Single Image | Ruijie Lu et.al. | 2507.05763 | null |
| 2025-07-08 | LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion | Yisu Zhang et.al. | 2507.05678 | null |
| 2025-07-08 | MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos | Rongsheng Wang et.al. | 2507.05675 | link |
| 2025-07-07 | EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling | Boyuan Wang et.al. | 2507.05198 | null |
| 2025-07-08 | StreamDiT: Real-Time Streaming Text-to-Video Generation | Akio Kodaira et.al. | 2507.03745 | null |
| 2025-07-03 | RefTok: Reference-Based Tokenization for Video Generation | Xiang Fan et.al. | 2507.02862 | null |
| 2025-07-03 | Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching | Xin Zhou et.al. | 2507.02860 | link |
| 2025-07-03 | LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion | Fangfu Liu et.al. | 2507.02813 | link |
| 2025-07-03 | DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation | Yunhan Yang et.al. | 2507.02299 | null |
| 2025-07-09 | LongAnimation: Long Animation Generation with Dynamic Global-Local Memory | Nan Chen et.al. | 2507.01945 | null |
| 2025-07-01 | Geometry-aware 4D Video Generation for Robot Manipulation | Zeyi Liu et.al. | 2507.01099 | link |
| 2025-07-01 | DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution | Zhe Kong et.al. | 2507.01012 | link |
| 2025-07-04 | Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations | Shivansh Patel et.al. | 2507.00990 | null |
| 2025-07-01 | Populate-A-Scene: Affordance-Aware Human Video Generation | Mengyi Shan et.al. | 2507.00334 | null |
| 2025-06-30 | FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion | Yu Lu et.al. | 2507.00162 | null |
| 2025-06-30 | Epona: Autoregressive Diffusion World Model for Autonomous Driving | Kaiwen Zhang et.al. | 2506.24113 | link |
| 2025-06-30 | VMoBA: Mixture-of-Block Attention for Video Diffusion Models | Jianzong Wu et.al. | 2506.23858 | link |
| 2025-06-30 | SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation | Shuai Tan et.al. | 2506.23690 | null |
| 2025-06-29 | Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis | Lei-lei Li et.al. | 2506.23263 | null |
| 2025-07-01 | Listener-Rewarded Thinking in VLMs for Image Preferences | Alexander Gambashidze et.al. | 2506.22832 | null |
| 2025-06-27 | Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy | Yuhao Liu et.al. | 2506.22432 | link |
| 2025-06-27 | RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation | Liudi Yang et.al. | 2506.22007 | link |
| 2025-06-27 | FairyGen: Storied Cartoon Video from a Single Child-Drawn Character | Jiayi Zheng et.al. | 2506.21272 | link |
| 2025-06-26 | Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models | Donggoo Kang et.al. | 2506.20946 | null |
| 2025-06-30 | StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation | Haodong Li et.al. | 2506.20756 | null |
| 2025-06-25 | Video Perception Models for 3D Scene Synthesis | Rui Huang et.al. | 2506.20601 | link |
| 2025-06-25 | Feature Hallucination for Self-supervised Action Recognition | Lei Wang et.al. | 2506.20342 | null |
| 2025-06-24 | Radial Attention: |
Xingyang Li et.al. | 2506.19852 | null |
| 2025-06-24 | AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models | Zehuan Huang et.al. | 2506.19851 | link |
| 2025-06-24 | GenHSI: Controllable Generation of Human-Scene Interaction Videos | Zekun Li et.al. | 2506.19840 | null |
| 2025-09-29 | Physics-Guided Motion Loss for Video Generation Model | Bowen Xue et.al. | 2506.02244 | null |
| 2025-05-30 | VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models | Xiangdong Zhang et.al. | 2505.23656 | link |
| 2025-05-01 | ReVision: High-Quality, Low-Cost Video Generation with Explicit 3D Physics Modeling for Complex Motion and Interaction | Qihao Liu et.al. | 2504.21855 | null |
| 2025-04-07 | VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior | Xindi Yang et.al. | 2503.23368 | link |
| 2025-03-28 | Exploring the Evolution of Physics Cognition in Video Generation: A Survey | Minghui Lin et.al. | 2503.21765 | null |
| 2025-11-05 | Rethinking Video Super-Resolution: Towards Diffusion-Based Methods without Motion Alignment | Zhihao Zhan et.al. | 2503.03355 | null |
| 2024-09-30 | PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation | Shaowei Liu et.al. | 2409.18964 | link |
| 2025-03-17 | Tora: Trajectory-oriented Diffusion Transformer for Video Generation | Zhenghao Zhang et.al. | 2407.21705 | link |
| 2024-10-28 | MotionCraft: Physics-based Zero-Shot Video Generation | Luca Savant Aira et.al. | 2405.13557 | null |
| 2024-11-19 | Video Diffusion Models: A Survey | Andrew Melnik et.al. | 2405.03150 | link |
| 2024-02-06 | Lumiere: A Space-Time Diffusion Model for Video Generation | Omer Bar-Tal et.al. | 2401.12945 | link |
| 2023-12-12 | Photorealistic Video Generation with Diffusion Models | Agrim Gupta et.al. | 2312.06662 | null |
| 2023-10-12 | VDT: General-purpose Video Diffusion Transformers via Mask Modeling | Haoyu Lu et.al. | 2305.13311 | link |
| 2023-07-11 | Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos | Kun Su et.al. | 2303.16897 | null |
| 2022-10-06 | Imagen Video: High Definition Video Generation with Diffusion Models | Jonathan Ho et.al. | 2210.02303 | null |