| WALL-WM: Carving World Action Modeling at the Event Joints |
2026-06-01 |
|
| AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance |
2026-05-31 |
|
| GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video |
2026-05-30 |
|
| minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models |
2026-05-28 |
|
| PlayClass: Automated Play Behaviour Classification in Poultry |
2026-05-26 |
Accep...Accepted at CV4Animals Workshop @ CVPR 2026 |
| World-R1: Reinforcing 3D Constraints for Text-to-Video Generation |
2026-05-26 |
ICML ...ICML 2026, Project Page: https://aka.ms/world-r1, Code: https://github.com/microsoft/World-R1 |
| EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation |
2026-05-22 |
|
| Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models |
2026-05-18 |
Accep...Accepted to CVPR 2026 Workshops CV4Smalls |
| Latent Video Prediction Learns Better World Models |
2026-05-15 |
|
| DriveCtrl: Conditioned Sim-to-Real Driving Video Generation |
2026-05-14 |
|
| Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception |
2026-05-12 |
v2: M...v2: Major revision. Retitled; expanded from TimeSformer alone to four backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, CLIP), with V-JEPA 2.1 nearly tripling pretrained performance. Adds zero-shot PHASE transfer, attention-rollout analysis, and a language-distillation control. Data (OOO sim. judgments) & core hybrid triplet+RSA LoRA method unchanged from v1. Prepared for NeurIPS 2026 submission |
| MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation |
2026-05-09 |
17 pages, 9 figues |
| Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios |
2026-05-07 |
|
| UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models |
2026-05-07 |
|
| Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation |
2026-04-16 |
Accep...Accepted at 2026 International Conference on Automatic Face and Gesture Recognition (FG) |
Please check the Github page for a better reading experience and more papers.
Unified
Accep...
Accepted at TMLR. Project page: https://unite-page.github.io/
Accep...
Accepted for publication in IEEE Communications Standards Magazine
Accep...
Accepted by KDD Ads Track 2026
Video Understanding
28 pa...
28 pages, 10 figures, 11 tables
Accep...
Accepted at the IEEE ITSC 2026
Prepr...
Preprint. 12 pages, 6 figures, 7 tables
ICML ...
ICML 2026 Camera-ready
Accep...
Accepted to CVPR 2026. Project page: https://sparsevideounderstanding.github.io
Accep...
Accepted by ICML 2026. Camera-ready version
World Model
Proje...
Project page: https://junjieye.com/RoboDream/
CVPR ...
CVPR 2026 Oral Presentation; 80 pages, 37 figures, 29 tables; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens
28 pa...
28 pages, 7 figures, 16 tables, Su
Proje...
Project page: https://gim-world.github.io/
Accep...
Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Main Conference
Accep...
Accepted by ICML 2026
Multimodal
https...
https://github.com/alexmartin1722/mirage
Softw...
Software: https://www.robots.ox.ac.uk/~vgg/software/wise/ , Online demos: https://www.robots.ox.ac.uk/~vgg/software/wise/demo/ , Example Queries: https://www.robots.ox.ac.uk/~vgg/software/wise/examples/
27 pa...
27 pages, 20 figures, Accepted to the Main Conference of ACL 2026
25 pa...
25 pages, 8 figures, 8 tables. Project page: https://zkangning.github.io/MMSkills_for_Visual_Agents/
Multimodal LLM
Accep...
Accepted to ICML 2026
Accep...
Accepted at ACL 2026 Main Conference. Camera-ready version
Video Foundation Model
Accep...
Accepted at CV4Animals Workshop @ CVPR 2026
ICML ...
ICML 2026, Project Page: https://aka.ms/world-r1, Code: https://github.com/microsoft/World-R1
Accep...
Accepted to CVPR 2026 Workshops CV4Smalls
v2: M...
v2: Major revision. Retitled; expanded from TimeSformer alone to four backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, CLIP), with V-JEPA 2.1 nearly tripling pretrained performance. Adds zero-shot PHASE transfer, attention-rollout analysis, and a language-distillation control. Data (OOO sim. judgments) & core hybrid triplet+RSA LoRA method unchanged from v1. Prepared for NeurIPS 2026 submission
Accep...
Accepted at 2026 International Conference on Automatic Face and Gesture Recognition (FG)