Skip to content

Latest 15 Papers - June 02, 2026 #215

Description

@github-actions

Please check the Github page for a better reading experience and more papers.

Unified

Title Date Comment
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning 2026-06-01
Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models 2026-06-01
Unified Context Evolution for LLM Agents 2026-06-01
CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation 2026-06-01
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments 2026-06-01 34 pages
Unified Semantic Transformer for 3D Scene Understanding 2026-06-01
Accep...

Accepted at TMLR. Project page: https://unite-page.github.io/

A Unified Framework for Structured Flow Modeling: From Continuous Fields to Data-Driven Representations 2026-06-01
A Unified E2E Energy Efficiency Testing Framework for Open RAN 2026-06-01
Accep...

Accepted for publication in IEEE Communications Standards Magazine

Tree-Guided Identify-Then-Exploit: A Unified Framework of Best Arm Identification and Regret Minimization for Dueling Bandits 2026-06-01
TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control 2026-06-01
A Unified Variational Design of Predictive Mirror Descent in Convex Games under Stochastic Feedback 2026-06-01
UniVocal: Unified Speech-Singing Code-Switching Synthesis 2026-06-01 accepted by ACL 2026
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking 2026-06-01
Accep...

Accepted by KDD Ads Track 2026

A Unified Evaluation-Instructed Framework for Query-Dependent Prompt Optimization 2026-06-01
A Unified Framework for Adversary-Aware Differential Privacy Bounds 2026-06-01

Video Understanding

Title Date Comment
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events 2026-06-01
28 pa...

28 pages, 10 figures, 11 tables

Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection 2026-06-01
Accep...

Accepted at the IEEE ITSC 2026

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models 2026-06-01 15 pages, 8 figures
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound 2026-06-01 24 pages, 9 figures
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection 2026-06-01
3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval 2026-06-01
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey 2026-06-01
An Efficient Streaming Video Understanding Framework with Agentic Control 2026-06-01
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models 2026-06-01
Prepr...

Preprint. 12 pages, 6 figures, 7 tables

APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention 2026-06-01 ACL 2026 main
VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding 2026-05-31
V-LynX: Token Interface Alignment for Video+X LLMs 2026-05-30
ICML ...

ICML 2026 Camera-ready

Towards Sparse Video Understanding and Reasoning 2026-05-30
Accep...

Accepted to CVPR 2026. Project page: https://sparsevideounderstanding.github.io

Linear Scaling Video VLMs for Long Video Understanding 2026-05-29
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding 2026-05-29
Accep...

Accepted by ICML 2026. Camera-ready version

World Model

Title Date Comment
RoboDream: Compositional World Models for Scalable Robot Data Synthesis 2026-06-01
Proje...

Project page: https://junjieye.com/RoboDream/

From Zero to Hero: Training-Free Custom Concept Spawning in World Models 2026-06-01
WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World 2026-06-01
CVPR ...

CVPR 2026 Oral Presentation; 80 pages, 37 figures, 29 tables; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens

Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation 2026-06-01
28 pa...

28 pages, 7 figures, 16 tables, Su

Geometry-Aware Implicit Memory for Video World Models 2026-06-01
Proje...

Project page: https://gim-world.github.io/

Policy and World Modeling Co-Training for Language Agents 2026-06-01 9 pages, 6 figures
TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks 2026-06-01
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents 2026-06-01
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation 2026-06-01
World-Task Factorization for Robot Learning 2026-06-01
Scaling Agentic Capabilities via Grounded Interaction Synthesis 2026-06-01
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning 2026-06-01
Accep...

Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Main Conference

Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects 2026-06-01
Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning 2026-06-01
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching 2026-06-01
Accep...

Accepted by ICML 2026

Multimodal

Title Date Comment
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling 2026-06-01 ICML 2026
ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning 2026-06-01
Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design 2026-06-01
CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning 2026-06-01
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation 2026-06-01
https...

https://github.com/alexmartin1722/mirage

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning 2026-06-01
Attention Dynamics and Adaptive Decision Support in C5ISR: A Recurrence Quantification Analysis of Visual and Multimodal Attention Guidance Effects on Mission Performance 2026-06-01 11 Figures, 3 Tables
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains 2026-06-01
WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata 2026-06-01
Softw...

Software: https://www.robots.ox.ac.uk/~vgg/software/wise/ , Online demos: https://www.robots.ox.ac.uk/~vgg/software/wise/demo/ , Example Queries: https://www.robots.ox.ac.uk/~vgg/software/wise/examples/

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality 2026-06-01
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics 2026-06-01
Context-Aware Workflow Decomposition for Automated Mobile UI Annotation Using Multimodal Large Language Models 2026-06-01
Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis 2026-06-01
Jailbreaking Multimodal Large Language Models using Multi-Clip Video 2026-06-01
27 pa...

27 pages, 20 figures, Accepted to the Main Conference of ACL 2026

MMSkills: Towards Multimodal Skills for General Visual Agents 2026-06-01
25 pa...

25 pages, 8 figures, 8 tables. Project page: https://zkangning.github.io/MMSkills_for_Visual_Agents/

Multimodal LLM

Title Date Comment
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling 2026-06-01 ICML 2026
DenseMLLM: Standard Multimodal LLMs for Dense Prediction 2026-06-01 ICML 2026
Enhancing the Socioeconomic Understanding of Foundation Models with Urban Mobility 2026-06-01
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference 2026-06-01
Accep...

Accepted to ICML 2026

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval 2026-06-01
Accep...

Accepted at ACL 2026 Main Conference. Camera-ready version

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification 2026-06-01
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning 2026-06-01 9 pages, 5 figures
AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection 2026-05-31 ICML 2026
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers 2026-05-30 Paper under review
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs 2026-05-30
LLMs Need Encoders for Semantic IDs Too 2026-05-29
The Regularizing Power of Language-Training Deepfake Detectors 2026-05-29
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs 2026-05-29
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages 2026-05-29
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding 2026-05-29 accept by iclm2026

Video Foundation Model

Title Date Comment
WALL-WM: Carving World Action Modeling at the Event Joints 2026-06-01
AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance 2026-05-31
GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video 2026-05-30
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models 2026-05-28
PlayClass: Automated Play Behaviour Classification in Poultry 2026-05-26
Accep...

Accepted at CV4Animals Workshop @ CVPR 2026

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation 2026-05-26
ICML ...

ICML 2026, Project Page: https://aka.ms/world-r1, Code: https://github.com/microsoft/World-R1

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation 2026-05-22
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models 2026-05-18
Accep...

Accepted to CVPR 2026 Workshops CV4Smalls

Latent Video Prediction Learns Better World Models 2026-05-15
DriveCtrl: Conditioned Sim-to-Real Driving Video Generation 2026-05-14
Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception 2026-05-12
v2: M...

v2: Major revision. Retitled; expanded from TimeSformer alone to four backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, CLIP), with V-JEPA 2.1 nearly tripling pretrained performance. Adds zero-shot PHASE transfer, attention-rollout analysis, and a language-distillation control. Data (OOO sim. judgments) & core hybrid triplet+RSA LoRA method unchanged from v1. Prepared for NeurIPS 2026 submission

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation 2026-05-09 17 pages, 9 figues
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios 2026-05-07
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models 2026-05-07
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation 2026-04-16
Accep...

Accepted at 2026 International Conference on Automatic Face and Gesture Recognition (FG)

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions