The following papers are my research interests (some papers I have recently read🥸).
- 2025.06.03 Video-XL-2
- 2025.06.02 ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- 2025.03.24 Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
- 2025.01.21 InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- 2024.12.31 Online Video Understanding: A Comprehensive Benchmark and Memory-Augmented Method
- 2024.09.27 From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
- 2024.09.22 Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
- 2024.09.02 TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
- 2024.05.31 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- 2024.03.22 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- 2023.06.08 Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- 2022.12.06 InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- 2024.04.08 MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
- 2024.04.04 LongVLM: Efficient Long Video Understanding via Large Language Models
- 2023.11.28 LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- 2023.07.31 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- 2025.04.22 From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
- 2025.03.15 Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
- 2025.01.30 SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
- 2025.01.16 Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
- 2025.01.12 A General Framework for Inference-time Scaling and Steering of Diffusion Models
- 2023.10.17 GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- 2025.03.31 A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- 2025.01.31 SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
- 2024.08.06 Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- 2024.07.31 Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- 2023.10.17 GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- 2023.07.12 T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation
- 2025.03.12 SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
- 2024.12.12 SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
- 2024.12.03 SNOOPI: Supercharged One-step Diffusion Distillation with Proper Guidance
- 2024.08 FLUX
- 2023.11.28 SDXL-Turbo github
- 2023.09.30 PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
- 2025.05.19 Mean Flows for One-step Generative Modeling
- 2025.05.26 ImgEdit: A Unified Image Editing Dataset and Benchmark
- 2025.05.22 KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
- 2025.04.29 In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- 2025.04.24 Step1X-Edit: A Practical Framework for General Image Editing
- 2025.02.24 KV-Edit: Training-Free Image Editing for Precise Background Preservation
- 2024.11.22 HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
- 2024.11.21 Stable Flow: Vital Layers for Training-Free Image Editing
- 2024.11.07 Taming Rectified Flow for Inversion and Editing
- 2022.08.02 Prompt-to-Prompt Image Editing with Cross Attention Control
- 2024.12.05 A Noise is Worth Diffusion Guidance
- 2024.07.19 Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
- 2024.06.06 ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
- 2023.12.13 Semantic-Driven Initial Image Construction for Guided Image Synthesis in Diffusion Model
- 2023.05.05 Guided Image Synthesis via Initial Image Editing in Diffusion Model
- 2023.04.27 Generating images of rare concepts using pre-trained diffusion models
- 2025.04.08 Transfer between Modalities with MetaQueries
- 2025.03.17 Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
- 2024.09.12 A Comprehensive Survey on Deep Multimodal Learning with Missing Modality
- 2021.02.26 Learning Transferable Visual Models From Natural Language Supervision (CLIP)
- 2023.03.15 GPT-4 Technical Report (GPT4)
- 2022.01.26 InstructGPT: Training Language Models to Follow Instructions with Human Feedback
- 2020.05.28 Language Models are Few-Shot Learners (GPT3)
- 2019.02.14 Language Models are Unsupervised Multitask Learners (GPT2)
- 2018.06.11 Improving Language Understanding by Generative Pre-Training (GPT)
- 2025.02.14 Large Language Diffusion Models
- 2024.10.24 Scaling up Masked Diffusion Models on Text
- 2024.10.23 Scaling Diffusion Language Models via Adaptation from Autoregressive Models (Adaptation)
- 2024.06.06 Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
- 2023.10.25 Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution (SEDD)
- 2023.03.12 Diffusion Models for Non-autoregressive Text Generation: A Survey (Survey)
- 2022.11.30 Score-based Continuous-time Discrete Diffusion Models (Score-based)
- 2022.11.02 Concrete Score Matching: Generalized Score Matching for Discrete Data
- The first to propose concrete score for modeling discrete data.
- 2022.05.30 A Continuous Time Framework for Discrete Denoising Models (CTMC)
- First complete continuous time framework for discrete diffusion.
- 2021.07.07 Structured Denoising Diffusion Models in Discrete State-Spaces (D3PM)
- Diffusion models with discrete state spaces as a competitive model class for large scale text or image generation, however, it trains and samples the model in discrete time.
- 2022.07.26 Classifier-Free Diffusion Guidance (classifier free guidance)
- 2021.05.11 Diffusion Models Beat GANs on Image Synthesis (classifier guidance)
- 2020.11.26 Score-Based Generative Modeling through Stochastic Differential Equations
- Unified framework for 19's score-based method and 20's DDPM.
- 2020.10.06 Denoising Diffusion Implicit Models (DDIM)
- 2020.06.19 Denoising Diffusion Probabilistic Models (DDPM)
- 2019.07.12 Generative Modeling by Estimating Gradients of the Data Distribution (Score-based method)
- 2015.03.12 Deep Unsupervised Learning using Nonequilibrium Thermodynamics
- 2024.07.08 Ada-adapter:Fast Few-shot Style Personlization of Diffusion Model with Pre-trained Image Encoder (Ada-adapter)
- 2023.08.13 IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models (IP-Adapter)
- 2023.02.10 Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)
- 2022.08.25 DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation (DreamBooth)
- 2022.08.02 An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion (Textual Inversion)
-
2023.12.06 On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm (RDED)
-
2023.06.22 Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective (SRe2L)
-
2022.07.20 DC-BENCH: Dataset Condensation Benchmark
-
2022.06.29 Beyond neural scaling laws: beating power law scaling via data pruning (scaling law)
-
2022.03.22 Dataset Distillation by Matching Training Trajectories (TM: matching trajectory)
-
2022.03.03 CAFE: Learning to Condense Dataset by Aligning Features (matching feature)
-
2021.10.08 Dataset Condensation with Distribution Matching (DM: matching distribution)
-
2021.06.09 Knowledge distillation: A good teacher is patient and consistent
-
2020.06.10 Dataset Condensation with Gradient Matching (DC: matching gradient)
-
2015.03.09 Distilling the Knowledge in a Neural Network
- 2020.06.13 Bootstrap your own latent: A new approach to self-supervised Learning (BYOL)
Transformer/...
AlexNet/GAN...
- Visual Generation zhihu
- Language Modeling by Estimating the Ratios of the Data Distribution (SEDD)