Skip to content

Latest commit

Β 

History

171 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ”­ Daily Papers (2026-08-21)

Document Parsing

Venue Name Primary affiliation Title GitHub Date
Paper ArmorOCR Ant Group ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation GitHub Stars Aug. 2026

ArmorOCR β€” A two-stage robust adversarial OCR framework (on-policy self-distillation + GRPO) that treats adversarial OCR as grounded perception, backed by AdvSpot, the first benchmark for grounded adversarial OCR evaluation with 390 images across 13 fine-grained adversarial text types.

Document Understanding

Venue Name Primary affiliation Title GitHub Date
Paper Q-Guide Amazon Question-Guided Evidence Acquisition for Multimodal Visual Question Answering - Aug. 2026

Q-Guide β€” A small evidence-acquisition agent for document VQA that reads the question, works out what evidence is still missing, and calls targeted tools (read text / zoom in / ground a region) to recover it, lifting DocVQA2026 from 40.0% to 65.0% over direct prompting.

Visual Text Generation

Venue Name Primary affiliation Title GitHub Date
Paper TextRefine Kuaishou TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters - Aug. 2026

TextRefine β€” A task-aligned post-training framework for text editing in product posters with operation-specific rewards (text-span-level for insertion, CTC glyph-level for replacement), plus OpenTextEdit, a 100K-image dataset for poster text editing.

πŸ“– Contents

Overview

A curated, continuously updated reading list of OCR in the era of large language models, covering document parsing and understanding, visual text generation, benchmarks, challenges, and future perspectives, with a focus on research around the past five years (2021–now).

Scope. This list tracks OCR in the LLM era: work that applies large vision-language or multimodal models to text-rich images and documents (parsing, understanding, benchmarks, and specialized text tasks). It is not a general document-AI list, a generic MLLM list, or a classical OCR-1.0 list; such work appears only when it directly bears on text-rich visual understanding.

A note on evaluation. Most recent systems are released as technical reports with self-reported numbers, private test sets, and inconsistent protocols, so cross-paper scores are rarely comparable in a rigorous sense. We list results as reported and, where known, indicate the evaluation basis. The field still lacks a unified, contamination-resistant, reproducible benchmark, and we see building one as a prerequisite for trustworthy leaderboard claims.

πŸŽ‰ News

  • [2026-2-11] πŸ”₯ We release an open-source resource to help the community easily track recent OCR research!

Contributing. PRs welcome. One row per model, newest first; please include venue/date, affiliation, and a code or model link.

πŸ” Emerging Trends

Emerging Trends (2023–2026)

  • End-to-end VLM-based parsing replaces modular OCR pipelines.
  • Reinforcement learning for layout and reading order modeling.
  • OCR-free document understanding models.
  • Scaling down: compact document VLMs under 1B parameters.
  • Long-doc OCR is a scaling problem of tokens and consistency, not just context length.
  • Structure (layout + logic) is the new accuracy.
  • Generative OCR shifts the core risk from β€œmisrecognition” to β€œhallucination”.
  • Benchmarks are moving toward executable evaluation.
  • Document agents and autonomous reasoning over PDFs.
  • Document Agents need recoverability, not one-shot perfection.

πŸ“„ Document Parsing

Document parsing focuses on converting visually complex documents into structured, machine-readable representations. In the LLM era, parsing is no longer a pipeline of isolated modules, but increasingly unified within end-to-end VLM architectures.

Venue Name Primary affiliation Title GitHub Date
Paper ArmorOCR Ant Group ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation GitHub Stars Aug. 2026
Paper NaviDC-OCR China Telecom AI NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents - Aug. 2026
Paper TongGuOCR SCUT TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents GitHub Stars Aug. 2026
Paper PaDoc Tsinghua University PaDoc: Layout-Grounded Parallel Decoding for Document Parsing GitHub Stars Aug. 2026
Logographic Pretraining QMUL Logographic Character Visual Pretraining via Semantic-based Contrastive Learning - Aug. 2026
SPIRAL SEU Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression GitHub Stars Aug. 2026
DocPO Tencent DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards - Aug. 2026
Paper DrawAI BUPT DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable HuggingFace GitHub Stars Aug. 2026
Paper LayoutLite Yuanli Technology LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR GitHub Stars Jul. 2026
Paper HPD-Parsing Baidu HPD-Parsing: Hierarchical Parallel Document Parsing HuggingFace GitHub Stars Jul. 2026
Paper OvisOCR2 Alibaba OvisOCR2 Technical Report HuggingFace GitHub Stars Jul. 2026
Paper DocOCR-Eval University of Melbourne DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth - Jul. 2026
Paper MonkeyOCRv2 HUST&Kingsoft MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Hugging Face GitHub Stars Jul. 2026
Paper Infinity-Parser2 INF Team Infinity-Parser2 Technical Report Hugging FaceGitHub Stars Jul. 2026
Paper HunyuanOCR-1.5 Tencent HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better HuggingFace Stars GitHub Stars Jul. 2026
Paper SAYRE Alibaba Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis - Jul. 2026
Paper P-MTP Baidu P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling - Jun. 2026
Paper RT-DocLayout Baidu RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild - Jun. 2026
Paper Unlimited-OCR Baidu Unlimited OCR Works GitHub Stars Jun. 2026
Paper Beaver Microsoft Research Building Agent Harnesses for Scientific Curation from Multimodal Sources - Jun. 2026
Paper Agents-K1 Shanghai AI Laboratory Agents-K1: Towards Agent-native Knowledge Orchestration HuggingFace Stars Jun. 2026
Paper PaddleOCR-VL-1.6 Baidu PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training GitHub Stars Jun. 2026
Paper PP-OCRv6 Baidu PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks GitHub Stars Jun. 2026
StrucTab IIE, CAS StrucTab: A Structured Optimization Framework for Table Parsing GitHub Stars Jun. 2026
Paper ExChart Zhejiang University Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework - Jun. 2026
Paper MinerU-Popo Shanghai AI Laboratory & OpenDataLab MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing GitHub Stars May. 2026
RTPrune DeepSeek RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference GitHub Stars May. 2026
Paper ABot-OCR Alibaba ABot-OCR Technical Report GitHub Stars May. 2026
Paper BabelDOC funstory.ai BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation GitHub Stars May. 2026
Paper Consensus Entropy Fudan University Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR GitHub Stars May. 2026
Paper FastOCR Tsinghua University & JD FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing - May. 2026
Paper MinerU2.5-Pro Shanghai AI Laboratory MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale GitHub Stars Apr. 2026
Paper PixelPrune OPPO PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding GitHub Stars Apr. 2026
Paper TexOCR Yale University & Zhejiang University TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction GitHub Stars Apr. 2026
Paper Falcon OCR Falcon Vision Team, TII Falcon Perception GitHub Stars Mar. 2026
Paper MinerU-Diffusion Shanghai AI Laboratory MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding GitHub Stars Mar. 2026
Paper Qianfan-OCR Baidu Qianfan-OCR: A Unified End-to-End Model for Document Intelligence GitHub Stars Mar. 2026
Paper dots.mocr HUST Multimodal OCR: Parse Anything from Documents GitHub Stars Mar. 2026
Paper FireRed-OCR Xiaohongshu Inc FireRed-OCR Technical Report GitHub Stars Mar. 2026
Paper AgenticOCR Shanghai AI Laboratory & OpenDataLab AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation GitHub Stars Mar. 2026
PTP Tencent Efficient Document Parsing via Parallel Token Prediction - Mar. 2026
Paper Logics-Parsing-Omni Alibaba Logics-Parsing-Omni Technical Report GitHub Stars Mar. 2026
Paper Agentar-Fin-OCR Ant Group Agentar-Fin-OCR - Mar. 2026
Paper PaddleOCR-VL Baidu Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing GitHub Stars Mar. 2026
Paper DODO Amazon & Technion DODO: Discrete OCR Diffusion Models - Feb. 2026
Paper HSD SCUT Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding - Feb. 2026
Paper MeDocVL Ping An Property MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing GitHub Stars Feb. 2026
Paper Dolphin-2.0 ByteDance Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting GitHub Stars Feb. 2026
Paper GLM-OCR Z.ai GLM-OCR Technical Report GitHub Stars Feb. 2026
Paper OCR-Agent Southwest Minzu University OCR-Agent: Agentic OCR with Capability and Memory Reflection GitHub Stars Feb. 2026

πŸ“„ See full list at Document-Parsing.md

πŸ“„ Document Understanding

Document understanding extends beyond structural parsing to semantic comprehension and reasoning over visually rich documents.

Venue Name Primary affiliation Title GitHub Date
Paper Q-Guide Amazon Question-Guided Evidence Acquisition for Multimodal Visual Question Answering - Aug. 2026
Paper DocClaw NTU DocClaw: A Unified Agentic System for Intelligent Document Processing GitHub Stars Aug. 2026
Paper ConceptFormer Northeastern University ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval GitHub Stars Aug. 2026
Hyper-M2RAG Hangzhou Dianzi University Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement GitHub Stars Aug. 2026
Paper Trident Emory University What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering - Aug. 2026
Paper D2-ScaleAgent Zhejiang University D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding - Aug. 2026
SEER UT Austin SEER: Long-Context Reasoning via Selective Visual-Text Compression GitHub Stars Aug. 2026
Paper HAM-RAG HKUST HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation GitHub Stars Aug. 2026
DRUF Shenzhen MSU-BIT University Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs GitHub Stars Aug. 2026
Paper DistilVDR Aalto University DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation GitHub Stars Aug. 2026
Paper InSight-doc HKUST InSight-doc: Agentic Visual Perception for Long-Document Understanding GitHub Stars Aug. 2026
Paper DocAtlas Wuhan University / Microsoft DocAtlas: Long-Document Understanding as Mutable-State Interaction - Aug. 2026
Paper DocMemo HIT, Shenzhen DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding GitHub Stars Aug. 2026
Paper ECF Beihang University Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? GitHub Stars Aug. 2026
Paper ADOPD 2026 Georgia Tech Thinking with Anchors: Grounded and Efficient Document Reasoning HuggingFace GitHub Stars Aug. 2026
Paper VTS MBZUAI When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning - Aug. 2026
Paper Q-CueGraph The University of Tokyo Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning - Aug. 2026
Paper DocTrace Baidu DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning - Aug. 2026
Paper CURV William & Mary CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning - Aug. 2026
Paper RAGOCR Peking University RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation - Aug. 2026
Paper VaRS-Doc SJTU VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval GitHub Stars Aug. 2026
Paper ET-Prune SJTU ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs GitHub Stars Aug. 2026
Paper HierDoc USTC HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering - Aug. 2026
DualG-MRAG BUAA DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation - Jul. 2026
Paper MMLDSum-LLM OPPO MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware - Jul. 2026
Paper TAP-RAG Tianjin University TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering Anonymous Code Jul. 2026
Perception-RFT Quantiphi Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment - Jul. 2026
CMDR NTT CMDR: Contextual Multimodal Document Retrieval HuggingFace Stars GitHub Stars Jul. 2026
Paper HiEvi-RAG USTC Hierarchical Evidence-Driven Reasoning for Long Document Understanding - Jul. 2026
Paper MultAttnAttrib β€” MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering - Jul. 2026
Paper OracleAnalyser NUDT OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training - Jun. 2026
Paper DocArena Adobe Research DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents - Jun. 2026
ViTexQA Meituan ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering GitHub Stars HuggingFace Dataset Jun. 2026
Paper PreciseDoc Tsinghua University An LMM for Precisely Grounding Elements in Documents - Jun. 2026
LightSTAR SJTU LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement GitHub Stars Jun. 2026
Paper SciLens HKUST SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding - Jun. 2026
Paper SAFE-Cascade Walmart SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering - Jun. 2026
Paper UMG-RAG Purdue University Uncertainty-Aware Hybrid Retrieval for Long-Document RAG - Jun. 2026
Paper MAGE-RAG BIT MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA GitHub Stars Jun. 2026
Paper MINARD University of Maryland Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures - Jun. 2026
Paper MM-BizRAG JPMorgan Chase & Co. MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A - Jun. 2026
Paper KG4VD National Taiwan University Multimodal Graph RAG for Long-range Visually Rich Document Understanding GitHub Stars Jun. 2026
Paper GeoSym127K SenseTime Research & CUHK Shenzhen GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning GitHub Stars May. 2026
Paper FinAgent-RAG Zhejiang University Agentic Retrieval-Augmented Generation for Financial Document Question Answering GitHub Stars May. 2026
Paper SMART UW-Madison Your Embedding Model is SMARTer Than You Think GitHub Stars May. 2026
Paper ReceiptBench Zhejiang University From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding GitHub Stars May. 2026
HiKEY Korea University HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering - May. 2026
Paper UniDoc-RL DeepGlint-AI & SJTU UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards GitHub Stars Apr. 2026
Doc-V* HUST Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA - Apr. 2026
DocSeeker HUST DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding GitHub Stars Apr. 2026

πŸ“„ See full list at Document-Understanding.md

πŸ“„ Visual Text Generation

Visual Text Generation focuses on generating or editing legible, visually harmonious, and semantically consistent text within images, serving as the creative inverse of OCR. In the LLM era, it is no longer a task reliant on specialized modules, but is emerging as a foundational skill for general-purpose generative models.

Venue Name Primary affiliation Title GitHub Date
Paper TextRefine Kuaishou TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters - Aug. 2026
Paper PosterText Wuhan University PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster - Aug. 2026
Paper TransAnyText Wuhan University TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation - Aug. 2026
onoff GIST Bridging Online and Offline Handwriting via Differentiable Physical Rendering GitHub Stars Aug. 2026
Paper PosterMELD Tsinghua University PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs GitHub Stars Aug. 2026
InnoText SYSU InnoText: A Unified Model for Visual Text Generation and Editing - Jul. 2026
Paper Boogu-Image-0.1 Boogu Team,Huawei Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation HuggingFace GitHub Stars Jul. 2026
Paper SciForma Microsoft & Peking University SciForma: Structure-Faithful Generation of Scientific Diagrams GitHub Stars Jul. 2026
Paper VecFontLLM Fuzhou University & Peking University VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts - Jul. 2026
Paper ArtChart Ant Group ArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text Rendering - Jul. 2026
Paper DataEvolver CSU DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation GitHub Stars Jul. 2026
Paper Qwen-Image-2.0-RL Alibaba Qwen-Image-2.0-RL Technical Report - Jun. 2026
UniTranslator IIE CAS UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation GitHub Stars Jun. 2026
Paper SteerVTE ByteDance & Peking University SteerVTE: Seamless Video Text Editing with Style and Glyph Control - Jun. 2026
Paper DiffMath SCUT & Huawei DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation GitHub Stars Jun. 2026
Paper PhyDrawGen University of Dhaka PhyDrawGen: Physically Grounded Diagram Generation from Natural Language - Jun. 2026
Paper NIV Reichman University NIV: Neural Axis Variations for Variable Font Generation GitHub Stars Jun. 2026
Paper Qwen-Image-2.0 Alibaba Group Qwen-Image-2.0 Technical Report GitHub Stars May. 2026
Paper MangaFlow The University of Tokyo MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation - May. 2026
Paper Wan-Image Alibaba Group Wan-Image: Pushing the Boundaries of Generative Visual Intelligence - Apr. 2026
PosterIQ PolyU PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation GitHub Stars Mar. 2026
Paper EfficientPosterGen Tsinghua University EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection GitHub Stars Mar. 2026
Paper TextFlow NJUST Towards Training-Free Scene Text Editing GitHub Stars Mar. 2026
Paper GlyphPrinter Fudan University GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering GitHub Stars Mar. 2026
Paper CTRL-S SJTU Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning - Mar. 2026
Paper LaDe Adobe Research LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition - Mar. 2026
Paper EchoGen USTC EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding - Mar. 2026
Paper WebVR StepFun WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics - Mar. 2026
Paper AutoFigure-Edit Westlake University AutoFigure-Edit: Generating Editable Scientific Illustration GitHub Stars Mar. 2026
Paper GlyphBanana SJTU & Xiaohongshu Inc. GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows GitHub Stars Mar. 2026
Paper InnoAds-Composer JD.com InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation - Mar. 2026
Paper Seeing is Improving USTC Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement GitHub Stars Mar. 2026
PosterOmni HKUST(GZ) PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback GitHub Stars Feb. 2026
TextPecker HUST TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering GitHub Stars Feb. 2026
Paper SSPT Tongji University Space Syntax-guided Post-training for Residential Floor Plan Generation - Feb. 2026
Paper ChatUMM Tsinghua University & Tencent Hunyuan ChatUMM: Robust Context Tracking for Conversational Interleaved Generation - Feb. 2026
PosterVerse SCUT PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography GitHub Stars Jan. 2026
- GPT-Image-1.5 OpenAI GPT-Image-1.5 --- Dec. 2025
- Gemini 3 Pro Image Google DeepMind Gemini 3 Pro Image (Nano Banana Pro) --- Nov. 2025
- Gemini 2.5 Flash Image Google DeepMind Gemini 2.5 Flash Image (Nano Banana) --- Oct. 2025
Paper Qwen-Image Qwen Team Qwen-Image Technical Report GitHub Stars Sep. 2025
Paper Seedream 4.0 ByteDance Seed Seedream 4.0: Toward Next-generation Multimodal Image Generation --- Aug. 2025
Paper Postergen Stony Brook University Postergen: Aesthetic-aware paper-to-poster generation via multi-agent llms GitHub Stars Aug. 2025
UniGlyph Tsinghua University UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis --- Jul. 2025
Paper X-Omni Tencent Hunyuan X X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again GitHub Stars Jul. 2025
Paper DreamPoster Intelligent Creation Lab, ByteDance DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design --- Jul. 2025
PosterCraft HKUST(GZ) PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework GitHub Stars Jun. 2025
Paper Bagel ByteDance Emerging Properties in Unified Multimodal Pretraining GitHub Stars May. 2025
Paper FLUX-Text Amap, Alibaba Group FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing GitHub Stars May. 2025
Paper TextFlux bilibili Inc. TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis GitHub Stars May. 2025

πŸ“„ See full list at Visual-Text-Generation.md

πŸ“„ Specialized Model

Beyond general document parsing and understanding, traditional OCR research remains essential for specialized visual-text structures, restoration, temporal analysis, and forensic reliability.

πŸ“„ Document Dewarping

Document dewarping restores photographed or scanned pages to a geometrically rectified and readable form.

Venue Name Primary affiliation Title GitHub Date
Paper BookNet HFUT BookNet: Dual-Page Book Image Rectification via Cross-Page Attention - Jan. 2026
Paper TADoc CAS TADoc: Robust Time-Aware Document Image Dewarping - Aug. 2025
Paper DocDewarpHV HIT Dual Dimensions Geometric Representation Learning Based Document Dewarping GitHub Stars Jul. 2025
DocMatcher FZI & KIT DocMatcher: Document Image Dewarping via Structural and Textual Line Matching GitHub Stars Mar. 2025
Paper - Shuya Branch of the Ivanovo State University Efficient Document Image Dewarping via Hybrid Deep Learning and Cubic Polynomial Geometry Restoration GitHub Stars Jan. 2025
DocScanner USTC DocScanner: Robust Document Image Rectification with Progressive Learning GitHub Stars Jan. 2025
DocRes SCUT DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks GitHub Stars Jun. 2024
DocNLC SCUT DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations GitHub Stars Feb. 2024
DocTr++ USTC Deep Unrestricted Document Image Rectification - 2024
LA-DocFlatten CAS Layout-aware Single-image Document Flattening GitHub Stars 2024
UVDoc ETH Zurich UVDoc: Neural Grid-based Document Unwarping GitHub Stars Oct. 2023
Foreground and Text-lines Aware Model HIT Shenzhen, China Foreground and text-lines aware document image rectification GitHub Stars Jun. 2023
DocMAE USTC & iFLYTEK DocMAE: Document Image Rectification via Self-supervised Representation Learning - 2023
Marior SCUT & IntSig Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild GitHub Stars Oct. 2022

πŸ“„ Physical Structure Analysis

Physical structure analysis identifies document regions, layouts, and spatial relationships before semantic reasoning.

Venue Name Primary affiliation Title GitHub Date
Paper PorTEXTO NOVA School of Science and Technology PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction - Jun. 2026
Paper - Insiders Technologies GmbH Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets - Jun. 2026
Paper IndustryBench-MIPU Multimodal and Industrial AI Team, Alibaba IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products GitHub Stars Jun. 2026
Paper ERN-Net Tamkang University ERN-Net : Evolving Reason Node-Net for Document Binarization - Jun. 2026
HiLEx IEM Kolkata HiLEx: Image-Based Hierarchical Layout Extraction from Question Papers GitHub Stars Mar. 2025
DCEM-ViT Sharda University, Greater Noida, India Devanagari character encoded mix-merge vision transformer for robust document layout analysis --- Mar. 2025
Efficient Additive Attention DLA Technical University of Kaiserslautern, Germany Efficient Additive Attention for Transformer-based Semi-supervised Document Layout Analysis --- Feb. 2025
DocSemi Technical University of Kaiserslautern, Germany DocSemi: Efficient Document Layout Analysis with Guided Queries --- Feb. 2025
FS-QCSNet China University of Mining and Technology Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation --- Jan. 2025
LayoutDETR Salesforce Research LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer - Sep. 2024
LayoutLLM Alibaba Group LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding GitHub Stars Jun. 2024
RoDLA KIT & Univ. of Oxford RoDLA: Benchmarking the Robustness of Document Layout Analysis Models GitHub Stars Jun. 2024
DocLLM JPMorgan Chase & Co. DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding GitHub Stars May 2024
SemiDocSeg Computer Vision Center, Barcelona, Spain SemiDocSeg: Harnessing Semi-Supervised Learning for Document Layout Analysis --- Mar. 2024
VGT Alibaba Group Vision Grid Transformer for Document Layout Analysis GitHub Stars Oct. 2023
GeoLayoutLM Alibaba Group GeoLayoutLM: Geometric Pre-training for Visual Information Extraction GitHub Stars Jun. 2023
M6Doc SCUT M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis GitHub Stars Jun. 2023
HRDoc USTC & iFLYTEK HRDoc: Dataset and Baseline Method Toward Hierarchical Reconstruction of Document Structures GitHub Stars Feb. 2023
LayoutLMv3 Microsoft LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking GitHub Stars Oct. 2022

πŸ“„ Reading Order Prediction

Reading order prediction recovers the intended sequence of text and visual elements in complex page layouts.

Venue Name Primary affiliation Title GitHub Date
Paper Orli University of Konstanz End-to-End Text Line Detection and Ordering GitHub Stars Jun. 2026
Paper FocalOrder Unisound AI Technology Co. Ltd. FocalOrder: Focal Preference Optimization for Reading Order Detection - Jan. 2026
Paper ROAP UESTC ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction GitHub Stars Jan. 2026
Paper XY-cut++ Tianjin University XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark - Apr. 2025
UniHDSA USTC & Microsoft Research Asia UniHDSA: A Unified Relation Prediction Approach for Hierarchical Document Structure Analysis GitHub Stars 2025
HanDoc-OrderOCR NCCU & Academia Sinica Reading between the Lines: Image-Based Order Detection in OCR for Chinese Historical Documents - Feb. 2024
Detect-Order-Construct USTC & Microsoft Research Asia Detect-Order-Construct: A Tree Construction Based Approach for Hierarchical Document Structure Analysis GitHub Stars 2024
Reading Order Matters Fudan University Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction GitHub Stars Dec. 2023
LayoutLMv3 SYSU LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking GitHub Stars Oct. 2022

πŸ“„ Mathematical Expression Recognition

Mathematical expression recognition converts printed or handwritten formulas into structured symbolic representations.

Venue Name Primary affiliation Title GitHub Date
Paper - University of Alicante Direct content-based retrieval from music scores images - May. 2026
Paper MusicSynth - MusicSynth: An Automated Pipeline for Generating Violin Fingerboard Animations from Sheet Music Using Optical Music Recognition - May. 2026
Paper Transcoda Heinrich Heine University DΓΌsseldorf Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training GitHub Stars May. 2026
Paper From Image to Music Language FindLab From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR - Apr. 2026
UniMERNet OpenDataLab & Shanghai AI Lab UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition GitHub Stars Jun. 2025
CDM OpenDataLab Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching GitHub Stars Jun. 2025
SSAN Inner Mongolia University SSAN: A Symbol Spatial-Aware Network for Handwritten Mathematical Expression Recognition GitHub Stars Apr. 2025
TAMER Peking University TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression Recognition GitHub Stars Apr. 2025
Journal VLPG UCAS Vision–language pre-training for graph-based handwritten mathematical expression recognition GitHub Stars Jan. 2025
BAT USTC & iFLYTEK Bidirectional Trained Tree-Structured Decoder for Handwritten Mathematical Expression Recognition GitHub Stars 2025
LattE Purdue University LattE: Improving LaTeX Recognition with Iterative Refinement GitHub Stars Sep. 2024
- TUAT & VGU & RIT & Nantes Univ. A Survey on Handwritten Mathematical Expression Recognition: The Rise of Encoder-Decoder and GNN Models - Sep. 2024
NAMER USTC & iFLYTEK Research NAMER: Non-Autoregressive Modeling for Handwritten Mathematical Expression Recognition - Sep. 2024
HMEG Hangzhou Dianzi Univ. Generating Handwritten Mathematical Expressions from Symbol Graphs: An End-to-End Pipeline - Jun. 2024
BPD SCUT A Tree-Based Model with Branch Parallel Decoding for Handwritten Mathematical Expression Recognition - 2024
SAN (Syntax-Aware Net) Tomorrow Advancing Life Syntax-Aware Network for Handwritten Mathematical Expression Recognition GitHub Stars Jun. 2022

πŸ“„ Table & Chart Understanding

Table and chart understanding covers structure recognition, data extraction, grounding, retrieval, and visual reasoning over structured graphics.

Venue Name Primary affiliation Title GitHub Date
Paper - Teamreboott Inc. Rethinking the Pointer Loss in Table Structure Recognition: Geometry-Aware Pointer Loss for Spatial Locality GitHub Stars Jun. 2026
Paper - Preferred Networks, Inc. Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition via Order-Independent Cell-Level Representations - Jun. 2026
Paper ChartLens Shandong University ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement GitHub Stars Jun. 2026
Paper POTATR Kensho Technologies (S&P Global) POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction - Jun. 2026
Journal ChartReLA University of Science, Ho Chi Minh city, Vietnam ChartReLA: A compact vision-language model for comprehensivechart reasoning via relationship modeling GitHub Stars Jan. 2026
Table-R1 XJTU-Liverpool University Can GRPO Boost Complex Multimodal Table Understanding? - Dec. 2025
TinyChart Alibaba & Tsinghua TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token Merging - Nov. 2024
ReachQA Fudan University Distill Visual Chart Reasoning Ability from LLMs to MLLMs GitHub Stars Oct. 2024
OmniParser Alibaba Group OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition GitHub Stars Jun. 2024
DePlot Google DePlot: One-shot visual language reasoning by plot-to-table translation GitHub Stars Jul. 2023
TableVLM Fudan University TableVLM: Multi-modal Pre-training for Table Structure Recognition - Jul. 2023
VAST Huawei Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling - Jun. 2023
LORE Alibaba Group LORE: Logical Location Regression Network for Table Structure Recognition GitHub Stars Feb. 2023
ChartQA York University ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning GitHub Stars Jul. 2022
LGPMA Hikvision LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment GitHub Stars Sep. 2021

πŸ“„ Scene Text Understanding

Scene text understanding unifies text detection, recognition, and spotting in natural images and related real-world settings.

Venue Name Primary affiliation Title GitHub Date
Paper Next-Generation Parallel Decoder for LPDR NUTECH Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation - Jun. 2026
Paper - Polytechnic University of Turin Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting - May. 2026
Paper MNSP Nankai University Masked Next-Scale Prediction for Self-supervised Scene Text Recognition GitHub Stars May. 2026
Paper - Afeka Academic College of Engineering Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing - Apr. 2026
DPText-DETR Wuhan Univ. DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer GitHub Stars Feb. 2023
DeepSolo Wuhan University DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting GitHub Stars Nov. 2022
ABINet++ USTC ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting GitHub Stars Nov. 2022
SwinTextSpotter SCUT SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition GitHub Stars Mar. 2022
ABCNet v2 SCUT ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting GitHub Stars May. 2021
PAN++ Nanjing University PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text GitHub Stars May. 2021
MANGO Hikvision Research Institute MANGO: A Mask Attention Guided One-Stage Scene Text Spotter GitHub Stars Dec. 2020
Mask TextSpotter v3 HUST Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting GitHub Stars Jul. 2020
Text perceptron Hikvision Research Institute Text perceptron: Towards end-to-end arbitrary-shaped text spotting GitHub Stars Apr. 2020
ABCNet SCUT ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network GitHub Stars Feb. 2020

πŸ“„ Text Removal & Editing

Text removal and editing modifies textual content in images while preserving visual consistency and surrounding appearance.

Venue Name Primary affiliation Title GitHub Date
Paper ProductConsistency Fractal Analytics ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL GitHub Stars Jun. 2026
Paper - Fudan University Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification - Jun. 2026
Paper UniDDT Nanjing University UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer - Jun. 2026
Paper TextWand Peking University TextWand: A Unified Framework for Scene Text Editing - Jun. 2026
TMIM USTC Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling GitHub Stars Sep. 2024
PowerPaint Tsinghua & Shanghai AI Lab PowerPaint: A Task is Worth One Word β€” Learning with Task Prompts for High-Quality Versatile Image Inpainting GitHub Stars Sep. 2024
MagicEraser HuaweiΒ  & SIAT, CAS MagicEraser: Erasing Any Objects via Semantics-Aware Control - Sep. 2024
TurboEdit Adobe TurboEdit: Instant Text-based Image Editing GitHub Stars Sep. 2024
SPDInv HKUST Source Prompt Disentangled Inversion for Boosting Image Editability GitHub Stars Sep. 2024
DARLING USTC Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing - Jun. 2024
SceneTextGen Meta & Rutgers Layout-Agnostic Scene Text Image Synthesis with Diffusion Models - Jun. 2024
- KAIST Prompt Augmentation for Self-supervised Text-guided Image Manipulation - Jun. 2024
ViTEraser SCUT ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining GitHub Stars Feb. 2024
FETNet Kyushu Univ. FETNet: Feature Erasing and Transferring Network for Scene Text Removal GitHub Stars 2023

πŸ“„ Text Image Super-Resolution

Text image super-resolution restores low-resolution text regions while preserving character identity and legibility.

Venue Name Primary affiliation Title GitHub Date
Paper FDF Anhui University Frequency Decoupled Framework for Screen Content Image Super-Resolution - Jun. 2026
Paper DTG-Restore Virginia Tech DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution - Jun. 2026
Paper Everything at Every Scale MIT Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution - May. 2026
Paper PRISM SJTU PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution GitHub Stars May. 2026
TextDiff BUPT TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution - 2025
PEAN Southeast Univ. PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-Resolution GitHub Stars Oct. 2024
DCDM IIT Roorkee DCDM: Diffusion-Conditioned-Diffusion Model for Scene Text Image Super-Resolution GitHub Stars Sep. 2024
DiffTSR BIT & SenseTime Diffusion-based Blind Text Image Super-Resolution GitHub Stars Jun. 2024
SGENet Fudan Univ. & Videt Tech. SGENet: Efficient Scene Text Image Super-Resolution with Semantic Guidance - Apr. 2024
STIRER Tongji Univ. & Fudan Univ. STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition GitHub Stars Oct. 2023
DocDiff BUPT DocDiff: Document Enhancement via Residual Diffusion Models GitHub Stars Oct. 2023
- HIT & NTU Learning Generative Structure Prior for Blind Text Image Super-Resolution GitHub Stars Jun. 2023
DPMN Southeast Univ. DPMN: Improving Scene Text Image Super-Resolution via Dual Prior Modulation Network GitHub Stars Feb. 2023
TPGSR Hong Kong PolyU Text Prior Guided Scene Text Image Super-Resolution GitHub Stars 2023

πŸ“„ Handwritten Document Analysis

Handwritten document analysis covers handwriting recognition, writer-related analysis, signatures, and document-level handwritten content.

Venue Name Primary affiliation Title GitHub Date
Paper - LTU Performance Gap Analysis between Latin and Arabic Scripts HTR - Jun. 2026
Paper Prototypical Signature Approach ETS / LIVIA Lab, Universite du Quebec A Prototypical Signature Approach for Writer-Independent Offline Signature Verification GitHub Stars Jun. 2026
Paper Stringalign - Stringalign: Moving beyond summary statistics with a transparent Unicode-aware tool for evaluating automatic transcription models - Jun. 2026
Paper - Offenburg UAS Intelligent Character Recognition of Handwritten Forms with Deep Neural Networks - Jun. 2026
MetaWriter Concordia Univ. & Mila MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning - Jun. 2025
- Univ. of Alicante On the Generalization of Handwritten Text Recognition Models - Jun. 2025
DetailSemNet NYCU, Taiwan DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration - Sep. 2024
TransOSV Xi'an Jiaotong Univ. TransOSV: Offline Signature Verification with Transformers - 2024

πŸ“„ Video Text Analysis

Video text analysis studies temporal text detection, recognition, spotting, tracking, and reasoning across video frames.

Venue Name Primary affiliation Title GitHub Date
Paper TraRA RIKEN TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance GitHub Stars Jun. 2026
Paper Beyond Detection Nankai University Beyond Detection: A Structure-Aware Framework for Scene Text Tracking - May. 2026
Paper VTAgent Wuhan University VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA - May. 2026
VimTS HUST VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-Domain Generalization GitHub Stars Apr. 2025
GoMatching Wuhan Univ. & NTU GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching GitHub Stars Dec. 2024
TransDeTR Zhejiang Univ. & BUPT TransDeTR: End-to-End Video Text Spotting with Transformer GitHub Stars 2024
BiRViT-1K CASIA Video Text Detection With Robust Feature Representation - 2024

πŸ“„ Historical Document Analysis

Historical document analysis addresses OCR, restoration, structure recovery, and interpretation for manuscripts and archival materials.

Venue Name Primary affiliation Title GitHub Date
AlphaOracle HUST AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning GitHub Stars Jul. 2026
Paper Urdu Katib Handwritten Dataset University of Gujrat, Pakistan Urdu Katib Handwritten Dataset: A Historical Document Dataset for Offline Urdu Handwritten Text Recognition with CRNN-Based Baseline Evaluation - Jun. 2026
Paper - UMRE, Italy A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts - Jun. 2026
Paper - Charles University Optical Music Recognition for Real-World Manuscripts with Synthetic Data - Jun. 2026
Paper - LIGM Leveraging Morphology for Historical Script Metrological Analysis - Jun. 2026
KaiRacters TU Wien KaiRacters: Character-Level-Based Writer Retrieval for Greek Papyri - Dec. 2024
- Univ. Modena e Reggio Emilia Binarizing Documents by Leveraging both Space and Frequency - Aug. 2024
CATMuS Medieval PSL Univ. & ENC CATMuS Medieval: A Multilingual Large-Scale Cross-Century Dataset in Latin Script for Handwritten Text Recognition - Aug. 2024
DELINE8K Brigham Young Univ. DELINE8K: A Synthetic Data Pipeline for Semantic Segmentation of Historical Documents GitHub Stars Aug. 2024
- Univ. Rennes, IRISA Training Transformer Architectures on Few Annotated Data: Application to Historical Handwritten Text Recognition - 2024
ColDBin DFKI ColDBin: Cold Diffusion for Document Image Binarization - Aug. 2023
- Univ. of Sousse Historical Document Image Segmentation Combining Deep Learning and Gabor Features - Aug. 2023
- TU Wien Feature Mixing for Writer Retrieval and Identification on Papyri Fragments GitHub Stars Aug. 2023

πŸ“„ Tampered Text Detection & Forensics

Tampered text detection and forensics studies document authenticity, manipulation localization, forgery detection, and evidence-grounded verification.

Venue Name Primary affiliation Title GitHub Date
Paper - Sichuan University A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 - Jun. 2026
Paper SynCred-Bench Tsinghua University SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation - Jun. 2026
Paper TextFake USTC TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images - Jun. 2026
Paper DocQT MAIF / La Rochelle Universite DocQT: Improving Document Forgery Localization Robustness via Diverse JPEG Quantization Tables GitHub Stars May. 2026
OSTF SCUT Revisiting Tampered Scene Text Detection in the Era of Generative AI GitHub Stars Feb. 2025
SAFIRE KAIST & NAVER WEBTOON AI SAFIRE: Segment Any Forged Image Region - Feb. 2025
FFDN Xiamen Univ. & Tencent YouTu Enhancing Tampered Text Detection Through Frequency Feature Fusion and Decomposition - Sep. 2024
AdaIFL Peking Univ. AdaIFL: Adaptive Image Forgery Localization via a Dynamic and Importance-aware Transformer Network GitHub Stars Sep. 2024
EditGuard Peking Univ. EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection GitHub Stars Jun. 2024
- SCUT & HUST Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel Methods GitHub Stars Jun. 2024
- HUST Toward Real Text Manipulation Detection: New Dataset and New Solution - 2024
- Sichuan Univ. Pre-Training-Free Image Manipulation Localization through Non-Mutually Exclusive Contrastive Learning - Oct. 2023
DTD SCUT Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution GitHub Stars Jun. 2023
TruFor Univ. Federico II & Google TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization GitHub Stars Jun. 2023
DGM4 HIT Shenzhen & NTU Detecting and Grounding Multi-Modal Media Manipulation GitHub Stars Jun. 2023

πŸ“„ See full list at Specialized-Model.md

πŸ“„ Benchmarks and Evaluation

Benchmarks play a critical role in shaping the evolution of OCR in the LLM era.

Venue Benchmark Name Description Link Date
BudgetDoc The first multimodal benchmark with explicit supervision for model-budget-performance trade-offs on document tasks, plus DRB, a 1B-parameter pre-flight estimator that matches or improves F1 in 9 of 15 configurations while drastically cutting reasoning cost. - Aug. 2026
OmniHandwritingOCR A diagnostic benchmark for handwritten OCR spanning HTR and HMER across six subtasks and twelve subsets (77.57K images), with difficulty-stratified multi-line formulas exposing sharp degradation and hallucinated corrections in current MLLMs. - Aug. 2026
CADP Southeast University Code as Representation: A Compilable Parsing Paradigm for Academic Documents GitHub Stars
Paper BEAR-Bench Yandex BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models -
Paper LongDocBench A benchmark for two document-level structure recovery tasks β€” Table-of-Contents Hierarchy Recovery and Contextual Relationship Recovery β€” built on 85 real-world financial reports, textbooks and academic papers spanning 2,582 pages with human-verified annotations for 3,937 heading nodes and 3,258 contextual relationships, showing that representative document parsers remain limited on both recovery tasks despite strong page-level performance. - Aug. 2026
JieZi Formalizes Ancient Chinese Character Exegesis (ACCE) as a four-level vision-language QA task (identification, glyph-form analysis, meaning exegesis, diachronic evolution), with JieZi-Dataset (500K expert-audited QA pairs) and scholar-curated JieZi-Bench. GitHub Stars Aug. 2026
Paper FormStruct-Bench A hierarchical and diagnostic benchmark that evaluates table-form document structure recognition at both the document level and progressively finer component levels, allowing aggregate performance to be traced back to specific structural failure modes. - Aug. 2026
Paper ChartAnno A benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. - Aug. 2026
Paper BanglaWild The first in-the-wild Bengali scene text recognition benchmark to pair dual verbatim/standard annotation with evaluation of both conventional OCR and generative VLMs on the same data. - Aug. 2026
Paper ConfBench The first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set. HuggingFace Aug. 2026
Paper VTC-Eval Reveals that existing VTC evaluation protocols relying on downstream task performance fail to measure text preservation fidelity due to strong MLLM linguistic priors, and introduces a decoupled evaluation framework that isolates semantic retention from visual encoding quality. - Aug. 2026
Paper Evidence-Risk Audit Shows that correct answers can survive even when no retained token is traceable to the supporting OCR region, and proposes an evidence-risk audit coupling answer behavior with geometric token-origin provenance to catch this blind spot in token pruning evaluation. GitHub Stars Aug. 2026
Paper LongChart Bench A new pipeline and benchmark designed to evaluate MLLMs’ visual reasoning and performance in multi-chart settings with complex computational relationships. - Aug. 2026
Paper XL-DocBench A fully human-verified benchmark for extra-long document understanding, with 1,519 retained questions from six professional domains and contexts up to 2,303 pages. - Aug. 2026
Paper ExtractBench The first benchmark for schema-guided enterprise document extraction that jointly scores value accuracy, record completeness at scale, grounding, and measured cost. HuggingFace GitHub Stars Aug. 2026
Paper FaithC4 A controlled multilingual perturbation benchmark for measuring transcription faithfulness in VLMs, and use it to evaluate 15 systems across English, Chinese, and Korean. - Jul. 2026
GDP.pdf A benchmark for multimodal reasoning grounded in professional PDF documents, built from expert-authored tasks with the original files and their domain semantics intact. HuggingFaceGitHub Stars Jul. 2026
Paper SynthDocBench A long-context visual document understanding benchmark, measuring VLM performance for the ability to locate multiple key information facts over long-contexts from a given document to reason and answer a multi-step question. - Jul. 2026
Paper WILDTRACE A source-internal long-context multi-hop reasoning benchmark built from natural evidence trails.481 tasks over 214 naturally occurring long-form sources (incident reports, literary narratives) with 7 source-internal evidence geometries Hugging Face Jul. 2026
MORE A comprehensive multilingual document parsing benchmark covering diverse scripts and layouts. GitHub Stars Jul. 2026
Paper ClinOCR-Bench The first public clinical OCR benchmark with 384 scanned medical documents across 6 artifact conditions (normal, handwriting, poor quality, rotation, tables, mixed). GitHub Stars Jul. 2026
HCSU A fine-grained benchmark for evaluating LVLMs on historical calligraphy style perception, addressing a critical gap in cultural heritage AI with hierarchical style annotations across multiple calligraphy traditions. HuggingFace Stars Jul. 2026
Paper PosterHarness An auditable harness that evaluates text-rich image models on scientific poster generation by checking legibility, aspect ratios, and hallucinated scientific figures. - Jul. 2026
β€” DocuBench A schema-guided structured extraction benchmark of 50 hard real-world documents spanning 10 file types and 11 languages (including RTL and CJK scripts), with hand-verified JSON labels, an open scorer, and committed per-document baseline outputs from six systems. GitHub Stars Jul. 2026
Paper S-OBI The first benchmark for sentence-level oracle bone inscription understanding with 695 QA pairs; evaluates MLLMs beyond isolated character recognition. GitHub Stars Jul. 2026
Paper OCR-Robust First systematic benchmark for evaluating VLM OCR reasoning robustness under visual perturbations. GitHub Stars Jun. 2026
Invoice Haystack A benchmark for visually homogeneous document retrieval targeting invoice embedding collapse analysis. GitHub Stars Jun. 2026
Paper AGORA A benchmark for agentic reasoning over large-scale workplace document archives. - Jun. 2026
Paper PuMVR A parallel-script benchmark isolating orthography to evaluate true multilingual capability in Vision-Language Models. GitHub Stars Jun. 2026
Paper TableVision A large-scale benchmark for spatially grounded reasoning over complex hierarchical tables, coupling multi-step table reasoning with pixel-accurate grounding trajectories. - Apr. 2026

πŸ“„ See full list at Benchmarks-and-Evaluation.md

Contributing

We welcome contributions from the community and encourage pull requests to help keep this project up to date. Suggestions, feedback, and corrections are always welcome β€” please feel free to share them anytime!

About

OCR in the Era of Large Language Models

Resources

Stars

703 stars

Watchers

26 watching

Forks

Releases

Packages

Contributors