Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

1 Commit
 
 
 
 

Repository files navigation

[ICLR 2026] Official Code for the Paper "DAVE: A VLM VISION ENCODER FOR DOCUMENT UNDERSTANDING AND WEB AGENTS"

Description

DAVE is a vision encoder built specifically for vision–language models to better handle document understanding and web agent tasks, addressing the weak structural/spatial features of standard vision encoders. It trains mostly on unlabeled data via self-supervised pretraining, then uses a small amount of high-quality supervised autoregressive data for parsing and localization.

DAVE teaser figure

Model Weights

  • Coming soon.

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors