[ICLR 2026] Official Code for the Paper "DAVE: A VLM VISION ENCODER FOR DOCUMENT UNDERSTANDING AND WEB AGENTS"
DAVE is a vision encoder built specifically for vision–language models to better handle document understanding and web agent tasks, addressing the weak structural/spatial features of standard vision encoders. It trains mostly on unlabeled data via self-supervised pretraining, then uses a small amount of high-quality supervised autoregressive data for parsing and localization.
- Coming soon.
