Replies: 2 comments
|
Grounding extracted values in the converted document's provenance is a good fit for docling, since conversion already records where every item sits on the page. A few things in the plan would bite in practice, though. #4201 overlaps more than the plan assumes. It rewrites the three files this touches ( Matching values against item text will produce wrong citations.
The citation would point at the whole item. "Confidence" would mean something else here. 1.0 for an exact match and 0.7 for a fuzzy one is a string-match score, and taking the min with A narrower first step could be a pure function over an |
|
@mvanhorn @wittjeff I really like the direction of this proposal! Adding native grounding layout metadata directly on However, as @wittjeff noted, we need to ensure this doesn't collide with the incoming
I'd be glad to help sketch out a clean string-matching lookup utility for |
Uh oh!
There was an error while loading. Please reload this page.
feat: grounded extraction citations and field confidence
Motivation
DocumentExtractor(shipped in PR #2138) returns per-page JSON from a VLM.ExtractedPageDataispage_no,extracted_data,raw_text,errorsonly (docling/datamodel/extraction.py:14-26). Conversion already stores layout provenance on everyDocItem(ProvenanceItem: page_no, bbox, charspan) and a document-levelConfidenceReport(layout/ocr/parse;table_scoreunimplemented). Extraction does not join either.LlamaExtract exposes
cite_sourcesandconfidence_scoreson extract: per-field citations (page/text, bbox except Turbo) plusparsing_confidence/extraction_confidence/confidence. RAG threads on r/LangChain ask how to keep provenance without blowing up cost.This plan adds the same two flags on Docling extract, filled by convert-then-ground: match extracted scalars to
DocItemtext and copyProvenanceItem. Do not wait on the VLM to emit citations. Do not wait on PR #4201: that PR is engines, formats, and payload channels, not field citations.This feature was proposed from code analysis without any hands-on usage.
Research Summary
General Project Dogfooding (Phase 0.5 Step 1e / Phase 1b.5)
890dd42.Feature-Specific Validation (Phase 2d)
ExtractedPageDatajson.loads,DocumentExtractor.extractsignature, LlamaExtract docsExtractionVlmPipeline._extract_data(extraction_vlm_pipeline.py:89-98) parses VLM text and constructsExtractedPageDatawith no grounding.extract()(document_extractor.py:129-149) has nocite_sources/confidence_scores. Live VLM tests intests/test_extraction.pyskip in CI (IS_CI). Grounding tests must not need a VLM.Project Context
docling/.agents/skills/docling/references/extraction.md)Related Issues & Community Signals
ghsearch for extraction citation / cite_sources / field grounding returned no issues or PRs.Social Search Results
Consumed from parent host_handoff. last30days was not run on this host. Path B was not taken.
cite_sources/confidence_scoresflagsextract_metadata.field_metadatacitationsCompetitive Analysis
LlamaExtract:
cite_sources: trueattaches per-field citations (page/text; bbox on non-Turbo).confidence_scores: trueattachesparsing_confidence,extraction_confidence,confidence.extraction_targetper_doc / per_page / per_table_row is a different gap (adjacent to #4201; SPECULATIVE here, not this PR).Docling conversion already has the layout data LlamaExtract has to re-derive from a parse. Extract ignores it.
Maintainer Analysis
prov, not a parallel citation type.What Gets Merged vs Rejected
feat(extraction):, tests that run without GPU, docs/examples, DCO.Proposed Solution
Add optional grounding on
ExtractedPageData:On
ExtractedPageData:field_citations: Optional[Dict[str, list[FieldCitation]]] = Nonefield_confidence: Optional[Dict[str, float]] = NoneFlags (default False, LlamaExtract names):
DocumentExtractor.extract(..., cite_sources: bool = False, confidence_scores: bool = False, converted_document: Optional[DoclingDocument] = None)extract_all.VlmExtractionPipelineOptionsascite_sources/confidence_scoresso anExtractionFormatOptioncan pin them. Kwargs override options.Grounding is a post-step after the VLM pipeline returns pages.
extraction_vlm_pipeline.pystaysjson.loadsonly. Do not prompt the VLM for citations.Algorithm in
docling/utils/extraction_grounding.py:DoclingDocument. Ifconverted_documentis passed, use it. Else ifcite_sourcesorconfidence_scores, runDocumentConverter().convert(source).documentonce per input (same path/stream the extractor already opened). If convert fails, leave citations empty and continue extract success; do not fail the VLM result.extracted_dataas a flat dict of scalars (str/int/float/bool). Nested dicts: join keys with.. Lists of scalars: one citation list per key, match each element. Skip dicts/lists of dicts in v1 (no per_table_row).DocItemtext (reading order). Match order: exact, then casefold+collapsed whitespace. First hit wins. Copyitem.prov[0]intoFieldCitation(page_no, bbox, charspan) plusself_ref.1.0exact,0.7fuzzy, omit key if no match. If aConversionResult.confidencepage score exists forpage_no, optionallymin(match_score, page.parse_score)whenconfidence_scoresis on. Do not invent a VLM logit.page_nowhenExtractedPageData.page_nois set.Not in this PR: VLM-emitted citations, per_table_row, whole-document merge (#4201), forms (#4055), word_cells (#3457).
Approach
docling/datamodel/extraction.pyFieldCitation, optional maps onExtractedPageData. Keep existing fields unchanged so current tests still pass.docling/utils/extraction_grounding.py(new)ground_extracted_page(page: ExtractedPageData, doc: DoclingDocument, *, cite_sources: bool, confidence_scores: bool, page_confidence: float | None = None) -> ExtractedPageDatagh, no VLM.docling/document_extractor.pyHappy to follow maintainer guidance before any PR (CLA noted).
All reactions