You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I maintain an accessibility-focused document pipeline (PDF/UA and EPUB 3 output, veraPDF
and EPUBCheck/Ace in the loop) and I run opendataloader-pdf as one of our backends,
including struct-tree mode. I want to float an idea that I think fills a gap in this
project specifically, and in the field generally.
The gap: there is no open reference for whether a tag tree is correct.
Conformance and correctness are different things, and only one of them is measurable
today. I auto-tagged a document with --format tagged-pdf on 2.5.8 and ran veraPDF
against PDF/UA-1: 97 of 106 rules pass, 16,953 checks pass. That is a good result, and
the nine failures are the enterprise export step by design, so this is not a complaint.
The point is what those 16,953 checks cannot tell you:
whether a /TD should have been a /TH (on a clean table page the output had 41 /TD and zero /TH, with no Scope anywhere);
whether a heading is at the right level;
whether the reading order a screen-reader user traverses matches the author's intent;
whether Alt text says anything (figures came out as "image 1", "image 2").
A file can pass PDF/UA-1 and be wrong in all four ways. No validator catches it, because
validators check conformance, not intent — and there is no open corpus of documents where
the right answer is recorded.
Where this touches odl-bench directly. The benchmark measures reading order, tables
and headings against extracted text, which is the right design for extraction. But as far
as I can tell none of the 200 documents carries a /StructTreeRoot, so there is no way to
score auto-tagging against an authored tree — including the obvious evaluation, which is to
take a well-tagged document, strip its tag tree, re-tag it automatically, and compare. That
is a clean and honest metric for a tagger and it is impossible without tagged documents.
I also noticed your roadmap lists structure validation — "verify PDF tag trees" — for
Q3 2026. That is the same problem from the other end: a validator needs a reference for
what the correct tree is. This is really why I am posting now rather than later.
The idea. A commissioned corpus: public-domain and government documents — rights-clean
by construction, no customer data, no scraping — stratified across the document types
remediation actually meets (forms, reports, letters, tenders, manuals, notices,
table-heavy, multi-column), each tagged to PDF/UA by professional remediation vendors
against a published rubric. Element roles, reading order, tag hierarchy, table structure,
alt-text presence, plus geometry. Phased, 1k → 10k pages.
Three things would make it usable rather than merely large:
A published annotation standard with equivalence classes — which tagging
differences are defensible variation (heading-level granularity, list-vs-paragraph
boundaries, table-header scope conventions) versus actual defects. Without this, a
tagging benchmark punishes correct work.
An inter-professional agreement study. The pilot triple-tagged by three independent
vendors, analysed per document type and complexity. As far as I know nobody has
measured how much professional PDF/UA remediators actually agree — a standalone result,
and the evidence base for those equivalence classes.
Multi-reference scoring. The benchmark subset stays fully multi-tagged and
adjudicated, so a tagger is scored against the set of acceptable taggings rather than
one annotator's house style — the same reasoning as multi-reference MT evaluation. This
is the piece that would let odl-bench carry a fair accessibility tier.
Licensing would be CDLA-Permissive for the data, CC-BY for the rubric, Apache for a
baseline model. No exclusivity.
What I am asking. Mostly: would this be useful to you, and would the rubric be worth
co-designing so the corpus is directly consumable by your tagger and your benchmark? I
would rather shape it around a real consumer than publish something academically tidy that
nobody's pipeline can ingest.
Happy to go into specifics — including how it would be funded and run — off this thread if
there is interest. And happy to be told the gap looks different from where you sit; that
would be useful too.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
I maintain an accessibility-focused document pipeline (PDF/UA and EPUB 3 output, veraPDF
and EPUBCheck/Ace in the loop) and I run
opendataloader-pdfas one of our backends,including struct-tree mode. I want to float an idea that I think fills a gap in this
project specifically, and in the field generally.
The gap: there is no open reference for whether a tag tree is correct.
Conformance and correctness are different things, and only one of them is measurable
today. I auto-tagged a document with
--format tagged-pdfon 2.5.8 and ran veraPDFagainst PDF/UA-1: 97 of 106 rules pass, 16,953 checks pass. That is a good result, and
the nine failures are the enterprise export step by design, so this is not a complaint.
The point is what those 16,953 checks cannot tell you:
/TDshould have been a/TH(on a clean table page the output had 41/TDand zero/TH, with noScopeanywhere);Alttext says anything (figures came out as"image 1","image 2").A file can pass PDF/UA-1 and be wrong in all four ways. No validator catches it, because
validators check conformance, not intent — and there is no open corpus of documents where
the right answer is recorded.
Where this touches
odl-benchdirectly. The benchmark measures reading order, tablesand headings against extracted text, which is the right design for extraction. But as far
as I can tell none of the 200 documents carries a
/StructTreeRoot, so there is no way toscore auto-tagging against an authored tree — including the obvious evaluation, which is to
take a well-tagged document, strip its tag tree, re-tag it automatically, and compare. That
is a clean and honest metric for a tagger and it is impossible without tagged documents.
I also noticed your roadmap lists structure validation — "verify PDF tag trees" — for
Q3 2026. That is the same problem from the other end: a validator needs a reference for
what the correct tree is. This is really why I am posting now rather than later.
The idea. A commissioned corpus: public-domain and government documents — rights-clean
by construction, no customer data, no scraping — stratified across the document types
remediation actually meets (forms, reports, letters, tenders, manuals, notices,
table-heavy, multi-column), each tagged to PDF/UA by professional remediation vendors
against a published rubric. Element roles, reading order, tag hierarchy, table structure,
alt-text presence, plus geometry. Phased, 1k → 10k pages.
Three things would make it usable rather than merely large:
differences are defensible variation (heading-level granularity, list-vs-paragraph
boundaries, table-header scope conventions) versus actual defects. Without this, a
tagging benchmark punishes correct work.
vendors, analysed per document type and complexity. As far as I know nobody has
measured how much professional PDF/UA remediators actually agree — a standalone result,
and the evidence base for those equivalence classes.
adjudicated, so a tagger is scored against the set of acceptable taggings rather than
one annotator's house style — the same reasoning as multi-reference MT evaluation. This
is the piece that would let
odl-benchcarry a fair accessibility tier.Licensing would be CDLA-Permissive for the data, CC-BY for the rubric, Apache for a
baseline model. No exclusivity.
What I am asking. Mostly: would this be useful to you, and would the rubric be worth
co-designing so the corpus is directly consumable by your tagger and your benchmark? I
would rather shape it around a real consumer than publish something academically tidy that
nobody's pipeline can ingest.
Happy to go into specifics — including how it would be funded and run — off this thread if
there is interest. And happy to be told the gap looks different from where you sit; that
would be useful too.
All reactions