Skip to content

v2.6.1

Choose a tag to compare

@lxman lxman released this 09 Sep 21:58

Fixed

  • Malformed classic cross-reference subsections are repaired during load. Some PDF producers
    declare a subsection as 1 N while emitting entries for objects 0..N-1, beginning with the
    object-0 free-list head. PdfLibrary previously shifted every offset by one object number and could
    resolve the trailer's /Root to /Pages or /Page instead of /Catalog. The reader now corrects
    the subsection only when both the object-0 free-head signature and the contradictory trailer
    /Size confirm the defect.
  • Text extraction now infers spaces from large TJ positioning gaps. PDFs that position each
    word separately without storing literal space characters no longer produce concatenated clauses;
    adjustments above 0.2 em insert a separator while ordinary kerning remains within a word.

Added

  • PdfPage.ExtractTextWithQuality() detects implausible Latin text layers. Its
    TextExtractionResult returns the extracted text, printable-ASCII ratio, unexpected-control and
    replacement-character counts, plus IsLikelyReadable. This lets callers route pages whose fonts
    lack a usable /ToUnicode mapping to OCR instead of trusting non-empty but garbled text.