Skip to content

v8.0.0

Choose a tag to compare

@github-actions github-actions released this 16 Sep 21:58
· 197 commits to master since this release

v8.0.0: 📄 A Ground-Up PDF Rewrite, Native DOCX & ODT Generation, and Encrypted Documents

I am pleased to announce the release of officeParser v8.0.0! This is the biggest step forward for PDF handling the library has taken. PDF text extraction has been rewritten from the ground up: instead of a flat page of lines, a PDF now yields real structure (headings with levels, tables with merged cells, lists, notes, multi-column reading order, and text colour), taken from the document's tags where they exist and reconstructed from geometry where they do not. Alongside the PDF work, two brand-new generators, to('docx') and to('odt'), turn any parsed source into a real Word or OpenDocument file; password-protected documents now decrypt and parse across every encryptable format; a new engine fills DOCX templates for mail-merge; a native pdf-lib engine generates PDFs with no headless browser (in Node and in the browser); and .odg (LibreOffice Draw) joins the parse list, bringing the count to 13 formats.

This is a major version, so there are breaking changes. Every one of them has a migration path: a better default with a flag to restore the old behavior, or a long-deprecated option finally removed. The critical ones are listed first below; the full changelog has the exhaustive list.


💥 Breaking Changes

  • Node.js >=22.13 is now required (was >=18), matching the bundled pdfjs-dist floor.
  • PDF output changed shape and content by default. A tagged PDF now yields heading / table / row / cell / list / note nodes and one paragraph per paragraph, not a page of flat lines. Word spacing, column reading order, hyphenation and super/subscripts are all corrected, so the plain text differs too. Every node also carries a bounds box by default; set ignorePageGeometry: true to omit it and restore the smaller AST.
  • ast.toText() was removed. Use (await ast.to('text')).value, which is the same content at its defaults and is fully configurable. The CLI's --toText flag is gone; use --to=text.
  • Long-deprecated config options were removed: outputErrorToConsole (use onWarning), ocrLanguage (use ocrConfig.language), ocrConfig.autoTerminateTimeout (use ocrConfig.timeout.autoTerminate), and putNotesAtLast (notes attach to node.notes). Each now raises UNRECOGNIZED_CONFIG_OPTION naming its replacement rather than being ignored in silence.
  • PDF images are now PNG, not BMP. Extracted images are pdf_image_p<page>_<n>.png with mimeType: 'image/png' (was .bmp / image/bmp), and roughly an order of magnitude smaller. Update any code that filters attachments by a .bmp suffix or the image/bmp type.
  • includeImages now defaults to 'image-only', which no longer leaks an image's recognized text into the Markdown fallback or the HTML alt. Pipelines that fed OCR text into RAG through Markdown/HTML should opt back in with 'image+ocr-text' or 'ocr-text-only'.
  • Markdown and fragment HTML no longer inline an image over maxInlineImageBytes (1.5 MB default). An over-cap image renders its recognized text (or a compact name reference) and raises IMAGE_NOT_INLINED; raise the cap (or set Infinity) to restore inlining. Standalone HTML always inlines.
  • ODF comments and Writer master-page headers/footers now parse by default (v7 dropped both entirely), so they surface on the node's .comments / in ast.auxiliary and flow into every generated output. Set ignoreComments / ignoreHeadersAndFooters to restore the old output.
  • PDF running headers and footers route to ast.auxiliary by default, so .to('text') and chunked output no longer repeat the running header on every page.
  • puppeteer (>=22) is now a declared optional peer dependency for the default to('pdf') engine, and pdf-lib an optional peer for the native engine. Projects that never generate PDFs are unaffected.

✨ What's New

1. A ground-up PDF extraction rewrite

PDF parsing no longer returns a page of flat paragraphs. When a PDF is tagged, headings, tables, lists and notes come from its structure tree; when it is not, they are reconstructed geometrically (a recursive XY-cut for multi-column and float-beside-text reading order, plus grid-table and list recovery). Merged cells recover colSpan / rowSpan, internal links resolve to the target section (not just the page), rotated runs are recovered instead of dropped, and run fill colour and highlight are extracted into formatting.color / formatting.backgroundColor by default (pdfParserConfig.extractTextColor). A new pdfParserConfig mirrors htmlParserConfig (useTags, detectColumns, headingDetection, pageRange, and more), and outline, page labels, permissions and AcroForm values are surfaced.

2. DOCX generation, to('docx')

A new generator writes a real WordprocessingML .docx from any parsed source: headings, styled runs, tables with merged cells, nested lists, images, hyperlinks, bookmarks, footnotes/endnotes/comments, code blocks and metadata. It adds no dependencies (built on the bundled fflate), runs identically in Node and the browser, is byte-reproducible, and round-trips cleanly back through the Word parser. Every value written into the XML is sanitized.

const docx = (await ast.to('docx')).value; // Uint8Array

3. ODT generation, to('odt')

The round-trip partner of the ODF parser writes a real OpenDocument Text package (headings, ODF whitespace encoding, nested lists, merged cells and header rows, images, links, footnotes, comments and metadata), with each distinct formatting bundle interned into one named automatic style for byte-stable output. Configurable via odtConfig (format, landscape, margin).

4. Password-protected documents, across every encryptable format

Encrypted OOXML (.docx / .xlsx / .pptx, agile and standard) and encrypted ODF (.odt / .ods / .odp / .odg) now decrypt and parse, and PDF decryption is exposed through the same options. One unified, top-level config drives every format, with unified error codes.

const ast = await parseOffice(file, { password: 'secret' });
// or supply it lazily/interactively:
const ast2 = await parseOffice(file, { onPassword: async () => promptUser() });

Decryption happens in a new zero-dependency crypto module, and untrusted input is bounded (key-stretch and second-layer decompression work is capped) so a hostile descriptor cannot hang the event loop.

5. DOCX templating and mail-merge

Fill a DOCX template's {{placeholder}} tags from a data object and get back a new .docx with all formatting intact; pass an array of objects to get one document per entry (a batch mail-merge). The substitution is run-aware, so a placeholder Word split across runs is still filled, and placeholders are matched in the body, headers, footers, notes and comments.

const filled = await renderTemplate(templateBytes, { data: { name: 'Ada', amount: '$42' } });
const batch  = await renderTemplate(templateBytes, { data: [ {name: 'Ada'}, {name: 'Alan'} ] }); // Uint8Array[]

Thank you Raymond Camden for suggesting this DOCX templating feature.

6. A native PDF generation engine (no browser, works client-side)

pdfConfig.engine: 'native' lays a PDF out directly with pdf-lib instead of printing HTML through a headless browser. It needs no Chrome, runs the same in Node and the browser, and is much lighter. A dedicated officeparser/browser-native-pdf entry keeps pdf-lib external so a self-bundling consumer can produce real PDF bytes entirely client-side.

const pdf = (await ast.to('pdf', { pdfConfig: { engine: 'native' } })).value;

7. ODG parsing and OCR layout recovery

.odg (and .otg templates) join the ODF family: each draw:page becomes a page node with its shape text, embedded tables and images. Separately, OCR output now rebuilds the recognized text's 2-D layout from Tesseract's word boxes (ocrConfig.preserveLayout, default on), so a scanned table or multi-column page keeps its columns instead of collapsing to a flat string.


🔧 Performance, Correctness & Security

  • PDF parsing is dramatically faster. The per-page dependency wait polled the wrong pdf.js object pool for font ids, so it burned a 500 ms idle timeout per font per page; each dependency is now routed to the pool that holds it. Two hot paths that were quadratic (buildLines and the geometric table/list scans) are now linear or bounded. This applies equally to the browser build.
  • Header rows are read from real DOCX and ODT files (the parsers ignored w:tblHeader and table:table-header-rows), a DOCX vertical-merge continuation cell no longer adds a stray empty paragraph on re-parse, and a per-PDF worker leak is fixed.
  • onNode now fires exactly once per node in HTML output, and ocr: true without extractAttachments no longer silently does nothing in any format (it raises OCR_REQUIRES_ATTACHMENTS).
  • Security hardening: global prototype pollution via list-numbering maps is closed (null-prototype maps), an ODF text:c / indentation OOM is clamped, and PDF image decoding is capped at 40 megapixels so a decompression-bomb image cannot exhaust memory.

🛠 Getting Started

npm install officeparser@8.0.0

🔗 Full Changelog: View v8.0.0 details
🔗 Documentation & Visualizer: officeparser.harshankur.com