v8.1.1: 🎞️ PowerPoint Slides in Deck Order & Faster LaTeX Parsing
I am pleased to announce the release of officeParser v8.1.1! This patch reads PowerPoint decks in the order the presentation shows them, each slide with its own notes, reads LaTeX text about twice as fast, and tidies what the HTML and Markdown generators write.
Warning
Behavior changes
- A table's place on the page is not its columns' alignment.
TableMetadata.align(HTML's<table data-align>) was written to Markdown as the alignment of every column. Columns are now written as their cells align them, and the table's place is an attribute list under it ({align=center}) in theextendeddialect. Thepandocpreset writes none, since Pandoc shows that line as text. To align columns, settext-alignon the cells. - A yellow highlight is a plain
<mark>in HTML, with nodata-colororstyle, and==text==in Markdown however the colour is written. Other colours are written as before.
Everything else is a fix: output only moves where it was wrong before.
🔧 What's Fixed
- PowerPoint slides are read in deck order, with their own notes. Order and notes were taken from the numbers in the parts' file names, so a deck whose slides had been moved or deleted came out in the wrong order, and notes could land on the slide next to their own.
slideNumberis now the slide's place in the deck. If you stored slide numbers or note ids from an earlier version, re-parse. - LaTeX text is read in runs, not a character at a time. Ordinary prose parses about twice as fast, and a long stretch of plain text many times faster. The tree a document parses to is unchanged.
- Windows-1252 text is right on every Node version. The
TextDecoderof Node 22.13.0 to 22.22.0 reads the curly quotes, dashes and euro sign of Windows-1252 as control characters (nodejs/node#56542). officeParser now decodes it itself, so HTML, RTF and LaTeX files read the same whatever the runtime. - The browser bundles load under a strict Content Security Policy. They no longer evaluate a string as they load, so a page with
script-src 'self'gets no violation report. - HTML header cells stay header cells. A
<th>, and any cell of a<thead>, was read as an ordinary cell, so a table could lose its header row in HTML, DOCX, ODT and LaTeX output. - HTML task items are marked. Each item of a task list carries
data-type="taskItem"besidedata-checked. - A column's alignment is written once in a Markdown table, not again as a
<div style="text-align: …">in each cell. - CSS colour keywords are not colours.
color: inherit,unset,currentcolor,initialand atransparentbackground are no longer read from HTML as colours and written back. - LaTeX output asks for the page a person would. The preamble is
\usepackage[a4paper,margin=1in]{geometry}, and the page is exactly the size it names: lengths written in TeX'sptcame out 0.4% small. A LaTeX length is also read in the unit it is written in, and a line break that ends a paragraph is kept. - Visualizer: the LaTeX preview looks like the LaTeX build. It is set in LaTeX's own fonts, at the sizes, on the paper and inside the margins the generated source asks for, where it used to be a web page in a sans-serif font.
- Inline DOCX markup has a fixed nesting limit. Past 512 levels inside a paragraph the parser raises
MAX_NESTING_DEPTH_EXCEEDED. Where it stopped used to depend on the runtime's call stack.
✍️ New: the officeParser blog
Seven posts on how document formats work inside and how officeParser reads them: officeparser.harshankur.com/blog (RSS).
🛠 Getting Started
npm install officeparser@8.1.1🔗 Full Changelog: View v8.1.1 details
🔗 Documentation & Visualizer: officeparser.harshankur.com