pandoc 3.12 #11913
jgm
announced in
Announcements
pandoc 3.12
#11913
Replies: 1 comment
|
Unofficial Linux/RISC-V (64-bit, |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I'm pleased to announce the release of pandoc 3.12,
available in the usual places:
Binary packages & changelog: https://github.com/jgm/pandoc/releases/tag/3.12
Source & API documentation: http://hackage.haskell.org/package/pandoc-3.12
This release focuses on performance. Here are some benchmarks showing
the improvement since pandoc 3.10.2:
There shoud be even more dramatic gains in image-heavy documents, because
of optimizations in ImageSize.
In addition to performance improvements, there are many other small
improvements and bug fixes. See the changelog for full details.
Most of the performance gains and and many of the bugs were found
with the help of Claude Fable.
API changes:
compactifyTable. FromText constraint added to the signatures
of htmlAttrs and tagWithAttrs.
Other changes of note:
-t xmlnow respects--standalone, and produces afragment (just the blocks) if it is not provided, like
-t native.Thanks to all who contributed, especially new contributors
Andonome, Clar Fon, Gaurav Vijay Jadhav, Robert Szarka,
Samuel Huang, Yusuf Efe, and zenor0.
Click to expand changelog
Markdown reader:
alertis now added to the produced Divs.base64DataURI.bareURLbefore trying uri/emailAddress.parseWithString'when parsing a block quote or list item. Otherwise it can happen that by the timeparseWithString'is called, the position has already been set to the next file on the command line. Fixes an odd bug withrebase_relative_paths(rebase_relative_paths extension: image reference inside final block object of file is incorrectly pathed to next input file of pandoc command #11888).rebase_relative_paths: recognize URLs with unknown schemes (rebase_relative_paths rewrites URLs whose scheme is not in Text.Pandoc.URI.schemes #11858).takeWhile1Pin the hot inline parsersstr,code,enclosure, andmmdShortSubscript.parseWithString'when parsing a block quote or list item (rebase_relative_paths extension: image reference inside final block object of file is incorrectly pathed to next input file of pandoc command #11888). This ensures thatrebase_relative_pathswill see the right source file.Typst reader:
form: "prose"citations to AuthorInText (feat(typst): mapform: "prose"citations to AuthorInText. #11846, Samuel Huang).pInline, only perform the label-target check forrefelements, not for every inline element.highlightas a mark span (Typst reader: handlehighlightas a mark span #11879, Samuel Huang).#par(explicit paragraph element).LaTeX reader:
\qedto produce U+00A0 (nbsp) instead of BEL.unescapeURLhandling of escaped backslash.\newifnames begin with “if”.##in tokenizer.retokenizeComment.untokenizelinear instead of quadratic.macroDeffast on non-macro-defining commands.peekTokinstead of going throughsatisfyTok.doMacrosstate update for non-macro tokens.inlineCommands, etc.) once per parse.\iftrueetc.Docx reader:
w:tooltip) #11869, Robert Szarka).comment-idinstead ofidin AST for comments.ODT reader:
textPropertieson paragraph styles (pandoc loses local formatting #2623). Previously these just got ignored, not applied to the paragraph’s text.HTML reader:
<input>is a void element.</tr>in tables. The</tr>closing tag is optional in HTML.<col>elements.raw_htmlfor inline<style>elements.<ol class="fancy lower-roman">got DefaultStyle, because the whole class attribute was compared against the known style names. Check each class individually. As a side effect, an unrecognized class no longer prevents falling back to the style attribute.htmlTag: Don’t copy the remaining input on each invocation. A space was appended to the remaining input to guarantee a TagPosition token after the parsed tag; since the input is a strict Text, this copied the entire remaining input every time htmlTag was called (e.g. for every inline HTML tag in a markdown document), giving quadratic behavior in tag-dense documents. Instead, handle the case where the tag is the final token by computing the end-of-input position directly.pSpanLike. The inline dispatcher already knows which span-like element it is looking at, so there is no need for pSpanLike to try a parser.pre/codeattribute precedence for first-wins dedup.pSatisfyas a single parsec primitive.pTagText: when the text contains no character that could parse as anything but Str, Space, or SoftBreak under the enabled extensions (and we are not in a pre element), return B.text directly.pTagContentsto trypStrandpSpacebefore the math, smart punctuation, and raw TeX parsers. This is safe because pStr cannot consume the special characters that start those parsers, and it avoids most guard checks on the slow path. This makes the html reader benchmark about 33% faster and halves its allocation.Muse reader:
strearlier in inline parser. This makes the reader 2x faster and reduces heap allocation by 60%.Org reader:
.pdfas an image format (Directory for intermediary results #11859). This matches the behavior of Emacs, which will render[[file:foo.pdf]]as an image.#+OPTIONS: ^:nildisable all sub-/superscript parsing.-and_in inline footnote labels. Org footnote labels may contain word-constituent characters, hyphens and underscores.RST reader:
lookupGEto find the next anonymous key. Parsing a document with 16000 anonymous links drops from 5.2s to 2.0s.CommonMark reader:
tex_math_gfmhandling.map idinsourceToToksin the common case where the source starts at line 1.AsciiDoc reader:
Djot reader:
Vimwiki reader:
Typst writer:
.typst:no-figure).RST writer:
nowrapto just footnote label, not body. This bug surfaced afternowrapwas fixed in doclayout.Docx writer:
_to start all bookmark names (BookmarkEnd tag placement in docx files #11845). This ensures that they are “hidden” and will not be read by screen readers.withDirection.withDirectionfor Space and SoftBreak.w:tooltip) #11869, Robert Szarka).convertSpacelinear instead of quadratic. On an ad hoc benchmark (4 paragraphs of 80K words each), conversion time drops from 3.2s to 1.1s, and time no longer depends on paragraph length (16 x 20K words previously took 1.8s, now also 1.2s).comment-idinstead ofidin AST for comments. Note that the Docx writer will still interpret anidattribute for legacy compatibility, so if you use markdown files that specifyid, they should still work.TEI writer:
rend, notrenditionattribute, on milestone ([markdown or djot reader, TEI writer] milestones (---) are rendered with wrong "rendition" attribute #11842, Yusuf Efe)renditiontakes pointers to rendition descriptions, whilerendis the free-text attribute, which is what a plain “line” value needs.LaTeX writer:
footnotehyperfor notes in longtable. Instead, generate them manually as we do for floating tables. This removes our dependency on footnotehyper, and resolves a compatibility problem withendfloat(LaTeX writer: the{\def\LTcaptype{none} ... }group around captionless tables breaksendfloat#11857).stInternalLinksstate field. The field has not been read since the writer switched to always adding hypertargets, but we were still doing a full-document query to populate it.sectionHeader. The note-free and link-free variants of the heading text were rendered for every heading, even though the former is only needed when the heading contains a note, image, or identified span, and the latter only for unnumbered, listed headings.inlineListToLaTeX. The strut-insertion and quote-kerning fixups were separate list traversals, allocating an intermediate list each, and this function is called at every level of inline nesting. Combine them into a single pass.lstinlinedelimiter.SOURCE_DATE_EPOCHis set.showHexinstead of (slower)printfintoLabel.RST writer:
\markers. With--reference-linksthis could make the inline reference and its definition render differently, producing a broken RST reference.Native writer:
ppDoc, which shows the document, tokenizes and re-parses the result into a generic Value, and lays that out via Text.PrettyPrint.HughesPJ. The layout step dominated the cost of the writer (and of any pipeline producing native output). We now build a width-cached layout tree directly from the AST and render it with a small renderer that reproduces HughesPJ’s layout algorithm exactly (including the ribbon computation with ribbonsPerLine = 1.2 and the treatment of glued closing delimiters), so the output is byte-for-byte identical to before. Verified against the old binary on all golden .native files and the markdown test corpus at many column widths, in both standalone and plain modes. The writer is about 4x faster. Also remove the pretty and pretty-show dependencies.XML writer:
--standalone. When standalone is selected, we get a full Pandoc element with xml header and metadata. When not, we get a fragment – just the blocks.ppcElement, using a configuration that treats elements with inline content as inline tags, so that no significant whitespace is added inside them. Consecutive text nodes are merged, SoftBreak is written as a literal newline, and whitespace runs that would not survive a roundtrip (e.g." \n"or"\n\n") are encoded as Space and SoftBreak elements.typst:property). Use the common convention of encoding such characters as_xHHHH_, whereHHHHis the hexadecimal code of the character: the writer encodes attribute names (foo:barbecomesfoo_x003A_bar) and the reader decodes them.("k","")disappeared and did not round trip.EPUB writer:
runPure (writeHtmlStringForEPUB ...). Since every one of these invocations starts with a fresh CommonState, the translations YAML file was re-read and re-parsed for every TOC item, which accounted for a significant part of the EPUB writer’s run time on documents with many sections.Powerpoint writer:
ANSI writer:
HTML writer:
intrinsicEventsHTML4. The list hadonmouseouttwice and was missingonmousemove.<p></p>for paragraphs with no rendered content.--id-prefixin EPUB3 footnote section id.strToHtmlmore efficient. Replace theT.groupBy-based implementation, which allocated a list of Text fragments and round-tripped through String, with a simpleT.breakscanner.Org writer:
#+begin_export htmlfor raw HTML blocks.#+begin_htmlwas removed in Org 9.0 (2016).Str "."andStr ")", sinceT.all isDigit ""is True. Require at least one digit.<<>>for spans with no id.=delimiter for inline code containing=. Org has no escape mechanism inside verbatim text, so=code with ==did not parse as verbatim. Fall back to the equivalent~...~delimiter when the content contains=(and no~).org-link-escapedoes: backslash-escape brackets and double backslash runs occurring before a bracket or at the end of the target. Link descriptions cannot contain escapes; instead, likeorg-link-make-string, insert a zero-width space between consecutive closing brackets and before a closing bracket at the end of the description.escapeStringallocated one Doc node per character for any string containing a non-alphanumeric character. Split on the (rare) special characters instead and emit intervening text as single literals. No change in output.length stNotes + 1walked the accumulated note list for every footnote, making note numbering quadratic in the number of notes.Org reader and writer:
^[ \t]*,*(\*|#\+), i.e. lines already starting with commas before*or#+get an additional comma; unescaping removes one comma from such lines. The writer previously left a literal,#+fooline unescaped, and the reader then stripped its comma when reading the result back, corrupting the code on round trips. Writer and reader now both handle runs of commas, matching Emacs.Text.Pandoc.ImageSize:
pdfSize. Treat malformed streams as a parse scanner (so we keep scanning the rest for a/MediaBox).viewBox.findSvgTagby using a single pass. Up to 60X faster on files with few<characters.writerDpifor AVIF images instead of hardcoding 72. With the default options this changes the assumed resolution from 72 to 96 dpi.numUnit. So e.g.width="3 cm"is now recognized.sizeInPoints.Text.Pandoc.SelfContained:
makeDataURI.@importfallback output.</scriptcheck case-insensitive.url(#...)occurrences in SVG attributes. Previously only the first got prefixed rewritten.isHtml5field from ConvertState.roleandaria-labelwhen inlining SVGs. Do not add them to other elements withsrcattributes.Text.Pandoc.UTF8:
toTextandtoTextLazyunconditionally ran a CR-removing filter over the input, allocating a full copy of the document even in the common case where no CRs are present. Check for a CR first (B.elem, a fast memchr) and reuse the input buffer unchanged if none is found; for the lazy variant, do this chunk-wise to preserve laziness. On a 10 MB LF-only input this makestoTextover 4x faster; when CRs are present the extra scan is not measurable.readFileexception-safe. UsewithFileinstead ofopenFileso the handle is closed even if reading throws.Text.Pandoc.Class:
toTextM. Skip the CR-filtering copy when the input contains no CRs, as already done in Text.Pandoc.UTF8.toText.runSilentlyerror-safe. Previously, if the action passed to runSilently threw an error that was later caught, the verbosity remained pinned at ERROR and all previously accumulated log messages were lost. Now the original log and verbosity are restored even when the action fails.isRelativeToParentDir. Compare the first path component rather than just looking at a prefix, to correctly handle paths like..foo/bar.yaml.extractURIData. The base64 indicator in a data URI is the final parameter of the media type and may follow other parameters, e.g.charset. Previously, the code expected;base64to be the only parameter.extractURIData.toTextMerrors. Scan for the first invalid UTF-8 sequence and report its actual position and byte.setNoCheckCertificate. The HTTP manager is created lazily with TLS settings based onstNoCheckCertificateand then cached in CommonState, so changing the option after the first request had no effect. Discard the cached manager when the option’s value changes.addToFileTree.openURLinto Text.Pandoc.Class.IO.HTTP. Commit 455bea9 added the new module but did not register it in pandoc.cabal or remove the original definitions.logOutput: avoid multiplehPutStrLn, which can cause confusing interleaving.Text.Pandoc.MediaBag:
data:andfile:URI schemes case-insensitively.mediaContentsis left lazy so contents need not be forced at insert time.Text.Pandoc.URI.isURIincanonicalize.Network.URI.isURItreats Windows drive-letter paths likec:/foo.pngas URIs...as a path component. The insertMedia check usedisInfixOf, so a harmless name like foo..bar.png was silently renamed to its content hash..and..components incanonicalize.normalisedoes not remove redundant path components, soimg/../a.pnganda.pngwere distinct keys. UsemakeCanonical(as PandocPure’s FileTree already does for its path-indexed map), which also handles duplicate and trailing slashes, replacing backslashes with slashes first.mediaPathcollisions between keys. The friendly mediaPath was derived by percent-unescaping the key, so distinct keys likea%20b.pnganda b.pngproduced the same mediaPath (“a b.png”) and silently clobbered each other on extraction (and inside docx/epub archives). Now the original name is only kept if the key contains no percent sign, somediaPathequals the key and distinct keys yield distinct paths; anything percent-encoded gets a content-hash name. Hashed names can only coincide for identical contents, which is harmless.Text.Pandoc.XML.Light:
Fix escaping of repeated
]]>inCDATA.Make
escStrmore efficient.Avoid round-trips in
ppCDataSprettify path. This makesshowCDataandppcCDataunused, so they are removed.Parse XML fragments from the event stream instead of using xml-conduit’s document parser, which requires a single root element. Our earlier woraround with a wrapper element was fragile. Behavior changes:
Use
text-builderpackage for rendering, instead of text’s lazy Text builder. On a large document this cuts docx conversion time by about 13%; output is byte-for-byte identical.Expose
ppcTopElementfrom Output.ConfigPP now has a field
inlineTagthat checks for inline tags. Inline tags are printed on one line and not indented, by default. ExportuseInlineTags,prettyConfigPP.Text.Pandoc.Sources:
takeWhileP/takeWhile1Pcombinators [API change]. Character streams over Sources previously had to be consumed one character at a time viasatisfy, at a cost of several allocations and monadic binds per character. The new combinators scan a whole run of matching characters with a single parser invocation usingT.span, while replicating the exact semantics ofT.pack <$> many/many1 (satisfy f), including empty-chunk handling, position updates at chunk boundaries, and parsec’s error messages.takeWhileP/takeWhile1Pcombinators across readers.Text.Pandoc.Translations:
setTranslationsnow keeps an already-loaded translation table when the language is unchanged, instead of unconditionally clearing the cache and forcing a re-read and re-parse of the translations YAML file on the nexttranslateTerm.Text.Pandoc.Chunks:
nav-pathattribute andrmNavAttrswalk. This no longer did anything; output is unaffected.compactifyTablefor tables produced by all readers. Remove old ad hocparaToPlainat the table cell level. This should ensure that we don’t get tables that mix Plain and Para (Docx writer: Para blocks in table cells get "Body Text" spacing (9pt before/after), bloating table rows #11864). Such tables tend to look funny when rendered in docx and other formats.Text.Pandoc.Data:
getDataFileNameswith-embed_data_files. Previously it was not looking in the right directory and not recursing.Text.Pandoc.Shared:
taskListItemFromAscii: Fix incorrect treatment of[ ]as checked.stringifyInlines, a single-passstringifyfor inlines [API change]. This is about 6x faster thanstringifyfor long inline sequencesstringifyby making it accumulate[Text]and concatenate once at the end, instead of mappending at every node.stringifyInlinesinstead ofstringifywhere possible.compactifyTable. [API change] This converts cells that consist in a single Para block to a Plain, provided the table contains only such cells (or empty cells).decrementTrailingRowSpans. Entries whose RowSpan fell to 0 were left in the map, relying on every consumer to guard against them. Delete them instead.toTaskListItem.endsWithPlainlook inside DefinitionList. endsWithPlain recursed into the last item of BulletList and OrderedList but ignored DefinitionList, so list items ending with a compact definition list were treated as loose by the RST, Org, and Haddock writers.htmlAttrs. This adds a FromText constraint tohtmlAttrsandtagWithAttrs[API change].ensureValidXmlIdentifiers.Text.Pandoc.Writers.Shared:
splitSentences.lookupMetaBool: treat empty block or inline list as False.htmlAttrs: escape id and class attributes, like the others.stripLeadingTrailingSpace- strip multiple Space, if present.toSubscript: handle minus sign.ensureValidXmlIdentifiersfor Figure and table sub-elements, resolving a bug that produced broken internal links in the HTML4/XHTML, EPUB, DocBook, TEI, ICML, FB2, and ODT writers.Text.Pandoc.Parsing:
uriSchemeby using a trie. Up to 15% faster in URL-heavy documents withautolink_bare_uris.$checks in math when delim is\(or\\(.anyOrderedListMarkermore efficient.HTML template:
reveal.js template:
flake.nix: parse allow-newer and allow-newer-deps in stack.yaml.
Make
embed_data_filesflag default to True. Remove flag settings from cabal.project. This makes it possible to override it on the command line.Depend on commonmark 0.3.1, commonmark-extensions 0.2.7.3, commonmark-pandoc 0.3.0.2 (major performance improvements).
Depend on released asciidoc 0.1.1 (major performance improvements).
Use released texmath 0.13.3 (major performance improvements).
Depend on released djot 0.1.4.3 (major performance improvements).
Use released citeproc 0.14 (major performance improvements).
Use released doclayout 0.6.
Use released zip-archive 0.5 (major performance improvements).
Use released skylighting-0.15 (major performance improvements).
Depend on released doctemplates 0.11.1.
Depend on released typst 0.12.
Require text >= 2.0.
Bump upper bound for unicode-data.
Allow crypton 2.0.x.
Allow Diff 2.0.
Add
tools/diff-golden-tests.sh.Fix
tools/diff-zip.shon non-Darwin.Add
tools/benchplot.js. This creates a nice graph comparing two benchmarks.typst-properties.md: fix fill syntax in Typst property examples (Fix fill syntax in Typst property documentation #11855, zenor0).Remove tested-with from cabal file. We tend not to keep it up to date.
Fix typo in Lua filter example (Docs: typo #11875, Andonome).
Fix a bug in jats-reader.xml (duplicate attribute)
This discussion was created from the release pandoc 3.12.
All reactions