I'm pleased to announce the release of pandoc 3.12,
available in the usual places:
Binary packages & changelog: https://github.com/jgm/pandoc/releases/tag/3.12
Source & API documentation: http://hackage.haskell.org/package/pandoc-3.12
This release focuses on performance. Here are some benchmarks showing
the improvement since pandoc 3.10.2:
There shoud be even more dramatic gains in image-heavy documents, because
of optimizations in ImageSize.
In addition to performance improvements, there are many other small
improvements and bug fixes. See the changelog for full details.
Most of the performance gains and and many of the bugs were found
with the help of Claude Fable.
API changes:
- Text.Pandoc.ImageSize: ImageType now derives Eq.
- Text.Pandoc.Sources: new functions takeWhileP, takeWhile1P.
- Text.Pandoc.Shared: new functions stringifyInlines,
compactifyTable. FromText constraint added to the signatures
of htmlAttrs and tagWithAttrs.
Other changes of note:
-t xmlnow respects--standalone, and produces a
fragment (just the blocks) if it is not provided, like-t native.- The default CSS now includes dark mode support.
Thanks to all who contributed, especially new contributors
Andonome, Clar Fon, Gaurav Vijay Jadhav, Robert Szarka,
Samuel Huang, Yusuf Efe, and zenor0.
Click to expand changelog
-
Markdown reader:
- Make alert keywords case-insensitive (#11836, Hendrik Erz). In addition, the class
alertis now added to the produced Divs. - Fix stream position handling in
base64DataURI. - Cheaply reject
bareURLbefore trying uri/emailAddress. - Reset sourcepos in
parseWithString'when parsing a block quote or list item. Otherwise it can happen that by the timeparseWithString'is called, the position has already been set to the next file on the command line. Fixes an odd bug withrebase_relative_paths(#11888). rebase_relative_paths: recognize URLs with unknown schemes (#11858).- Use
takeWhile1Pin the hot inline parsersstr,code,enclosure, andmmdShortSubscript. - Reset sourcepos in
parseWithString'when parsing a block quote or list item (#11888). This ensures thatrebase_relative_pathswill see the right source file.
- Make alert keywords case-insensitive (#11836, Hendrik Erz). In addition, the class
-
Typst reader:
- Map
form: "prose"citations to AuthorInText (#11846, Samuel Huang). - Make the handler maps monomorphic by wrapping the handlers in newtypes with polymorphic fields (BlockHandler, InlineHandler).
- Store document labels in a Set instead of a list.
- In
pInline, only perform the label-target check forrefelements, not for every inline element. - Handle
highlightas a mark span (#11879, Samuel Huang). - Collapse citations around soft break (#11897).
- Handle block content in inline element bodies (#11881, Samuel Huang).
- Handle
#par(explicit paragraph element).
- Map
-
LaTeX reader:
- Support LaTeX3 (xparse) document commands (#7540). Support the features described in the LaTeX usrguide.
- Fix
\qedto produce U+00A0 (nbsp) instead of BEL. - Fix
unescapeURLhandling of escaped backslash. - Require that
\newifnames begin with “if”. - Fix position off-by-one after
##in tokenizer. - Fix doubled source positions in
retokenizeComment. - Make raw token capture O(1) per token.
- Make
untokenizelinear instead of quadratic. - Fail
macroDeffast on non-macro-defining commands. - Peek at next token directly in
peekTokinstead of going throughsatisfyTok. - Skip
doMacrosstate update for non-macro tokens. - Build command dispatch maps (
inlineCommands, etc.) once per parse. - Keep ASCII quotes when ligatures are disabled.
- Don’t discard state changes made in optional arguments.
- Keep nested conditionals balanced in
\iftrueetc.
-
Docx reader:
- Read ScreenTips as link titles (#11869, Robert Szarka).
- Use
comment-idinstead ofidin AST for comments.
-
ODT reader:
- Rewrite in monadic style. Net -994 lines. No changes to test output.
- Handle
textPropertieson paragraph styles (#2623). Previously these just got ignored, not applied to the paragraph’s text.
-
HTML reader:
- Don’t require a closing tag for checkbox inputs.
<input>is a void element. - Allow omitted
</tr>in tables. The</tr>closing tag is optional in HTML. - Don’t drop cells when there are too few
<col>elements. - Let the first of duplicate attributes win. HTML specifies that the first occurrence wins.
- Respect
raw_htmlfor inline<style>elements. - Use a Map for the note table.
- Report the noteref position for unresolved notes. The ReferenceNotFound warning was logged after the whole document had been parsed, so it always pointed at the end of the input. Record the position of the first reference to each note and use it in the warning.
- Find list style anywhere in the class attribute. Previously
<ol class="fancy lower-roman">got DefaultStyle, because the whole class attribute was compared against the known style names. Check each class individually. As a side effect, an unrecognized class no longer prevents falling back to the style attribute. - Limit iframe nesting depth.
htmlTag: Don’t copy the remaining input on each invocation. A space was appended to the remaining input to guarantee a TagPosition token after the parsed tag; since the input is a strict Text, this copied the entire remaining input every time htmlTag was called (e.g. for every inline HTML tag in a markdown document), giving quadratic behavior in tag-dense documents. Instead, handle the case where the tag is the final token by computing the end-of-input position directly.- Pass the tag name to
pSpanLike. The inline dispatcher already knows which span-like element it is looking at, so there is no need for pSpanLike to try a parser. - Fix
pre/codeattribute precedence for first-wins dedup. - Implement
pSatisfyas a single parsec primitive. - Add a fast path to
pTagText: when the text contains no character that could parse as anything but Str, Space, or SoftBreak under the enabled extensions (and we are not in a pre element), return B.text directly. - Reorder
pTagContentsto trypStrandpSpacebefore the math, smart punctuation, and raw TeX parsers. This is safe because pStr cannot consume the special characters that start those parsers, and it avoids most guard checks on the slow path. This makes the html reader benchmark about 33% faster and halves its allocation.
- Don’t require a closing tag for checkbox inputs.
-
Muse reader:
- Try
strearlier in inline parser. This makes the reader 2x faster and reduces heap allocation by 60%. - Check raw line before parsing table rows. Table row parsers are speculatively tried at every paragraph line boundary via the terminator continuation. This change gives another 2x speed improvement.
- Try
-
Org reader:
- Fix hard parse failure on bare backslash.
- Recognize
.pdfas an image format (#11859). This matches the behavior of Emacs, which will render[[file:foo.pdf]]as an image. - Make
#+OPTIONS: ^:nildisable all sub-/superscript parsing. - Allow
-and_in inline footnote labels. Org footnote labels may contain word-constituent characters, hyphens and underscores. - Fix swapped arguments in exportSettings parser.
- Don’t lowercase meta keyword twice.
- Avoid quadratic complexity when parsing table rows. Parsing a 32000-row table drops from 7.8s to 0.9s (and scales linearly instead of quadratically); maximum residency also improves.
- Use a Set for anchor ids. Reading a document with 64000 anchors and as many internal links drops from 23.5s to 5.2s and now scales linearly.
-
RST reader:
- Use a predicate to test for for special characters.
- Avoid double-parsing lines preceding a non-underline.
- Fail fast when no link can start at this position.
- Use
lookupGEto find the next anonymous key. Parsing a document with 16000 anonymous links drops from 5.2s to 2.0s. - Replace association list with Map for resolving note references. Parsing a document with 16000 named notes drops from 5.8s to 3.5s, and scaling is now linear.
- Fix a bug that led to an empty class for code blocks with no specified language.
-
CommonMark reader:
- Avoid nested walks in
tex_math_gfmhandling. - Avoid rebuilding the token list with
map idinsourceToToksin the common case where the source starts at line 1.
- Avoid nested walks in
-
AsciiDoc reader:
- Resolve footnotes, stem, and icons during conversion. Previously the reader made three separate passes over the parsed AsciiDoc AST (for footnotes, stem math types, and icons) before converting to the pandoc AST. Instead, resolve all three during the toPandoc conversion, which already traverses everything once, threading footnote state and document attributes through a StateT layer. Table headers are now converted before body rows so that footnote resolution follows document order. This makes the reader about twice as fast on typical documents.
-
Djot reader:
- Fix a bug that led to an empty class for code blocks with no specified language.
-
Vimwiki reader:
- Split class attribute and put it in the proper slot in pandoc’s Attr.
-
Typst writer:
-
RST writer:
- Apply
nowrapto just footnote label, not body. This bug surfaced afternowrapwas fixed in doclayout.
- Apply
-
Docx writer:
- Fix sizing for images (#11838). Setting size to a percent now scales image to percent of page width. Previously a percent width or height would just provide a maximum bound rather than scaling.
- Include default style even if paragraph has non-style props (#11867). Previously a paragraph that was, e.g. center-aligned would be missing its default Body Text style. Also ensure that any block sequence (list items, table cells, block quote, div), First Paragraph is set for the first paragraph in the item.
- Use
_to start all bookmark names (#11845). This ensures that they are “hidden” and will not be read by screen readers. - Honor CSL hanging-indent and spacing hints (#11871, Samuel Huang).
- Add a fast path to
withDirection. - Avoid double
withDirectionfor Space and SoftBreak. - Map link/image titles to ScreenTips.
- Write link titles as ScreenTips (#11869, Robert Szarka).
- Properly signal error if reference docx can’t be parsed. Previously this led to an incomprehensible error at a later phase.
- Make
convertSpacelinear instead of quadratic. On an ad hoc benchmark (4 paragraphs of 80K words each), conversion time drops from 3.2s to 1.1s, and time no longer depends on paragraph length (16 x 20K words previously took 1.8s, now also 1.2s). - Cache token-type styles instead of rebuilding per Code. On an ad hoc benchmark with 100K inline code spans, conversion time drops from 2.4s to 1.5s.
- Add FirstParagraph class after display math (#11900). This way the continuation text can be styled flush-left in a style where normal paragraphs are indented.
- Treat a “mixed” widths table as all-default (#11899). Previously if some columns were ColWidthDefault and others ColWidth 0.x, we would get columns with a specified width of 0 for the default ones. Instead, treat all columns as default in this case.
- Use
comment-idinstead ofidin AST for comments. Note that the Docx writer will still interpret anidattribute for legacy compatibility, so if you use markdown files that specifyid, they should still work.
-
TEI writer:
- Use
rend, notrenditionattribute, on milestone (#11842, Yusuf Efe)renditiontakes pointers to rendition descriptions, whilerendis the free-text attribute, which is what a plain “line” value needs.
- Use
-
LaTeX writer:
- Avoid nested walk in table cell line-break handling.
- Don’t use
footnotehyperfor notes in longtable. Instead, generate them manually as we do for floating tables. This removes our dependency on footnotehyper, and resolves a compatibility problem withendfloat(#11857). - Remove unused
stInternalLinksstate field. The field has not been read since the writer switched to always adding hypertargets, but we were still doing a full-document query to populate it. - Avoid needless conversions in
sectionHeader. The note-free and link-free variants of the heading text were rendered for every heading, even though the former is only needed when the heading contains a note, image, or identified span, and the latter only for unnumbered, listed headings. - Fuse preprocessing passes in
inlineListToLaTeX. The strut-insertion and quote-kerning fixups were separate list traversals, allocating an intermediate list each, and this function is called at every level of inline nesting. Combine them into a single pass. - Avoid String round-trip for highlighted inline code.
- Don’t unpack code string to choose
lstinlinedelimiter. - Only compute PDF trailer ID when
SOURCE_DATE_EPOCHis set. - Use
showHexinstead of (slower)printfintoLabel.
-
RST writer:
- Don’t re-transform stored labels and alt text. Reference labels, image substitution labels, and alt text are stored after the document-wide walk in inlineListToRST has already applied transformInlines (flattening, backslash-space insertion, etc.). Rendering them with inlineListToRST applied these non-idempotent transformations a second time, inserting duplicate
\markers. With--reference-linksthis could make the inline reference and its definition render differently, producing a broken RST reference.
- Don’t re-transform stored labels and alt text. Reference labels, image substitution labels, and alt text are stored after the document-wide walk in inlineListToRST has already applied transformInlines (flattening, backslash-space insertion, etc.). Rendering them with inlineListToRST applied these non-idempotent transformations a second time, inserting duplicate
-
Native writer:
- Render directly instead of using pretty-show. Previously writeNative used pretty-show’s
ppDoc, which shows the document, tokenizes and re-parses the result into a generic Value, and lays that out via Text.PrettyPrint.HughesPJ. The layout step dominated the cost of the writer (and of any pipeline producing native output). We now build a width-cached layout tree directly from the AST and render it with a small renderer that reproduces HughesPJ’s layout algorithm exactly (including the ribbon computation with ribbonsPerLine = 1.2 and the treatment of glued closing delimiters), so the output is byte-for-byte identical to before. Verified against the old binary on all golden .native files and the markdown test corpus at many column widths, in both standalone and plain modes. The writer is about 4x faster. Also remove the pretty and pretty-show dependencies.
- Render directly instead of using pretty-show. Previously writeNative used pretty-show’s
-
XML writer:
- Respect
--standalone. When standalone is selected, we get a full Pandoc element with xml header and metadata. When not, we get a fragment – just the blocks. - Use the pretty-printer instead of manual newlines. Render the document with
ppcElement, using a configuration that treats elements with inline content as inline tags, so that no significant whitespace is added inside them. Consecutive text nodes are merged, SoftBreak is written as a literal newline, and whitespace runs that would not survive a roundtrip (e.g." \n"or"\n\n") are encoded as Space and SoftBreak elements. - Encode attribute names that are not valid XML names. Pandoc attribute names may contain characters that are not allowed in XML attribute names, such as colons (
typst:property). Use the common convention of encoding such characters as_xHHHH_, whereHHHHis the hexadecimal code of the character: the writer encodes attribute names (foo:barbecomesfoo_x003A_bar) and the reader decodes them. - Don’t drop attributes with empty values. The writer dropped any attribute with an empty value, so user key-value attributes like
("k","")disappeared and did not round trip. - Improve performance: the XML writer is now on par with the JSON writer.
- Respect
-
EPUB writer:
- Render TOC item titles in the host monad. Previously each TOC item title in the nav entry was rendered with a separate
runPure (writeHtmlStringForEPUB ...). Since every one of these invocations starts with a fresh CommonState, the translations YAML file was re-read and re-parsed for every TOC item, which accounted for a significant part of the EPUB writer’s run time on documents with many sections. - Speed up MathML/SVG detection for the manifest.
- Render TOC item titles in the host monad. Previously each TOC item title in the nav entry was rendered with a separate
-
Powerpoint writer:
- Add archive entries in a single pass.
- Cache the parsed slide master in WriterEnv.
- Skip speaker-notes walk when there are no notes.
-
ANSI writer:
- Fix missing bar on blockquote’s first line (#11804, Gaurav Vijay Jadhav D.).
-
HTML writer:
- Fix typo in
intrinsicEventsHTML4. The list hadonmouseouttwice and was missingonmousemove. - Don’t emit
<p></p>for paragraphs with no rendered content. - Improve email obfuscation. Preserve formatting (e.g. emphasis) in the link text of obfuscated mailto links. Preserve link attributes in reference- and JavaScript-obfuscated links. Avoid double-escaping the already-rendered link text in the fallback branch for unparseable mailto URLs.
- Respect
--id-prefixin EPUB3 footnote section id. - Respect incremental/nonincremental classes inside columns.
- Make KaTeX CSS URL handling consistent with the JS URL.
- Use truncate consistently for table width percentages. Previously the table width could exceed the sum of column widths.
- Avoid walking section contents when not producing slides.
- Hoist attribute set unions to top level.
- Make
strToHtmlmore efficient. Replace theT.groupBy-based implementation, which allocated a list of Text fragments and round-tripped through String, with a simpleT.breakscanner.
- Fix typo in
-
Org writer:
- Don’t render table body rows twice. Table rows were converted once to compute column widths and then again to produce the output. Since blockListToOrg is stateful, any footnote in a table cell was registered twice, yielding duplicate footnote definitions and skewed numbering; it also doubled the rendering work. Reuse the first conversion.
- Use
#+begin_export htmlfor raw HTML blocks.#+begin_htmlwas removed in Org 9.0 (2016). - Don’t treat bare punctuation as a list marker. The check that keeps ordered list markers from ending up at the beginning of a line matched
Str "."andStr ")", sinceT.all isDigit ""is True. Require at least one digit. - Don’t emit
<<>>for spans with no id. - Avoid
=delimiter for inline code containing=. Org has no escape mechanism inside verbatim text, so=code with ==did not parse as verbatim. Fall back to the equivalent~...~delimiter when the content contains=(and no~). - Escape square brackets in links. Link targets containing square brackets broke the bracket link syntax. Escape targets the way Emacs’
org-link-escapedoes: backslash-escape brackets and double backslash runs occurring before a bracket or at the end of the target. Link descriptions cannot contain escapes; instead, likeorg-link-make-string, insert a zero-width space between consecutive closing brackets and before a closing bracket at the end of the description. - Build escaped strings from chunks, not characters.
escapeStringallocated one Doc node per character for any string containing a non-alphanumeric character. Split on the (rare) special characters instead and emit intervening text as single literals. No change in output. - Emit definitions for footnotes nested in footnotes.
- Use a counter for footnote numbers. Computing the reference number as
length stNotes + 1walked the accumulated note list for every footnote, making note numbering quadratic in the number of notes.
-
Org reader and writer:
- Make code line comma-escaping match Emacs’ behavior. Org’s escaping rule for code in src/example blocks adds a comma to lines matching
^[ \t]*,*(\*|#\+), i.e. lines already starting with commas before*or#+get an additional comma; unescaping removes one comma from such lines. The writer previously left a literal,#+fooline unescaped, and the reader then stripped its comma when reading the result back, corrupting the code on round trips. Writer and reader now both handle runs of commas, matching Emacs.
- Make code line comma-escaping match Emacs’ behavior. Org’s escaping rule for code in src/example blocks adds a comma to lines matching
-
Text.Pandoc.ImageSize:
- ImageType now derives Eq [API change].
- Fix typo in bare JPEG signature detection.
- Fix size detection for lossless (VP8L) WebP.
- Make EMF parsing more robust.
- Don’t let zlib errors escape as exceptions in
pdfSize. Treat malformed streams as a parse scanner (so we keep scanning the rest for a/MediaBox). - Fix AVIF detection and parsing.
- Take bounding box origin into account for EPS. The size was computed from the upper corner alone, giving wrong dimensions for EPS files whose bounding box has a non-zero origin.
- Handle commas and fractional numbers in SVG
viewBox. - Add tests for image type and size detection (Tests.ImageSize).
- Determine JPEG size without decoding the image (which can allocate huge amounts of memory). If the header scan fails, we still fall back to the full decoder.
- Scan PDF object streams in chunks, not byte by byte.
- Speed up
findSvgTagby using a single pass. Up to 60X faster on files with few<characters. - Determine PNG size without decoding the image. If the header scan fails, we still fall back to the full decoder.
- Use
writerDpifor AVIF images instead of hardcoding 72. With the default options this changes the assumed resolution from 72 to 96 dpi. - Allow whitespace between number and unit in
numUnit. So e.g.width="3 cm"is now recognized. - Handle largesize and size-0 boxes in AVIF parser.
- Make checkDpi default to 72 for negative dpi values, not 0; a negative dpi would produce negative dimensions in
sizeInPoints.
-
Text.Pandoc.SelfContained:
- Fix inverted charset condition in
makeDataURI. - Keep semicolon in
@importfallback output. - Make
</scriptcheck case-insensitive. - Prefix all
url(#...)occurrences in SVG attributes. Previously only the first got prefixed rewritten. - Handle gzip decompression errors gracefully.
- Remove unused
isHtml5field from ConvertState. - Cache fetched resources. Previously every occurrence of a resource was fetched, decompressed, and CSS-rewritten independently, so a document referencing the same image N times triggered N network requests. Failed fetches are cached too, so each missing resource is now reported only once.
- Only add
roleandaria-labelwhen inlining SVGs. Do not add them to other elements withsrcattributes. - Escape only what is needed in textual data URIs.
- Fix inverted charset condition in
-
Text.Pandoc.UTF8:
- Avoid copying input when it contains no CRs.
toTextandtoTextLazyunconditionally ran a CR-removing filter over the input, allocating a full copy of the document even in the common case where no CRs are present. Check for a CR first (B.elem, a fast memchr) and reuse the input buffer unchanged if none is found; for the lazy variant, do this chunk-wise to preserve laziness. On a 10 MB LF-only input this makestoTextover 4x faster; when CRs are present the extra scan is not measurable. - Make
readFileexception-safe. UsewithFileinstead ofopenFileso the handle is closed even if reading throws.
- Avoid copying input when it contains no CRs.
-
Text.Pandoc.Class:
- Avoid copying input in
toTextM. Skip the CR-filtering copy when the input contains no CRs, as already done in Text.Pandoc.UTF8.toText. - Make
runSilentlyerror-safe. Previously, if the action passed to runSilently threw an error that was later caught, the verbosity remained pinned at ERROR and all previously accumulated log messages were lost. Now the original log and verbosity are restored even when the action fails. - Fix
isRelativeToParentDir. Compare the first path component rather than just looking at a prefix, to correctly handle paths like..foo/bar.yaml. - Fix base64 detection in
extractURIData. The base64 indicator in a data URI is the final parameter of the media type and may follow other parameters, e.g.charset. Previously, the code expected;base64to be the only parameter. - Fix percent-decoding in
extractURIData. - Report accurate offset in
toTextMerrors. Scan for the first invalid UTF-8 sequence and report its actual position and byte. - Reset HTTP manager in
setNoCheckCertificate. The HTTP manager is created lazily with TLS settings based onstNoCheckCertificateand then cached in CommonState, so changing the option after the first request had no effect. Discard the cached manager when the option’s value changes. - Don’t follow symlink cycles in
addToFileTree. - Finish factoring
openURLinto Text.Pandoc.Class.IO.HTTP. Commit 455bea9 added the new module but did not register it in pandoc.cabal or remove the original definitions. logOutput: avoid multiplehPutStrLn, which can cause confusing interleaving.
- Avoid copying input in
-
Text.Pandoc.MediaBag:
- Use hashlazy to avoid copying media contents.
- Treat
data:andfile:URI schemes case-insensitively. - Make MediaItem mime type and path strict.
mediaContentsis left lazy so contents need not be forced at insert time. - Use
Text.Pandoc.URI.isURIincanonicalize.Network.URI.isURItreats Windows drive-letter paths likec:/foo.pngas URIs. - Only reject
..as a path component. The insertMedia check usedisInfixOf, so a harmless name like foo..bar.png was silently renamed to its content hash. - Collapse
.and..components incanonicalize.normalisedoes not remove redundant path components, soimg/../a.pnganda.pngwere distinct keys. UsemakeCanonical(as PandocPure’s FileTree already does for its path-indexed map), which also handles duplicate and trailing slashes, replacing backslashes with slashes first. - Prevent
mediaPathcollisions between keys. The friendly mediaPath was derived by percent-unescaping the key, so distinct keys likea%20b.pnganda b.pngproduced the same mediaPath (“a b.png”) and silently clobbered each other on extraction (and inside docx/epub archives). Now the original name is only kept if the key contains no percent sign, somediaPathequals the key and distinct keys yield distinct paths; anything percent-encoded gets a content-hash name. Hashed names can only coincide for identical contents, which is harmless.
-
Text.Pandoc.XML.Light:
-
Fix escaping of repeated
]]>inCDATA. -
Make
escStrmore efficient. -
Avoid round-trips in
ppCDataSprettify path. This makesshowCDataandppcCDataunused, so they are removed. -
Parse XML fragments from the event stream instead of using xml-conduit’s document parser, which requires a single root element. Our earlier woraround with a wrapper element was fragile. Behavior changes:
- Content fragments with an XML declaration or DOCTYPE followed by multiple root elements, and text-only or empty input, now parse instead of erroring.
- Attributes now preserve document order instead of being sorted alphabetically (the document parser stored them in a Map).
- Errors for unresolved entities and mismatched tags now report source positions.
-
-
Use
text-builderpackage for rendering, instead of text’s lazy Text builder. On a large document this cuts docx conversion time by about 13%; output is byte-for-byte identical. -
Expose
ppcTopElementfrom Output. -
ConfigPP now has a field
inlineTagthat checks for inline tags. Inline tags are printed on one line and not indented, by default. ExportuseInlineTags,prettyConfigPP. -
Text.Pandoc.Sources:
- Add bulk
takeWhileP/takeWhile1Pcombinators [API change]. Character streams over Sources previously had to be consumed one character at a time viasatisfy, at a cost of several allocations and monadic binds per character. The new combinators scan a whole run of matching characters with a single parser invocation usingT.span, while replicating the exact semantics ofT.pack <$> many/many1 (satisfy f), including empty-chunk handling, position updates at chunk boundaries, and parsec’s error messages. - Use the new bulk
takeWhileP/takeWhile1Pcombinators across readers.
- Add bulk
-
Text.Pandoc.Translations:
setTranslationsnow keeps an already-loaded translation table when the language is unchanged, instead of unconditionally clearing the cache and forcing a re-read and re-parse of the translations YAML file on the nexttranslateTerm.- Term names in translation files are now parsed with a precomputed Map lookup instead of the derived Read instance, which is very slow for a 22-constructor enum.
-
Text.Pandoc.Chunks:
- Remove vestigial
nav-pathattribute andrmNavAttrswalk. This no longer did anything; output is unaffected. - Use
compactifyTablefor tables produced by all readers. Remove old ad hocparaToPlainat the table cell level. This should ensure that we don’t get tables that mix Plain and Para (#11864). Such tables tend to look funny when rendered in docx and other formats.
- Remove vestigial
-
Text.Pandoc.Data:
- Fix
getDataFileNameswith-embed_data_files. Previously it was not looking in the right directory and not recursing.
- Fix
-
Text.Pandoc.Shared:
taskListItemFromAscii: Fix incorrect treatment of[ ]as checked.- Add
stringifyInlines, a single-passstringifyfor inlines [API change]. This is about 6x faster thanstringifyfor long inline sequences - Speed up
stringifyby making it accumulate[Text]and concatenate once at the end, instead of mappending at every node. - Use
stringifyInlinesinstead ofstringifywhere possible. - Add new function
compactifyTable. [API change] This converts cells that consist in a single Para block to a Plain, provided the table contains only such cells (or empty cells). - Drop expired entries in
decrementTrailingRowSpans. Entries whose RowSpan fell to 0 were left in the map, relying on every consumer to guard against them. Delete them instead. - Recognize empty task list items in
toTaskListItem. - Make
endsWithPlainlook inside DefinitionList. endsWithPlain recursed into the last item of BulletList and OrderedList but ignored DefinitionList, so list items ending with a compact definition list were treated as loose by the RST, Org, and Haddock writers. - Avoid Text -> String round-trips in
htmlAttrs. This adds a FromText constraint tohtmlAttrsandtagWithAttrs[API change]. - Fuse traversals in
ensureValidXmlIdentifiers.
-
Text.Pandoc.Writers.Shared:
- Fix logic bug in
splitSentences. lookupMetaBool: treat empty block or inline list as False.htmlAttrs: escape id and class attributes, like the others.stripLeadingTrailingSpace- strip multiple Space, if present.toSubscript: handle minus sign.- Fix
ensureValidXmlIdentifiersfor Figure and table sub-elements, resolving a bug that produced broken internal links in the HTML4/XHTML, EPUB, DocBook, TEI, ICML, FB2, and ODT writers.
- Fix logic bug in
-
Text.Pandoc.Parsing:
- Improve performance of
uriSchemeby using a trie. Up to 15% faster in URL-heavy documents withautolink_bare_uris. - Minor code cleanup.
- Remove
$checks in math when delim is\(or\\(. - Make
anyOrderedListMarkermore efficient.
- Improve performance of
-
HTML template:
- Dark mode support (#11831, Clar Fon).
-
reveal.js template:
- Fix plugin paths (#11907).
-
flake.nix: parse allow-newer and allow-newer-deps in stack.yaml.
-
Make
embed_data_filesflag default to True. Remove flag settings from cabal.project. This makes it possible to override it on the command line. -
Depend on commonmark 0.3.1, commonmark-extensions 0.2.7.3, commonmark-pandoc 0.3.0.2 (major performance improvements).
-
Depend on released asciidoc 0.1.1 (major performance improvements).
-
Use released texmath 0.13.3 (major performance improvements).
-
Depend on released djot 0.1.4.3 (major performance improvements).
-
Use released citeproc 0.14 (major performance improvements).
-
Use released doclayout 0.6.
-
Use released zip-archive 0.5 (major performance improvements).
-
Use released skylighting-0.15 (major performance improvements).
-
Depend on released doctemplates 0.11.1.
-
Depend on released typst 0.12.
-
Require text >= 2.0.
-
Bump upper bound for unicode-data.
-
Allow crypton 2.0.x.
-
Allow Diff 2.0.
-
Add
tools/diff-golden-tests.sh. -
Fix
tools/diff-zip.shon non-Darwin. -
Add
tools/benchplot.js. This creates a nice graph comparing two benchmarks. -
typst-properties.md: fix fill syntax in Typst property examples (#11855, zenor0). -
Remove tested-with from cabal file. We tend not to keep it up to date.
-
Fix typo in Lua filter example (#11875, Andonome).
-
Fix a bug in jats-reader.xml (duplicate attribute)