v3.3.0 Add extraction as MarkDown, include text from included XObjects
·
19 commits
to main
since this release
Immutable
release. Only release title and notes can be modified.
What's Changed
- Apply the font /Differences array to every byte of a text segment by @splitbrain in #453
- Normalize white-space and odd digit counts in hexadecimal strings by @splitbrain in #452
- Switch to static license badge to fix pool issues by @PrinsFrank in #473
- Rename "content" to "text" in samples to prepare for markdown extraction by @PrinsFrank in #474
- Don't parse content of inline images by @PrinsFrank in #476
- Add simple markdown document sample by @PrinsFrank in #478
- Add extracted markdown to samples by @PrinsFrank in #479
- Move text extraction code out of document parsing namespaces by @PrinsFrank in #480
- Clean up line grouping strategies by @PrinsFrank in #481
- Clean up page/document arguments as page already has access to the document by @PrinsFrank in #482
- Add sample from issue #457 by @PrinsFrank in #483
- Bump actions/checkout from 7.0.0 to 7.0.1 by @dependabot[bot] in #485
- Bump zizmorcore/zizmor-action from 0.5.7 to 0.6.0 by @dependabot[bot] in #486
- Add markdown extraction methods by @PrinsFrank in #488
- Update phpstan/phpstan requirement from 2.2.5 to 2.2.6 by @dependabot[bot] in #489
- Bump zizmorcore/zizmor-action from 0.6.0 to 0.6.1 by @dependabot[bot] in #490
- Add sample with content in xobjects by @PrinsFrank in #492
- Extract text from xObjects by @PrinsFrank in #491
- Bump zizmorcore/zizmor-action from 0.6.1 to 0.6.2 by @dependabot[bot] in #493
- Include benchmark stats as text in README by @PrinsFrank in #495
- Update comparison table with more benchmark information by @PrinsFrank in #496
- Exit recursion in xObject includes by @PrinsFrank in #498
- Fix profiling script to work without bootstrapping composer by @PrinsFrank in #499
- Track object numbers in decorated objects to detect recursion by @PrinsFrank in #500
- Use correct xObject dictionary when in xObject context by @PrinsFrank in #501
- Flush inMemoryStream to string in ContentStreamParser to reduce method call overhead by @PrinsFrank in #502
- Add operator char hashmap to reduce getOperator calls by @PrinsFrank in #503
- Don't parse contentstreams in textObjects that don't contain text when extracting text only by @PrinsFrank in #504
- Remove extra round of dictionary parsing that was then discarded by @PrinsFrank in #505
- Remove rollingCharBuffer from dictionaryParser and only keep track of two last characters by @PrinsFrank in #506
- Optimize content stream parsing by @PrinsFrank in #507
- Optimize dictionary parsing by @PrinsFrank in #508
- Remove usage of infiniteBuffer by @PrinsFrank in #509
- Remove unnecessary clones on readonly classes by @PrinsFrank in #510
- Extract bold/italic markdown information by @PrinsFrank in #511
- Extract headings in markdown by @PrinsFrank in #512
- Document supported markdown features by @PrinsFrank in #513
Full Changelog: v3.2.0...v3.3.0