Skip to content

0.27.16

Latest

Choose a tag to compare

@aballman aballman released this 05 Oct 22:34
· 7 commits to main since this release
0242693

This release rolls up all changes since 0.27.10; versions 0.27.11–0.27.15 were not published to PyPI individually.

Fixes

  • Reject images whose frames decode to too many pixels instead of exhausting memory. partition_image() measures an image before partitioning and raises UnprocessableEntityError above IMAGE_MAX_TOTAL_PIXELS (default 500,000,000); with hi_res, every frame is charged. (0.27.16 — 0242693, #4520)
  • Reject spreadsheets whose worksheets span too many cells instead of exhausting memory. partition_xlsx() measures each worksheet's span by streaming the file before reading it, and raises UnprocessableEntityError above XLSX_MAX_CELLS (default 5,000,000). Subtable detection builds its graph from populated cells only. (0.27.12 — 74e2fea, #4515)
  • Reject CSV and TSV files that span too many cells instead of exhausting memory. partition_csv() / partition_tsv() measure the file's span by streaming it before Pandas reads it, and raise UnprocessableEntityError above CSV_MAX_CELLS (default 5,000,000); a lone \r line ending is normalized to \n. (0.27.15 — 15c6302, #4518)
  • Bound the work DOCX tables can demand through their declared layout grid. Grid size is computed from w:gridBefore / w:gridAfter / w:gridSpan without expansion; past DOCX_TABLE_MAX_CELLS (default 5,000,000) text_as_html is omitted with a warning while text is still extracted. (0.27.14 — 25f405c, #4519)
  • Keep the labels of auto-numbered DOCX lists. partition_docx() prefixes computed list labels (1., a), iv.) to ListItem text, resolved per level in document order; bullets stay unprefixed. (0.27.11 — d9f976e, #4508)

Maintenance

  • Bound CI dependency downloads with inactivity timeouts and bounded retries. (0.27.13 — 727db66, #4514)

Full changelog: https://github.com/Unstructured-IO/unstructured/blob/0.27.16/CHANGELOG.md