Repository navigation
This release rolls up all changes since 0.27.10; versions 0.27.11–0.27.15 were not published to PyPI individually.
Fixes
- Reject images whose frames decode to too many pixels instead of exhausting memory.
partition_image()measures an image before partitioning and raisesUnprocessableEntityErroraboveIMAGE_MAX_TOTAL_PIXELS(default 500,000,000); withhi_res, every frame is charged. (0.27.16 — 0242693, #4520) - Reject spreadsheets whose worksheets span too many cells instead of exhausting memory.
partition_xlsx()measures each worksheet's span by streaming the file before reading it, and raisesUnprocessableEntityErroraboveXLSX_MAX_CELLS(default 5,000,000). Subtable detection builds its graph from populated cells only. (0.27.12 — 74e2fea, #4515) - Reject CSV and TSV files that span too many cells instead of exhausting memory.
partition_csv()/partition_tsv()measure the file's span by streaming it before Pandas reads it, and raiseUnprocessableEntityErroraboveCSV_MAX_CELLS(default 5,000,000); a lone\rline ending is normalized to\n. (0.27.15 — 15c6302, #4518) - Bound the work DOCX tables can demand through their declared layout grid. Grid size is computed from
w:gridBefore/w:gridAfter/w:gridSpanwithout expansion; pastDOCX_TABLE_MAX_CELLS(default 5,000,000)text_as_htmlis omitted with a warning while text is still extracted. (0.27.14 — 25f405c, #4519) - Keep the labels of auto-numbered DOCX lists.
partition_docx()prefixes computed list labels (1.,a),iv.) toListItemtext, resolved per level in document order; bullets stay unprefixed. (0.27.11 — d9f976e, #4508)
Maintenance
- Bound CI dependency downloads with inactivity timeouts and bounded retries. (0.27.13 — 727db66, #4514)
Full changelog: https://github.com/Unstructured-IO/unstructured/blob/0.27.16/CHANGELOG.md