Repository navigation
Replies: 1 comment
|
This has been added :) Exposed in the CLI + all bindings https://developers.llamaindex.ai/liteparse/guides/extraction/#layout-blocks |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Python users currently receive a flat
List[TextItem]per page with no semantic structure. Themarkdown_layoutpipeline already classifies every projected line into semantic Block objects (Heading, Table, Figure, Paragraph, ListItem, CodeBlock, HorizontalRule, GridFallback) to produce the markdown output, but those Block objects are discarded after rendering.ProjectedLine.spans: Vec<TextItem>carries the source items for each line, but is also discarded.Could you surface this classification per blocks to Python as ParsedPage.document_blocks: List[DocumentBlock], from which users can build a structured document with labeled, item-grouped blocks ?
All reactions