Skip to content

v1.0 – Chunking Visualizer & Unstructured Playground

Choose a tag to compare

@thomast8 thomast8 released this 17 Nov 09:38
· 274 commits to main since this release
61cb139

v1.0 – Chunking Visualizer & Unstructured Playground

This first stable release ships a complete local playground plus a web-based visualizer for experimenting with PDF ingestion using the open-source Unstructured stack.

Highlights

  • Chunking Visualizer (FastAPI + static UI)

    • Interactive UI to inspect PDFs, table matching, and chunker performance.
    • Metrics view with run config recap, overall metrics, and per-table cards (coverage, cohesion, F1, selected chunk count) with “Highlight all/best” and “Details”.
    • Inspect view with separate Chunks and Elements tools, overlay toggles, type filters, and drilldowns (HTML table preview, chunk lists, element cards).
    • Stable overlay colours per element type and per-row table boxes so multi-chunk highlights don’t stack.
  • Chunking / QA scripts

    • process_unstructured.py for interactive full-document runs.
    • scripts/preview_unstructured_pages.py for focused page-slice runs, gold-table matching against dataset/gold.jsonl, and detailed per-table + overall metrics.
  • Deployment & infra

    • Dockerfile tuned for faster rebuilds: system deps + Python deps are installed via uv using pyproject.toml and uv.lock before copying the app, with a simple uv run python main.py start command.
    • Railway-ready: the app binds to HOST/PORT env vars and supports volume-backed output directories via DATA_DIR / OUT_DIR / OUTPUT_DIR and PDF_DIR.

How to run locally

uv sync
uv run uvicorn main:app --host 127.0.0.1 --port 8765
# then open http://127.0.0.1:8765/
  • PDFs default to res/.
  • Outputs default to outputs/unstructured/ unless you set:
    • DATA_DIR=/some/path (then outputs go under $DATA_DIR/outputs/unstructured and PDFs under $DATA_DIR/pdfs), or
    • OUT_DIR / OUTPUT_DIR and PDF_DIR explicitly.

For targeted script-based QA, use for example:

uv run python scripts/preview_unstructured_pages.py \
  --input res/V3.0_Reviewed_translation_EN_full\ 4.pdf \
  --pages 4-6 \
  --only-tables \
  --output outputs/unstructured/V3_0_EN_4.pages4-6.tables.jsonl \
  --gold dataset/gold.jsonl \
  --emit-matches outputs/unstructured/V3_0_EN_4.matches.json