This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v1.10.0: Table Extraction Benchmarking
🔢 Table Extraction Benchmarking
This release adds comprehensive table extraction benchmarking capabilities with separate workflows to avoid contaminating general text extraction benchmarks.
✨ New Features
- Separate Table Extraction Workflow: Isolated benchmarking with specialized timeouts (2-8 hours)
- GMFT Support for Kreuzberg: Advanced table extraction with
kreuzberg[gmft] - Table Analysis Module: Structure preservation and content detection scoring
- Table-Specific CLI:
--table-extraction-onlyflag for targeted benchmarking - Extended Test Documents: HTML tables with simple and complex structures
🚀 Framework-Specific Optimizations
- Kreuzberg: GMFT integration for advanced table processing (8h timeout)
- Docling: Built-in ML table extraction (6h timeout)
- Unstructured: Hi-res strategy for table detection (5h timeout)
- MarkItDown: Basic table conversion (3h timeout)
- Extractous: Fast but limited table support (2h timeout)
📊 Analysis Capabilities
- Table structure preservation scoring
- Content detection and accuracy metrics
- Framework performance rankings
- Format support matrix analysis
- Comprehensive reporting (CSV, Markdown, JSON)
🔧 Usage
# Run table extraction benchmarks
uv run python -m src.cli benchmark --table-extraction-only --framework all
# Analyze table extraction results
uv run python -m src.cli table-analysis --results-dir resultsBoth general text extraction and table extraction workflows will trigger automatically on this release.