Skip to content
This repository was archived by the owner on Sep 28, 2025. It is now read-only.

v1.10.0: Table Extraction Benchmarking

Choose a tag to compare

@Goldziher Goldziher released this 11 Jul 22:29
· 104 commits to main since this release
v1.10.0
2b37330

🔢 Table Extraction Benchmarking

This release adds comprehensive table extraction benchmarking capabilities with separate workflows to avoid contaminating general text extraction benchmarks.

✨ New Features

  • Separate Table Extraction Workflow: Isolated benchmarking with specialized timeouts (2-8 hours)
  • GMFT Support for Kreuzberg: Advanced table extraction with kreuzberg[gmft]
  • Table Analysis Module: Structure preservation and content detection scoring
  • Table-Specific CLI: --table-extraction-only flag for targeted benchmarking
  • Extended Test Documents: HTML tables with simple and complex structures

🚀 Framework-Specific Optimizations

  • Kreuzberg: GMFT integration for advanced table processing (8h timeout)
  • Docling: Built-in ML table extraction (6h timeout)
  • Unstructured: Hi-res strategy for table detection (5h timeout)
  • MarkItDown: Basic table conversion (3h timeout)
  • Extractous: Fast but limited table support (2h timeout)

📊 Analysis Capabilities

  • Table structure preservation scoring
  • Content detection and accuracy metrics
  • Framework performance rankings
  • Format support matrix analysis
  • Comprehensive reporting (CSV, Markdown, JSON)

🔧 Usage

# Run table extraction benchmarks
uv run python -m src.cli benchmark --table-extraction-only --framework all

# Analyze table extraction results  
uv run python -m src.cli table-analysis --results-dir results

Both general text extraction and table extraction workflows will trigger automatically on this release.