Skip to content
This repository was archived by the owner on Sep 28, 2025. It is now read-only.

v1.5.0: Simplified 2-Tier Format Support

Choose a tag to compare

@Goldziher Goldziher released this 11 Jul 13:03
· 130 commits to main since this release
v1.5.0
dfea640

๐ŸŽ‰ Release v1.5.0

๐ŸŽฏ Simplified Format Support System

This release simplifies the format support system from 3 tiers to 2 tiers, making benchmarking choices clearer and more focused on document extraction use cases.

Format Tiers (Simplified)

Tier 1: Universal Support (5/5 frameworks) โœ…

  • 7 formats: PDF, PPTX, XLSX, PNG, BMP, HTML, CSV
  • 100% framework support for fair performance comparison
  • Use: --format-tier universal

Tier 2: Common Support (4/5 frameworks) ๐ŸŒŸ RECOMMENDED

  • 11 formats: Tier 1 + XLS, MD, JPEG, TXT
  • 80%+ framework support for real-world scenarios
  • Use: --format-tier common (default)

All Formats ๐Ÿ“Š

  • 18 formats: Complete capability assessment
  • Use: --format-tier all

๐Ÿ”„ What Changed

Removed Tier 3 (Partial Support):

  • Eliminated formats with only 3/5 framework support
  • Removed: JPG variant, EML, MSG, JSON, YAML
  • These specialized formats (email, data) don't align with document extraction benchmarking

๐Ÿ’ก Why This Change?

  1. Simpler Choice: Only decide between fair comparison vs real-world testing
  2. Better Focus: Concentrates on actual document formats
  3. Clearer Purpose: Document extraction benchmarks shouldn't include email/data formats
  4. User-Friendly: Reduces confusion about which tier to choose

๐Ÿ“Š Expected Success Rates

  • Tier 1: All frameworks achieve ~100% success
  • Tier 2: Most frameworks achieve 85-95% success
  • All: Shows full capabilities including edge cases

๐Ÿš€ Usage

# Fair performance comparison
uv run python -m src.cli benchmark --framework all --format-tier universal

# Real-world benchmarking (default and recommended)
uv run python -m src.cli benchmark --framework all --format-tier common

# Complete framework evaluation
uv run python -m src.cli benchmark --framework all --format-tier all

๐Ÿ”ง Technical Details

  • Updated config.py to remove PARTIAL_FORMATS
  • CLI now accepts only: universal, common, all
  • CI workflow updated to reflect 2-tier system
  • Documentation simplified throughout

This release makes the benchmarking suite more intuitive while maintaining its effectiveness for comparing text extraction frameworks on relevant document formats.