This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v1.5.0: Simplified 2-Tier Format Support
๐ Release v1.5.0
๐ฏ Simplified Format Support System
This release simplifies the format support system from 3 tiers to 2 tiers, making benchmarking choices clearer and more focused on document extraction use cases.
Format Tiers (Simplified)
Tier 1: Universal Support (5/5 frameworks) โ
- 7 formats: PDF, PPTX, XLSX, PNG, BMP, HTML, CSV
- 100% framework support for fair performance comparison
- Use:
--format-tier universal
Tier 2: Common Support (4/5 frameworks) ๐ RECOMMENDED
- 11 formats: Tier 1 + XLS, MD, JPEG, TXT
- 80%+ framework support for real-world scenarios
- Use:
--format-tier common(default)
All Formats ๐
- 18 formats: Complete capability assessment
- Use:
--format-tier all
๐ What Changed
Removed Tier 3 (Partial Support):
- Eliminated formats with only 3/5 framework support
- Removed: JPG variant, EML, MSG, JSON, YAML
- These specialized formats (email, data) don't align with document extraction benchmarking
๐ก Why This Change?
- Simpler Choice: Only decide between fair comparison vs real-world testing
- Better Focus: Concentrates on actual document formats
- Clearer Purpose: Document extraction benchmarks shouldn't include email/data formats
- User-Friendly: Reduces confusion about which tier to choose
๐ Expected Success Rates
- Tier 1: All frameworks achieve ~100% success
- Tier 2: Most frameworks achieve 85-95% success
- All: Shows full capabilities including edge cases
๐ Usage
# Fair performance comparison
uv run python -m src.cli benchmark --framework all --format-tier universal
# Real-world benchmarking (default and recommended)
uv run python -m src.cli benchmark --framework all --format-tier common
# Complete framework evaluation
uv run python -m src.cli benchmark --framework all --format-tier all๐ง Technical Details
- Updated
config.pyto removePARTIAL_FORMATS - CLI now accepts only: universal, common, all
- CI workflow updated to reflect 2-tier system
- Documentation simplified throughout
This release makes the benchmarking suite more intuitive while maintaining its effectiveness for comparing text extraction frameworks on relevant document formats.