This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v2.1.0 - Complete PDF Specialist Removal & Clean Architecture
🎯 Major Architecture Cleanup
This release completes the removal of PDF specialist frameworks and cleans up all related filtering logic for a simpler, more maintainable codebase.
✅ What's Changed
- Complete removal: All PDF specialist frameworks (PyMuPDF, PDFPlumber, Playa) entirely removed from codebase
- Simplified architecture: Removed unnecessary filtering logic since PDF specialists no longer exist
- Cleaner CI/CD: Streamlined workflows to only run multi-format frameworks
- Updated configurations: Fresh AI assistant rules and documentation
🔧 Current Framework Support (6 frameworks)
- Kreuzberg (sync/async + OCR variants) - 71MB installation
- Extractous - Rust-based performance - ~100MB
- Unstructured - Enterprise solution - 146MB
- MarkItDown - Microsoft converter - 251MB
- Docling - IBM Research ML - 1GB+
🚀 Technical Improvements
- Removed
filter_pdf_specialists()function - no longer needed - Simplified
reporting.pyinitialization logic - Cleaner
generate_index.pywithout filtering calls - Updated GitHub Actions workflows to exclude PDF specialists entirely
📊 Fresh Benchmark Results Coming
This release will trigger a complete benchmark run generating:
- Clean visualization charts without PDF specialists
- Updated performance rankings for multi-format frameworks only
- Fresh GitHub Pages deployment with corrected data
🎯 Impact
- Fairer comparisons: Only multi-format frameworks with similar scope
- Cleaner results: No more mixed single-format vs multi-format data
- Simpler maintenance: Less code, fewer edge cases
- Better insights: Focus on frameworks users actually compare
Breaking Change: PDF specialist frameworks are completely removed. Use previous releases if you need PDF-only framework benchmarks.