This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v1.2.1: Fix Extractous API Usage
π Critical Bug Fix
Extractous API Correction
- Fixed API Usage: Corrected
extract_file_to_string()call to handle tuple return value - Proper Text Extraction: Now extracts full text content instead of just 2 characters
- Performance Validation: Confirmed Extractous working correctly with proper extraction
π Performance Update
Local Benchmark Results
- Text File: Kreuzberg 0.001s vs Extractous 0.046s
- PDF File: Kreuzberg 0.033s vs Extractous 0.123s
- Kreuzberg Speed: 3.8x - 50x faster than Extractous on test files
Key Findings
- Kreuzberg uses pypdfium2 (Google's PDFium library) - highly optimized
- Extractous 18x speed claims may be context-dependent or for different document types
- Both frameworks now working correctly for comprehensive benchmarking
π What's Fixed
- β Extractous properly extracts full text content
- β API compatibility issues resolved
- β Ready for comprehensive benchmarking
- β Performance baselines established
π¬ Next Steps
This release triggers a new comprehensive benchmark run with the corrected Extractous implementation. Results will show accurate performance comparisons across all frameworks and document types.
The benchmark pipeline will now provide reliable data comparing:
- Kreuzberg (pypdfium2-based)
- Extractous (Rust-based)
- Unstructured, MarkItDown, Docling
π¦ Testing
All frameworks validated with proper text extraction and performance measurement.