Skip to content
This repository was archived by the owner on Sep 28, 2025. It is now read-only.

v1.2.1: Fix Extractous API Usage

Choose a tag to compare

@Goldziher Goldziher released this 11 Jul 08:04
· 137 commits to main since this release
5c6e52e

πŸ› Critical Bug Fix

Extractous API Correction

  • Fixed API Usage: Corrected extract_file_to_string() call to handle tuple return value
  • Proper Text Extraction: Now extracts full text content instead of just 2 characters
  • Performance Validation: Confirmed Extractous working correctly with proper extraction

πŸ“Š Performance Update

Local Benchmark Results

  • Text File: Kreuzberg 0.001s vs Extractous 0.046s
  • PDF File: Kreuzberg 0.033s vs Extractous 0.123s
  • Kreuzberg Speed: 3.8x - 50x faster than Extractous on test files

Key Findings

  • Kreuzberg uses pypdfium2 (Google's PDFium library) - highly optimized
  • Extractous 18x speed claims may be context-dependent or for different document types
  • Both frameworks now working correctly for comprehensive benchmarking

πŸš€ What's Fixed

  • βœ… Extractous properly extracts full text content
  • βœ… API compatibility issues resolved
  • βœ… Ready for comprehensive benchmarking
  • βœ… Performance baselines established

πŸ”¬ Next Steps

This release triggers a new comprehensive benchmark run with the corrected Extractous implementation. Results will show accurate performance comparisons across all frameworks and document types.

The benchmark pipeline will now provide reliable data comparing:

  • Kreuzberg (pypdfium2-based)
  • Extractous (Rust-based)
  • Unstructured, MarkItDown, Docling

🚦 Testing

All frameworks validated with proper text extraction and performance measurement.