This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v0.1.0 - Initial Release
🎉 Initial Release
Features
- Comprehensive benchmarking suite for Python text extraction libraries
- Support for 5 frameworks: Kreuzberg (sync/async), Docling, MarkItDown, Unstructured
- 94 test documents across multiple formats and languages
- Detailed performance profiling (CPU, memory, I/O)
- Framework isolation in CI/CD for reliable results
- Automatic aggregation and reporting with partial failure handling
- Multiple report formats: Markdown, HTML, JSON
Tested Frameworks
- Kreuzberg v3.3.0+ (sync & async)
- Docling v2.15.0+ (with performance warnings)
- MarkItDown v0.0.1a2+
- Unstructured v0.16.11+
CI/CD
- Automated benchmarks on release (2-hour timeout per framework)
- Manual workflow dispatch for testing
- Graceful handling of timeouts and failures
Known Issues
- Docling has significant performance issues on medium-sized PDFs (40+ seconds per MB)
- Some frameworks may timeout on medium documents with 2-hour limit
This release triggers the first automated benchmark run!
Full Changelog: https://github.com/Goldziher/python-text-extraction-libs-benchmarks/commits/v0.1.0