This repository was archived by the owner on Sep 28, 2025. It is now read-only.
Repository navigation
v1.8.0 - Corrected Kreuzberg Format Support & Deployment Fixes
π What's New
π Corrected Kreuzberg Format Support
- Fixed format exclusions: Kreuzberg now correctly supports DOCX, ODT, RST, and ORG formats
- Accurate testing: Kreuzberg will be tested on 18/20 formats instead of 14/20
- Better comparisons: More fair framework comparisons with correct format support
π§ Deployment Pipeline Fixes
- Resolved circular imports: Fixed Python module path conflicts in GitHub Actions
- Dynamic HTML generation: Framework versions now display correctly on website
- Stable deployments: GitHub Pages deployment pipeline working reliably
π Enhanced Visualizations
- All 6 frameworks: Every chart now includes all frameworks (Kreuzberg sync/async, Docling, MarkItDown, Unstructured, Extractous)
- Kreuzberg branding: Added logo and favicon to benchmark website
- Correct success rates: Fixed success rate calculations based on actual tested files
π οΈ Technical Improvements
- Fixed framework version extraction from pyproject.toml
- Improved error handling in deployment pipeline
- Updated documentation to reflect actual format support
π Expected Impact
With corrected format support, Kreuzberg's success rates should improve significantly in the next benchmark run, providing more accurate performance comparisons.
π Live Results
View the latest benchmark results at: https://goldziher.github.io/python-text-extraction-libs-benchmarks/
Full Changelog: v1.7.0...v1.8.0