Skip to content
This repository was archived by the owner on Sep 28, 2025. It is now read-only.

v1.8.0 - Corrected Kreuzberg Format Support & Deployment Fixes

Choose a tag to compare

@Goldziher Goldziher released this 11 Jul 17:46
· 113 commits to main since this release
ba28b09

πŸš€ What's New

πŸ“‹ Corrected Kreuzberg Format Support

  • Fixed format exclusions: Kreuzberg now correctly supports DOCX, ODT, RST, and ORG formats
  • Accurate testing: Kreuzberg will be tested on 18/20 formats instead of 14/20
  • Better comparisons: More fair framework comparisons with correct format support

πŸ”§ Deployment Pipeline Fixes

  • Resolved circular imports: Fixed Python module path conflicts in GitHub Actions
  • Dynamic HTML generation: Framework versions now display correctly on website
  • Stable deployments: GitHub Pages deployment pipeline working reliably

πŸ“Š Enhanced Visualizations

  • All 6 frameworks: Every chart now includes all frameworks (Kreuzberg sync/async, Docling, MarkItDown, Unstructured, Extractous)
  • Kreuzberg branding: Added logo and favicon to benchmark website
  • Correct success rates: Fixed success rate calculations based on actual tested files

πŸ› οΈ Technical Improvements

  • Fixed framework version extraction from pyproject.toml
  • Improved error handling in deployment pipeline
  • Updated documentation to reflect actual format support

πŸ“ˆ Expected Impact

With corrected format support, Kreuzberg's success rates should improve significantly in the next benchmark run, providing more accurate performance comparisons.

🌐 Live Results

View the latest benchmark results at: https://goldziher.github.io/python-text-extraction-libs-benchmarks/

Full Changelog: v1.7.0...v1.8.0