Skip to content
This repository was archived by the owner on Sep 28, 2025. It is now read-only.

v0.1.0 - Initial Release

Choose a tag to compare

@Goldziher Goldziher released this 03 Jul 17:29
· 176 commits to main since this release
4d6587c

🎉 Initial Release

Features

  • Comprehensive benchmarking suite for Python text extraction libraries
  • Support for 5 frameworks: Kreuzberg (sync/async), Docling, MarkItDown, Unstructured
  • 94 test documents across multiple formats and languages
  • Detailed performance profiling (CPU, memory, I/O)
  • Framework isolation in CI/CD for reliable results
  • Automatic aggregation and reporting with partial failure handling
  • Multiple report formats: Markdown, HTML, JSON

Tested Frameworks

  • Kreuzberg v3.3.0+ (sync & async)
  • Docling v2.15.0+ (with performance warnings)
  • MarkItDown v0.0.1a2+
  • Unstructured v0.16.11+

CI/CD

  • Automated benchmarks on release (2-hour timeout per framework)
  • Manual workflow dispatch for testing
  • Graceful handling of timeouts and failures

Known Issues

  • Docling has significant performance issues on medium-sized PDFs (40+ seconds per MB)
  • Some frameworks may timeout on medium documents with 2-hour limit

This release triggers the first automated benchmark run!

Full Changelog: https://github.com/Goldziher/python-text-extraction-libs-benchmarks/commits/v0.1.0