v0.1.0 — Experimental PDF Extraction Evaluation
Pre-release
Pre-release
First experimental release of PDF Batch Extractor, focused on evaluating PDF parsing technologies and establishing the initial project direction.
Purpose
This version explores different Python libraries for PDF content extraction in order to determine the most suitable foundation for the project.
Libraries Evaluated
pdfminer.sixpdfplumberPyMuPDF(MuPDF)
Evaluation Criteria
The analysis focused on:
- Ease of use
- Text extraction quality
- Support for images and tables (exploratory)
- Overall performance
Outcome
This exploratory phase guided the decision to adopt a block-based extraction approach and laid the groundwork for future pipeline development.
Notes
This version represents an early-stage experimental baseline and is not intended for production use. Subsequent releases introduce structured extraction, configuration, and architectural improvements.