Skip to content

v0.1.0 — Experimental PDF Extraction Evaluation

Pre-release
Pre-release

Choose a tag to compare

@MGsBueno MGsBueno released this 08 Apr 12:34
· 36 commits to main since this release
7598132

First experimental release of PDF Batch Extractor, focused on evaluating PDF parsing technologies and establishing the initial project direction.

Purpose

This version explores different Python libraries for PDF content extraction in order to determine the most suitable foundation for the project.


Libraries Evaluated

  • pdfminer.six
  • pdfplumber
  • PyMuPDF (MuPDF)

Evaluation Criteria

The analysis focused on:

  • Ease of use
  • Text extraction quality
  • Support for images and tables (exploratory)
  • Overall performance

Outcome

This exploratory phase guided the decision to adopt a block-based extraction approach and laid the groundwork for future pipeline development.


Notes

This version represents an early-stage experimental baseline and is not intended for production use. Subsequent releases introduce structured extraction, configuration, and architectural improvements.