New Features
This release introduces a major improvement to metadata and acknowledgements extraction from research papers:
GROBID integration for extracting structured metadata directly from PDFs:
- Title
- Authors
- Abstract
- Accurate acknowledgements extraction using PDFMiner, which performs better in capturing free-form text sections than GROBID for this use case.
Named Entity Recognition (NER) using the dslim/bert-base-NER model to identify:
- Persons
- Organizations mentioned in the acknowledgements section.
Technical Changes
-
Refactored src/preprocess.py to combine GROBID and PDFMiner:
-
GROBID is used for structured fields (title, authors, abstract). -
PDFMiner is used for extracting the acknowledgements section as raw text.
-
-
The ner_ack.py script remains compatible and processes the improved acknowledgements content to extract named entities.
-
Output is saved in a structured format, ready for knowledge graph construction and enrichment with external data sources (e.g., Wikidata, ROR).