Skip to content

v1.0.1 - GROBID Integration and Improved Acknowledgements Extraction

Latest

Choose a tag to compare

@javisiierra javisiierra released this 28 May 10:35
· 1 commit to main since this release

New Features

This release introduces a major improvement to metadata and acknowledgements extraction from research papers:

GROBID integration for extracting structured metadata directly from PDFs:

  • Title
  • Authors
  • Abstract
  • Accurate acknowledgements extraction using PDFMiner, which performs better in capturing free-form text sections than GROBID for this use case.

Named Entity Recognition (NER) using the dslim/bert-base-NER model to identify:

  • Persons
  • Organizations mentioned in the acknowledgements section.

Technical Changes

  • Refactored src/preprocess.py to combine GROBID and PDFMiner:

    • GROBID is used for structured fields (title, authors, abstract).
      
    • PDFMiner is used for extracting the acknowledgements section as raw text.
      
  • The ner_ack.py script remains compatible and processes the improved acknowledgements content to extract named entities.

  • Output is saved in a structured format, ready for knowledge graph construction and enrichment with external data sources (e.g., Wikidata, ROR).