Skip to content

v0.10.0

Latest

Choose a tag to compare

@rsgoncalves rsgoncalves released this 13 Sep 19:28
· 1 commit to main since this release

Change Log

User-facing changes:

  • Add a column DiseaseLocation to the table efo_labels containing anatomical locations associated with each EFO term through the property EFO:0000784 (has_disease_location); if any exist.
    • When a term does not have such a relation, the code tries to find location relations inherited from the most immediate parent(s) in the class hierarchy.
    • Populate DiseaseLocation column with NA when no locations exist.
  • Add a table called version_info to document details such as versions of ontologies & the generated DB and the download date of the OpenGWAS metadata used to build the database.
  • Add functionality to include additional, user-provided tables in the generated database.
  • Add functionality to include additional, user-provided ontologies in the generated database.
  • Output a compressed tar.xz archive instead of a .db file.
  • Add UBERON ontology tables to the database.
  • Update to latest EFO v3.57.0.

Developer-facing changes:

  • Fix issue where ontology mappings were being undercounted, because the mapping identifiers we generated were colliding in the event of multiple mappings per trait. Switched to generating random UUIDs for mapping identifiers to avoid such collisions.
  • Fix issue with mismatching CURIEs that resulted in mapping count discrepancies.
    • the CURIEs generated by bioregistry were different than those in the SemanticSQL database.
    • updated the bioregistry package version (and some other dependencies) to a more current version that produces CURIEs as expected to match properly.
  • Use the SemanticSQL .gz archives instead of the .db files.
    • In the case of EFO, the .db version of EFO was much older than the one in the .gz archive.
  • Save efo_labels table with the mapping counts to disk.
  • Update query_database.py to work directly with the tar file.
  • Updated some dependencies: text2term, pandas, and bioregistry.