Skip to content

bioextract 0.1.0

Choose a tag to compare

@FuqingZh FuqingZh released this 05 Aug 05:08
· 39 commits to main since this release
e61b676

Highlights

bioextract 0.1.0 is the first release of the converged domain-access contract for local biological-database snapshots.

  • Added stable top-level database handles for ChEBI, Rhea, GO, KEGG, Reactome, WikiPathways, eggNOG, InterPro, UniProt, STRING, and OmniPath.
  • Added query-ready domain access for ChEBI, Rhea, KEGG metabolic data, UniProtKB reviewed entries, GO ontologies, Reactome, WikiPathways, InterPro/Pfam, and publication inspection.
  • Added native read-only DuckDB access through XDatabase.from_duckdb(...).connect() for unrestricted SQL over validated publications.
  • Added inspect_publication() for lightweight provenance and schema inspection without selecting a resource-specific reader.

Breaking changes

This release intentionally converges the pre-1.0 API and is not drop-in compatible with the 0.0.x helper API.

  • Use complete *Database names and the top-level import form, for example from bioextract import RheaDatabase.
  • Replace abbreviated *Db handles, resource limits, write_tidy()/directory writers, canonical Parquet writers, and legacy score-filter names with the 0.1.0 database and DuckDB APIs.
  • Use write_duckdb(destination) for materialized publications and XDatabase.from_duckdb(path) to reopen them.
  • GO no longer emits the synthetic subcell projection; query the ontology's declared cellular-component terms instead.
  • Selection and result implementation types are no longer stable resource-package exports; use database methods and the documented output schemas.

Storage and provenance

  • Materialized bioextract products use one DuckDB publication per logical product, whether the publication contains one relation or many.
  • Official/native representations remain directly consumable where they are already suitable, including eggNOG SQLite and the configured STRING/OmniPath sources.
  • Publication provenance is embedded in the DuckDB _bioextract schema; it is not supplied by a sidecar manifest.
  • Publication writes use validation, staging, bounded execution, and atomic replacement. Official two-dimensional source headers are retained when possible; derived fields use stable snake_case names and column lineage is recorded when mapping is required.

Installation

pip install bioextract==0.1.0

Python 3.13 or newer is required. Runtime dependencies include Polars, DuckDB, and PyArrow.

Verification

The release tag workflow completed successfully on the exact 0.1.0 tag, including dependency synchronization, the complete repository check, distribution build, twine check, and PyPI Trusted Publishing.

Full changelog: 0.0.8...0.1.0