bioextract 0.1.0
Highlights
bioextract 0.1.0 is the first release of the converged domain-access contract for local biological-database snapshots.
- Added stable top-level database handles for ChEBI, Rhea, GO, KEGG, Reactome, WikiPathways, eggNOG, InterPro, UniProt, STRING, and OmniPath.
- Added query-ready domain access for ChEBI, Rhea, KEGG metabolic data, UniProtKB reviewed entries, GO ontologies, Reactome, WikiPathways, InterPro/Pfam, and publication inspection.
- Added native read-only DuckDB access through
XDatabase.from_duckdb(...).connect()for unrestricted SQL over validated publications. - Added
inspect_publication()for lightweight provenance and schema inspection without selecting a resource-specific reader.
Breaking changes
This release intentionally converges the pre-1.0 API and is not drop-in compatible with the 0.0.x helper API.
- Use complete
*Databasenames and the top-level import form, for examplefrom bioextract import RheaDatabase. - Replace abbreviated
*Dbhandles, resource limits,write_tidy()/directory writers, canonical Parquet writers, and legacy score-filter names with the 0.1.0 database and DuckDB APIs. - Use
write_duckdb(destination)for materialized publications andXDatabase.from_duckdb(path)to reopen them. - GO no longer emits the synthetic subcell projection; query the ontology's declared cellular-component terms instead.
- Selection and result implementation types are no longer stable resource-package exports; use database methods and the documented output schemas.
Storage and provenance
- Materialized bioextract products use one DuckDB publication per logical product, whether the publication contains one relation or many.
- Official/native representations remain directly consumable where they are already suitable, including eggNOG SQLite and the configured STRING/OmniPath sources.
- Publication provenance is embedded in the DuckDB
_bioextractschema; it is not supplied by a sidecar manifest. - Publication writes use validation, staging, bounded execution, and atomic replacement. Official two-dimensional source headers are retained when possible; derived fields use stable
snake_casenames and column lineage is recorded when mapping is required.
Installation
pip install bioextract==0.1.0Python 3.13 or newer is required. Runtime dependencies include Polars, DuckDB, and PyArrow.
Verification
The release tag workflow completed successfully on the exact 0.1.0 tag, including dependency synchronization, the complete repository check, distribution build, twine check, and PyPI Trusted Publishing.
Full changelog: 0.0.8...0.1.0