Releases: FuqingZh/bioextract
Release list
bioextract 0.8.0
What's Changed
Full Changelog: 0.7.1...0.8.0
bioextract 0.7.1
Fix publication-backed Reactome unmatched_ids schema; full local gate passed with 848 tests, and PR #57 passed the required py check. The 0.7.1 tag triggers the package release workflow for PyPI publication.
What's Changed
- docs: close out Reactome v0.5 release by @FuqingZh in #56
- fix(reactome): preserve unmatched selection schema by @FuqingZh in #57
Full Changelog: 0.7.0...0.7.1
bioextract 0.7.0
Reactome mapping capability expansion through v0.5.
- Adds the complete Reactome v96 mapping matrix, human Complex/EWAS entity-pathway relations, and human GMT gene sets.
- Publishes the Reactome v0.5 publication contract with 17 biological tables.
- The package was published successfully by the release workflow; see PyPI for distribution assets.
- Formal resource and catalog activation are recorded separately in the repository release evidence.
What's Changed
- feat: expand Reactome mapping API for v0.2 by @FuqingZh in #54
- feat: complete Reactome v0.5 capability set by @FuqingZh in #55
Full Changelog: 0.6.0...0.7.0
bioextract 0.5.0
Highlights
- Converges supported resource queries on public, request-aware
polars.LazyFrameboundaries. - Keeps compact normalized relation anchors and defers large one-to-many expansion until explicitly requested.
- Adds species-safe and publication-aware behavior across WikiPathways, InterPro/Pfam, UniProt, Rhea, and KEGG mapping/metabolic access.
- Adds execution-integrity validation plus bounded P0-P5 performance evidence and implementation documentation.
Verification
BIOEXTRACT_TEST_THREADS=1 pdm run check: 807 passed.- Release workflow
31983492022: build and PyPI publish jobs succeeded. - Published artifacts: bioextract 0.5.0 on PyPI.
bioextract 0.3.0
Highlights
- Added
GODatabase.select_ancestors()with canonical and alternate GO ID resolution,is_a/part_ofancestor traversal, optional GO slim projection, self mappings, obsolete policies, and unmatched reasons. - Added DuckDB query pushdown for publication-backed
select_terms(),list_subsets(), and ancestor selection without temporary input tables or complete-frame materialization. - Preserved source/publication parity, pinned-publication identity checks, existing nine-table publication compatibility, and downstream ownership of protein membership and enrichment analysis.
Validation
- PR #48 merged at
fa72dbbafter review-thread fixes. - Main
py-cipassed on the merge head. py-publishrun31350254199passed canonical tag validation,pdm run check, distribution build,twine check, and PyPI trusted publishing.- Published wheel SHA-256:
2e460c53ab7cfa8ee96eb57babc259d367aa85cdd5d619d99ff7efb1579a2f72. - Clean-environment install verified version
0.3.0,GODatabase.select_ancestors, and removed legacy constructors.
bioextract 0.1.0
Highlights
bioextract 0.1.0 is the first release of the converged domain-access contract for local biological-database snapshots.
- Added stable top-level database handles for ChEBI, Rhea, GO, KEGG, Reactome, WikiPathways, eggNOG, InterPro, UniProt, STRING, and OmniPath.
- Added query-ready domain access for ChEBI, Rhea, KEGG metabolic data, UniProtKB reviewed entries, GO ontologies, Reactome, WikiPathways, InterPro/Pfam, and publication inspection.
- Added native read-only DuckDB access through
XDatabase.from_duckdb(...).connect()for unrestricted SQL over validated publications. - Added
inspect_publication()for lightweight provenance and schema inspection without selecting a resource-specific reader.
Breaking changes
This release intentionally converges the pre-1.0 API and is not drop-in compatible with the 0.0.x helper API.
- Use complete
*Databasenames and the top-level import form, for examplefrom bioextract import RheaDatabase. - Replace abbreviated
*Dbhandles, resource limits,write_tidy()/directory writers, canonical Parquet writers, and legacy score-filter names with the 0.1.0 database and DuckDB APIs. - Use
write_duckdb(destination)for materialized publications andXDatabase.from_duckdb(path)to reopen them. - GO no longer emits the synthetic subcell projection; query the ontology's declared cellular-component terms instead.
- Selection and result implementation types are no longer stable resource-package exports; use database methods and the documented output schemas.
Storage and provenance
- Materialized bioextract products use one DuckDB publication per logical product, whether the publication contains one relation or many.
- Official/native representations remain directly consumable where they are already suitable, including eggNOG SQLite and the configured STRING/OmniPath sources.
- Publication provenance is embedded in the DuckDB
_bioextractschema; it is not supplied by a sidecar manifest. - Publication writes use validation, staging, bounded execution, and atomic replacement. Official two-dimensional source headers are retained when possible; derived fields use stable
snake_casenames and column lineage is recorded when mapping is required.
Installation
pip install bioextract==0.1.0Python 3.13 or newer is required. Runtime dependencies include Polars, DuckDB, and PyArrow.
Verification
The release tag workflow completed successfully on the exact 0.1.0 tag, including dependency synchronization, the complete repository check, distribution build, twine check, and PyPI Trusted Publishing.
Full changelog: 0.0.8...0.1.0