Releases: HERRY423/BioNexus
Releases · HERRY423/BioNexus
Release list
BioNexus v1.0.0-rc.2
[1.0.0-rc.2] - 2026-08-21
🛡️ Added (Cryptographic Provenance & Standalone Transparency Proofs)
- Standalone Rekor Transparency Proofs (
evidence/rekor_transparency_proof.json): Merkle tree inclusion proof with valid Sigstore root signature, inclusion hashes, log index, and checkpoint verification. - RFC 3161 TSA Timestamp Evidence (
evidence/tsa_timestamp_token.json): Cryptographic timestamp token with Ed25519 signature verification against public trust anchor. - Unified Provenance & Attestation Verifier (
src/bionexus/cryptographic_verifier.py,src/bionexus/attestation_authority.py): Strict, fail-closed Merkle root and attestation verification. - Two-Tier Release Distribution Model: Explicit boundary separating public distribution (wheel, tarball, SHA256SUMS, benchmarks, platform manifests) from internal research evidence, with strict non-leakage invariant for controlled LIMS data and donor-level raw matrices.
- GitHub Artifact Attestations (
.github/workflows/release.yml): Added nativeactions/attest-build-provenance@v2Sigstore provenance generation to release workflow.
🧹 Fixed (Code Quality & SSOT Synchronization)
- Ruff Lint & Import Order Hardening: Resolved unused imports, missing newlines, and E402 script import order issues across tests, scripts, and source modules.
- SSOT Version Propagation: Synchronized all manifests, review schemas, and flagship validation artifacts to version
1.0.0-rc.2.
BioNexus v1.0.0-rc.1
[1.0.0-rc.1] - 2026-08-20
🛡️ Added (Data Governance & Data Egress Contract — BNS-SEC-001..010)
-
Runtime Egress Guard Engine (
src/bionexus/egress_guard.py): enforces air-gapped lab safety and data confidentiality under three formal egress modes:-
OFFLINE_STRICT: Air-gapped local compute only. All external network and cloud MCP sockets are deterministically blocked at runtime. -
ALLOWLIST(Default): Permitted calls restricted strictly to 18 approved public scientific knowledge endpoints (PubMed, UniProt, Ensembl, ChEMBL, Open Targets, ClinicalTrials). Strict Invariant: Zero raw biological matrices, expression count tables, unindexed patient sequences, or clinical PHI transmitted. Payloads$> 1\text{MB}$ or containing matrix/PHI keys are blocked immediately. -
CONNECTED: External calls permitted with mandatory cryptographic audit logging.
-
-
Cryptographic Audit Ledger (
logs/egress_audit.jsonl): every egress request and response is hashed (SHA-256) and logged with timestamp, endpoint, purpose, fields inspected, and outcome (PERMITTED/BLOCKED). -
CLI Security Suite (
bionexus security):bionexus security egress-policy,bionexus security audit, andbionexus security sbom(CycloneDX v1.5 JSON). -
Institutional Security Documentation Surface:
-
SECURITY.md: Vulnerability reporting (48h response), supported versions, and security architecture. -
docs/security/THREAT_MODEL.md: High-value assets, threat actors, prompt injection, MCP poisoning, and supply-chain mitigations. -
docs/security/DATA_CLASSIFICATION.md: 4-tier data classification (PUBLIC_BENCHMARK,PROPRIETARY_UNPUBLISHED,CONTROLLED_ACCESS_GENOMIC,RESTRICTED_CLINICAL_PHI). -
docs/security/SECRET_HANDLING.md: Zero hardcoded secrets invariant and pre-commit scanning. -
docs/security/SBOM.md&scripts/generate_sbom.py: CycloneDX SBOM generator. -
docs/security/RELEASE_SIGNING.md: Sigstore Cosign keyless release signing & GitHub Artifact Attestations.
-
🧭 Changed (Context-Conditioned Epistemic Ladder — Rejecting "Magic Number" Refusals)
-
6-Stage Epistemic Decision Ladder: replaced simplistic
$N < 3 \to \text{refuse}$heuristics with a rigorous statistical ladder:
Design Identifiability?$\to$ Dispersion Estimability?$\to$ Uncertainty Quantified?$\to$ Power & Effect-Size Regime?$\to$ Claim Class Evaluated?$\to$ Evidence Ceiling Assigned. -
Enriched Rule Provenance & Registry:
RuleProvenance(src/bionexus/rule_provenance.py),src/bionexus/data/rule_registry.json, andreview/SCIENTIFIC_RULE_CATALOG.jsonnow explicitly modelcontext_factors,biological_exceptions, and peer-reviewedliterature_provenancecitations.
🔬 Added (Flagship Capabilities Empirical Credibility Closed Loop — 12/14 Criteria / VALIDATED Tier)
-
10-Dimensional Spatial Validity Confounder Benchmark (
evals/spatial_stress_test.py): actively tests 10 spatial confounder mechanisms: baseline, segmentation leakage, cell density, cell area morphology, nuclear eccentricity, tissue boundary effects, neighborhood radius sweep (15–100$\mu m$ ), transcript spillover, FOV batch confounding, and coordinate permutation null. -
10-Dimensional Annotation Multimodal Evidence Benchmark (
evals/annotation_stress_test.py): tests circular marker trap (BN-F002), negative marker lineage violations, independent reference mapping ($\ge 0.70$ ), CITE-seq surface protein concordance ($\ge 0.75 \to \text{ROBUST}$ ), discordant modalities (CONFLICTED), open-set gating (ABSTAIN), doublet artifacts, clustering resolution sweep, and adversarial overclaim interception. -
Elevation to VALIDATED Tier: elevated
scrna.annotation_evidenceandspatial.inference_validityalongsidescrna.pseudobulk_detoVALIDATEDtier with 12/14 criteria satisfied.
🚦 Changed (CI Matrix Overhaul — Zero || true, Explicit Reliability Tiers)
-
Eliminated all
|| trueerror suppression in.github/workflows/ci.yml. -
Three Structured Matrix Tiers:
-
core-matrix: Python 3.10–3.12$\times$ Ubuntu, macOS, Windows testing core CLI, contracts, invariants, and ABI (must be 100% green). -
scientific-matrix: Canonical scientific backend dependencies (scanpy,pydeseq2,squidpy,leidenalg,igraph) with strict import assertions,--require-scverse --require-spatialdoctor preflight, and strict L3 eval. -
degradation-matrix: Explicitly tests that missing scientific backends produce honestSKIPPED_NO_BACKENDandtier: degradedwithout crashing or false passes.
-
🏛️ Added (Community Governance & 7-Stage Closed-Loop Rule Challenge Lifecycle)
-
7-Stage Closed-Loop Rule Challenge Lifecycle (
docs/governance/RULE_CHALLENGE_LIFECYCLE.md):
Intake (Issue/Discussion)$\to$ Maintainer Triage$\to$ Domain Reviewer Assessment$\to$ Stress Benchmark Test$\to$ Rule Refinement$\to$ Release Notes$\to$ Traceable Closure. -
Cleaned all
file:///local paths acrossCONTRIBUTING.md,README.md, anddocs/plugin-development.mdinto repository-relative links.
⚖️ Added (Evidence Model — Evidence Strength ≠ Intended Use Requirement)
- Third warrant-engine decoupling (
src/bionexus/evidence_model.py): purpose decides the evidence requirement, never the evidence value. A study with 10 donors/group, pre-registration, adequate power, and an independent replication carries ROBUST evidence whether the researcher calls it exploratory or confirmatory; weak data does not acquire a REPLICATED standing because someone declares a clinical purpose. - Three explicit objects:
EvidenceAssessment(how strong the evidence IS — computed only from declared evidence factorsreplication / sample_design / effect_stability / external_validation / sensitivity_analysis / confound_controls / backend_fidelity / provenanceand active violations; purpose- and policy-independent by construction),ClaimContext(nine claim classes, descriptive → clinical_actionability, each with its own minimum bar viaCLAIM_REQUIREMENTS), andUseRequirement(purpose + claim class composed — the only place purpose enters). evaluate_sufficiency()compares evidence against the composed bar:WARRANTED·WARRANTED_WITH_LIMITS(documented ack; the bar never moves) ·NOT_SUFFICIENT_FOR_INTENDED_USEwith an explicit gap list. Undeclared intended use is never sufficient for any use.research_purpose.py:PURPOSE_EVIDENCE_CEILINGis reinterpreted asPURPOSE_EVIDENCE_REQUIREMENT(same numbers, new semantics; the old name survives as a deprecated alias).PurposeContext.required_evidencereplacesevidence_ceiling(deprecated).assess_warrant()accepts anEvidenceAssessmentand starts the ceiling from what the evidence is worth;evaluate_viability_with_purpose()threadsevidence_factors/claim_context/documented_extrasand attachesevidence_assessment+sufficiencyto the EvidenceCard.- 16 new theory-invariant tests (
tests/unit/test_evidence_model.py), including the two canonical examples: ROBUST + population_effect + confirmatory → WARRANTED; SUPPORTED + clinical → NOT_SUFFICIENT_FOR_INTENDED_USE.
🛡️ Added (Backend Identity Conformance — BNS-EF-012..016 / BN-F010)
src/bionexus/backend_conformance.py+ CLIbionexus backend-identity: every canonical capability now answers a machine-checkable identity audit — claimed backend, observed executed backend, entry points, version, execution fingerprint, and fallback flag.declared_backend == observed_backendis verified via the installed-distribution witness (importlib.metadata.packages_distributions).
🚀 Changed (Release Pipeline Automation & Dynamic Pre-Release Tagging)
-
.github/workflows/release.ymlautomatically detects-rc,-alpha,-betatags and setsprerelease: trueon GitHub Releases. - Verified clean-venv wheel execution gate: Version SSOT
$\to$ ruff$\to$ unit tests$\to$ full scientific backend$\to$ strict benchmark$\to$ build$\to$ clean venv$\to$ wheel install$\to$ doctor$\to$ registry check$\to$ backend identity$\to$ strict eval$\to$ manifest validation$\to$ SHA256$\to$ GitHub Release.
BioNexus v0.10.0
[0.10.0] - 2026-08-17
🛡️ Epistemic Honesty & Fail-Closed by Default (BNS-EF-002 / BNS-CC-012)
- Fail-closed by default for Tangram, GEARS, and NicheFormer:
run_tangram_spatial_mapping(),predict_gears_perturbation(), andforecast_spatial_niche()now default toallow_fallback=False. Missing backends trigger immediate deterministic refusal (REFUSAL_BACKEND_UNAVAILABLE). Grade C heuristic baselines run only with explicit caller opt-in (allow_fallback=True) and are transparently labeled asGrade C Experimentalwithout masquerading as official neural network models. - Frontier capability segregation: Segregated experimental foundation models and closed-loop exploration (
scfm.geneformer_canonical,scfm.scgpt_canonical,scfm.rank_proxy_embedding,perturbation.gears_prediction,spatial.nicheformer_forecasting,closed_loop.perturbation_to_niche) intoFRONTIER_CAPABILITIES, maintaining the stable core of 13 certified canonical capabilities inCANONICAL_CAPABILITIES. - Version SSOT & Zero-Drift Guard: Unified versioning across
pyproject.toml,bionexus.registry.yaml,src/bionexus/versions.py,src/bionexus/__init__.py,plugin.json,marketplace.json, and all client manifests. Addedscripts/sync_version.pyandtests/unit/test_version_ssot.pyCI enforcement.
🌐 Added (Standards Interoperability — BNS-016: no proprietary data-standard island)
src/bionexus/interop.py: BioNexus exports through published community standards instead of inventing a proprietary research bundle format. Run capsules project to RO-Crate 1.1 following Workflow Run Crate conventions (capability →ComputationalWorkflowwith the Workflow RO-Crate profile; execution → schema.orgCreateActionwith instrument/object/result/startTime/endTime and the Process Run Crate 0.5 profile; evidence maturity rides inside the crate as a contextual entity). Ledgers project to RO-Crate with claims/evidence as contextual entities (isBasedOnsupport edges). Run capsules project to BioCompute Objects (IEEE 2791-2020) with all six domains (provenance, usability, description, execution, io, parametric) and a content-computedetag. Deterministic, offline projections; fail-closed exports (a projection failing structural validation is never written); structural validators with disclosed scope. CLI:bionexus interop ro-crate|bco|check.src/bionexus/standards.py+bionexus standards: machine-readable standards alignment registry with a closed, honest status vocabulary —implemented(RO-Crate, Workflow Run Crate, BCO, PROV-O) /aligned(Bioschemas typing, nf-core schemas) /proposal(GA4GH AI Work Stream) /tracked(ELIXIR, scverse, Bioconductor, WorkflowHub) — and the mandatory verbatim disclaimer: BioNexus is not an industry standard and does not claim to be one.docs/standards-engagement.md: the GA4GH AI standardization window strategy — a concrete mapping of BioNexus artifacts (failure taxonomy, capability-contract schema, refusal semantics, host conformance, BioFailureBench) onto the AI Work Stream focus areas, engagement venues, and the contribution rule: offer vocabulary, schemas, and tests; never announce; let adoption invert the direction.- New spec
spec/BNS-016-standards-interop.md(BNS-IO-001..013); teststest_interop.py,test_standards.py.
🧭 Added (Product matrix & scope boundary — BNS-IO-012)
docs/product-matrix.md: the four-layer matrix (bionexus-core / bionexus-audit / bionexus-conformance / reference capability packs) with a test-enforced module mapping and the explicit non-goals list — no planner, memory, multi-agent, chat UI, cloud workspace, notebook replacement, compute service, or agent marketplace. README gains the matrix section;test_product_matrix.pyguards the documented mapping against drift and enforces downward-only layering (core never imports the audit layer).
🔬 Added (Why-install case on the front page)
- README opens with the one case a computational biologist understands immediately: the before/after of "Run DE between these two clusters" — an agent that returns "153 significant genes" vs BioNexus blocking with BN-F002 Pseudoreplication and the pseudobulk remedy. One case beats a hundred features.
🧱 Added (Scientific Assertion Firewall — BNS-013)
Product repositioning: BioNexus catches biological analyses that should not have been run. Three researcher-facing entry points, usable without any host agent:
bionexus preflight(src/bionexus/preflight.py): runs BEFORE compute. Resolves the declared intent onto a capability contract, inspects the actual data state (matrix semantics, donor/condition structure, confounding, spatial provenance — reading.h5addirectly when anndata is present), and renders the seven-section verdict block: INTENT / DATA STATE / RISKS / DECISION / ALLOWED / FORBIDDEN CLAIM / REMEDY. Decision vocabulary is the fail-closed table (BNS-AD-014) verbatim; ALLOWED under a prevented decision comes only from taxonomyacceptable_degradation; FORBIDDEN CLAIM is mechanically derived from the capability's forbidden-claim catalog + evidence ceiling. Exit codes: 0 proceed (incl. capped/degraded), 1 refused/blocked, 2 missing evidence.bionexus audit <notebook|script>(src/bionexus/analysis_audit.py): deterministic static rule engine (BFA-001..BFA-013) over.ipynb/.py/.R/.Rmd/.qmdscreening the canonical trap classes: cell-level pseudoreplication, raw/log matrix confusion, missing FDR, batch/condition confounding, inappropriate statistical unit, annotation without evidence, circular marker validation, missing negative controls, spatial coordinate substitution (incl.obsm['spatial'] = obsm['X_umap']), parameter instability, overclaimed causality (via the prohibited-claims auditor), backend substitution, and unexecuted code claims. Findings cite rule id + BN-Fxxx + evidence line + remedy; the mandatory disclaimer states that absence of findings is NOT proof of validity. Data-file audit behavior (.h5ad/csv matrix semantics) is preserved unchanged.bionexus verify <results>(src/bionexus/verification.py): verifies final results against their Claim–Evidence Ledger (BNS-012): fail-closed re-resolution per claim, evidence lines with honest symbols, ceiling cross-check, and not-warranted flagging of causal/mechanistic language beyond the evidence class. Exit non-zero on ABSTAIN/CONFLICTED/unwarranted claims; honest intermediate maturities do not fail.- New spec
spec/BNS-013-scientific-assertion-firewall.md(BNS-FW-001..014); teststest_preflight.py,test_analysis_audit.py,test_result_verify.py.
🏆 Added (Flagship Certification Track + two flagship capabilities — BNS-015)
- Flagship principle: three CERTIFIED capabilities with independent external validation outweigh ten self-tested certifications. Flagship set:
scrna.pseudobulk_de(A: pseudoreplication),scrna.annotation_evidence(B: annotation evidence),spatial.inference_validity(C: spatial inference validity). The M4 10-CERTIFIED target is unchanged (BNS-CF-006); the flagship track is prioritization, never weakening. External criteria (public dataset, independent ground truth, cross-host, external reviewer) cannot be implementer-satisfied.bionexus certificationnow publishes the flagship section with per-capability external-criteria-remaining. src/bionexus/annotation_evidence.py+ contractscrna.annotation_evidence: not another CellTypist — assesses how much evidence backs a candidate cell-type label (reference mapping, marker consistency, negative markers, doublet risk, ontology compatibility, open-set detection, cross-method agreement) and returns per-label verdicts SUPPORTED / TENTATIVE / ABSTAIN with published deterministic thresholds.src/bionexus/spatial_inference.py+ contractspatial.inference_validity: not a Squidpy reimplementation — tests whether a spatial conclusion survives its alternative explanations (12-control canonical registry: cell size, transcript density, segmentation uncertainty, nuclear eccentricity, local density, spot composition, spatial autocorrelation, batch/FOV, ligand/receptor abundance, contact geometry, neighborhood radius, permutation null). Verdict ladder ROBUST / SUPPORTED / FRAGILE / ABSTAIN; ceiling FRAGILE without orthogonal validation.- New spec
spec/BNS-015-flagship-certification.md(BNS-FC-001..008); teststest_flagship_capabilities.py.
🪤 Added (BioFailureBench: the Scientific Trap Corpus — BNS-014)
evals/datasets/biofailurebench.yaml: 26 traps (23 gating — all passing deterministically; 3 frontier known limitations) covering all twelve taxonomy modes. Each trap is a complete record with eight fields: data, intended analysis, hidden flaw (BN-Fxxx), expected detection, allowed computation, forbidden claim, remediation, reference. Includes a positive control (BF-024) so the bench cannot degrade into an all-refusal benchmark. Host-agnostic: Claude, Codex, Cursor, Biomni, and future agents run the identical suite viabionexus eval --suite biofailurebench.evals/biofailurebench.py+bionexus bench validate: machine-checked corpus integrity (field completeness, taxonomy linkage, gating/frontier prefixes, ID resolution, mode coverage); invalid corpora fail CI.- Three taxonomy open gaps CLOSED with wired detection (BNS-FT-008): BN-F004 identifier mismatch (router stage-3.5 namespace screen, BF-008/BF-025), BN-F005 missing FDR (
abi.enforce_statistical_warrantcaps warrant at PRELIMINARY, BF-005), BN-F008 cross-database contradiction (router trap screen, BF-016). BN-F009 embedding-coordinate substitution is now refused at routing (BF-007/BF-020). Perfect condition-donor confounding and open-set/annotation-evidence traps are screened deterministically (BF-003/BF-013, BF-004/BF-022). - New spec
spec/BNS-014-biofailurebench.md(BNS-BF-001..009)...
BioNexus v0.9.0
[0.9.0] - 2026-08-16
🌐 Fixed (FastMCP Dynamic Tool Registration & Fallback Routing Disambiguation)
- Fixed FastMCP default tool leakage in
scripts/local_mcp_server.py: wrapped the 6 hosted-overlap fallback tools (search_pubmed,get_pubmed_article,search_biorxiv,search_chembl,search_opentargets,search_clinical_trials) andsearch_cosmicbehindBIONEXUS_LOCAL_HOSTED_FALLBACKS=1. By default, FastMCP now registers exactly 9 local unique tools (GTEx, GEO, STRING, UniProt, Ensembl, gnomAD, PDB, AlphaFold, Reactome) plus 6 Resources and 6 Prompts. - Eliminated Agent routing ambiguity: AI coding agents querying literature, targets, or clinical trials will cleanly route to dedicated cloud-hosted MCP endpoints without duplicate tool confusion.
- Synced test coverage: updated
tests/unit/test_mcp_server.pyto assert 9 default unique tools and 16 tools upon opt-in.
🎖️ Added (Capability Certification Program — BNS-010)
src/bionexus/certification.py: 14 evidence criteria and four tiers (CERTIFIED / VALIDATED / EXPERIMENTAL / CONNECTOR-ONLY). Tiers are computed from recorded evidence, never asserted (BNS-CF-002); structural cross-checks re-verify contract-derived criteria against the live ABI, preconditions, and taxonomy. Honest current state: 0 CERTIFIED, 7 VALIDATED, 1 EXPERIMENTAL — the per-capability blocking-criteria list is the published roadmap to the M4 target of 10 CERTIFIED (evidence must be produced, criteria never weakened, BNS-CF-006). CLI:bionexus certification.- New spec
spec/BNS-010-capability-certification.md; teststests/unit/test_certification.py(CERTIFIED requires all 14 — structurally un-gameable).
🧯 Added (Scientific Failure Taxonomy — BNS-011)
src/bionexus/failures.py: twelve normative failure modes (BN-F001 assay-state confusion, BN-F002 pseudoreplication, BN-F003 unsupported annotation, BN-F004 identifier mismatch, BN-F005 missing multiple-testing correction, BN-F006 invalid model assumption, BN-F007 parameter instability, BN-F008 cross-database contradiction, BN-F009 missing spatial provenance, BN-F010 backend degradation masquerading, BN-F011 claim inflation, BN-F012 unexecuted maturity claim). Each record: definition, canonical example, affected capabilities, detection rule, fail-closed required behavior, acceptable degradation, benchmark cases. Three modes are honestly flagged as open gaps (no benchmark coverage yet).classify_violation()tags runtime violations with taxonomy IDs. CLI:bionexus failures list|show.- New spec
spec/BNS-011-failure-taxonomy.md; tests verify record shape, vocabulary, and that every benchmark-case reference resolves to a real eval case.
🚫 Added (Fail-Closed Gate — BNS-005 §6)
src/bionexus/failclosed.py:prevent_invalid_run()— the canonical gate implementing knowing when not to compute is a scientific capability: missing evidence → ABSTAIN (request data), invalid input → REFUSE, backend unavailable → DEGRADE WITH DISCLOSURE, assumption violated → BLOCK CLAIM, claim beyond warrant → BLOCK CLAIM, external validation absent → CAP EVIDENCE LEVEL. No row resolves to silent execution. ReturnsPreventionDecisionwith failure-mode IDs, remedies, and the underlying routing decision. CLI:bionexus prevent "<query>".- New spec requirements BNS-AD-013..015; tests cover all six rows plus the clean RUN PERMITTED exit.
📒 Added (Claim–Evidence Ledger — BNS-012)
src/bionexus/ledger.py: claims as auditable dependency graphs —ClaimRecord(supported_by / contradicted_by / depends_on) over closed-vocabularyEvidenceRefnodes (dataset, transformation, method_run, statistical_result, database, cross_method). Fail-closed resolution: any contradiction forces CONFLICTED; no support forces ABSTAIN; otherwise the weakest supporting warrant, clamped by the capability's ABI evidence ceiling (database/cross-method support counts as external validation). JSON round-trip + PROV-O JSON-LD projection; append-only (duplicate IDs rejected). Deliberately a data structure, not a graph platform. CLI:bionexus ledger show|jsonld.- New spec
spec/BNS-012-claim-evidence-ledger.md; tests include the CLAIM-017 reference scenario. - ABI clamp refinement: warning states (FRAGILE / CONFLICTED / ABSTAIN / UNASSESSED) are never rewritten by evidence ceilings — only ascending-ladder warrant levels (PRELIMINARY→REPLICATED) are clamped.
📜 Added (BioNexus Scientific Contract Specification — BNS series)
spec/normative specification tree: nine RFC 2119-style documents (BNS-001..BNS-009) plus index, defining the scientific contract that binds BioNexus and any connected host agent — capability contract & Scientific ABI, input semantic invariants, execution fidelity, evidence maturity, abstention & degradation, provenance, cross-method validation, host conformance, and capability lifecycle. Every requirement carries a stable ID (BNS-XX-nnn) with a live verification hook (unit test, eval category, or runtime refusal).tests/unit/test_spec_conformance.pyenforces document presence, RFC 2119 keyword usage, and cross-document reference integrity.
🧬 Added (Biological Capability ABI — bionexus.abi, ABI v1.0)
- Capability contracts upgraded from metadata to a Scientific ABI: every canonical capability now projects to a machine-readable ABI record (
input_contractwith allowed matrix states and coordinate types,preconditions,forbidden_claims,executionreference backend/algorithm,validationpolicy,evidence_ceiling,provenancerequirements), generated from the canonical contract so it cannot drift.forbidden_claimsandevidence_ceiling_without_external_validationare new normative fields onCapabilityContract. - Normative forbidden-claim taxonomy (
FORBIDDEN_CLAIM_CATALOG): 11 claim families (causal interaction, cell-cell communication, cell-type identity without reference, clinical diagnosis, treatment recommendation, model substitution, hazard causation, true-expression recovery, sensor calibration, regulatory compliance, pipeline results without execution) with detection patterns. - Routing-time forbidden-claim interception (BNS-AD-009): requests asking a capability for a claim on its forbidden list are now deterministically refused with the scientific reason and reformulation remedy (e.g. "use Moran's I to prove cell-cell communication" → ABSTAIN).
- Evidence-ceiling clamping (
enforce_evidence_ceiling): over-warranted maturity claims are clamped to the capability's ceiling (spatial SVG → FRAGILE without external validation; exploratory clustering → PRELIMINARY; REPLICATED requires external truth sets). - CLI:
bionexus abi list|show <id>|audit-claims <id> --claims ...|conformance. - New unit tests:
tests/unit/test_abi.py(10 tests: projection completeness, single-source-of-truth, claim audits, ceiling clamps, router interception, no-false-positive controls, CLI surface).
🎯 Changed (Calibration honesty — frontier track, BNS-LC-004..007)
- Frontier calibration track: new
evals/datasets/calibration_edge.yaml(11 probes,known_limitation: true) exploring adjacent-rank maturity discrimination, coordinate-substitution detection, statistical-power auditing, multi-intent routing, and ABI ceiling clamps. Frontier cases are executed and reported with honest pass/fail but excluded from gating metrics until graduation. - Honest benchmark reporting: reports now separate gating (guaranteed behavior) from frontier (known limitations) and state the union accuracy — gating-only 100% is explicitly labeled NOT a calibration claim. Current honest state: gating 42/42 (CRI 100%), frontier 7/11, union 49/53 = 92.5%, union calibration verdict UNDERCONFIDENT (macro-F1 96.3%, OCE 0.041). The four open known limitations are published by name in every report.
- Calibration metrics extended: adjacent-rank error rate, within-one accuracy, per-class precision/recall/F1, calibration verdict, skipped-no-backend accounting, and cross-host consistency (single-host runs reported as not evaluated rather than trivially consistent).
- L2/L3 maturity attribution made honest: claim audits warrant at most PRELIMINARY (they verify absence of overclaim, not statistical support); L3 outcome cases attest SUPPORTED only when the gold pipeline actually recovered the planted signal, and are excluded from calibration (disclosed count) when optional backends are absent. Fixes the previously un-reproducible "100% calibration" committed report (the prior numbers masked 12 L2/L3 maturity mismatches).
- Spec-strengthened gating cases (BNS-AD-009):
claim-gxppart11-001andclaim-acmg-clinical-001now expect ABSTAIN — requesting an FDA Part 11 certified audit trail or an official clinical diagnostic report is refused at routing time instead of being permitted and audited post-hoc. Three new forbidden-claim refusal cases added (refuse-forbidden-*). - Loader fix:
expected_maturityis now actually read from YAML suites (was silently dropped). - New unit tests:
tests/unit/test_calibration_frontier.py(8 tests) and updatedtest_eval_harness.py.
🔒 Fixed (Single source of truth for the plugin mirror trees)
- Dual skill-tree drift eliminated:
skills/single-cell-rna-qc/scripts/scrna_pipeline.pyhad diverged by 49 lines between the canonical root tree and theplugins/bionexus/skills/copy (the mirror lacked--run-dir/ Run Capsule support). The mirror is now regenerated from the root; both trees are byte-identical. - Drift detection now covers code trees, not just JSON manifests:
bionexus registry --checkandscripts/registry_compiler.py --checkverifyskills/andscripts/are byte-identical to theirplugins/bionexus/mirrors (content edits, missing files, and stale mirror-only files all fail CI);--generateresynchronizes the mirrors automatically. Ignored artifacts (__pycache__, logs, docto...
BioNexus v0.8.0
[0.8.0] - 2026-08-15
🚀 Added
-
Machine-Readable Capability Contracts (
bionexus.capabilities):- Implemented formal
CapabilityContract,SemanticInputType,Precondition,RefusalTrigger,EvidenceRequirement, andCapabilityEvaluationResult. - Registered 8 canonical capabilities:
scrna.pseudobulk_de,scrna.exploratory_clustering,spatial.morans_svg,survival.kaplan_meier,scvi.probabilistic_vae,allotrope.format_conversion,nextflow.pipeline_launch,variant.acmg_classification. - Added CLI subcommands:
bionexus capability [list|show|check].
- Implemented formal
-
6-Stage Scientific Intent & Invariant Router (
bionexus.intent_router):- Implemented 6-stage routing pipeline: Intent Extraction
$\to$ Data Semantics$\to$ Preconditions$\to$ Capability Matching$\to$ Backend Probe$\to$ Decision. - Added authoritative routing statuses:
PERMITTED,NEEDS_DATA,ABSTAIN,DEGRADED_ADVISORY. - Added CLI command:
bionexus route "<query>" [--data <path>] [--min-replicates <N>].
- Implemented 6-stage routing pipeline: Intent Extraction
-
BioNexus Eval: Agent Behavior & Scientific Reliability Benchmark (
evals/):- Implemented 8 Core Reliability Metrics: Routing Accuracy, Unsafe Invocation Rate, Abstention Precision/Recall, Capability Hallucination Rate, Backend Fidelity, Scientific Semantic Error Rate, Evidence Calibration, and Composite Reliability Index (CRI).
- Created 6 evaluation datasets across 29 structured prompts:
routing.yaml,refusal.yaml,capability_claim.yaml,scientific_semantics.yaml,backend_failure.yaml,adversarial.yaml. - Added CLI command:
bionexus eval [--suite <name>] [--report <path>] [--json].
-
Single Source of Truth (SSOT) Multi-Platform Manifest Compiler (
bionexus.registry):- Canonical
bionexus.registry.yamlcompiling into Agent Plugins 1.0 (plugin.json,mcp.json), Claude Plugin (.claude-plugin/plugin.json), and OpenAI Codex (.codex/config.json). - Added CLI command:
bionexus registry [--generate|--check|--validate-endpoints].
- Canonical
-
Skill Scaffolding Tool (
bionexus.scaffold):- Added CLI generator
bionexus create-plugin <name>generating Gold Reference skill directories, pipelines, and offline unit test fixtures.
- Added CLI generator
-
Ecosystem Governance Documentation:
- Added
docs/versioning-policy.md,docs/compatibility-matrix.md,docs/migration-guide.md,docs/deprecation-policy.md. - Added automated GitHub release workflow
.github/workflows/release.yml.
- Added
🔄 Changed
- EvidenceCard 2.0 (Three-Layer Epistemic Model):
- Decoupled into
ExecutionState(EXECUTED,DEGRADED,REFUSED,FAILED),DimensionGrade(A,B,C,UNTESTED,NOT_APPLICABLE,INSUFFICIENT,CONFLICTED), andConclusionMaturity(ABSTAIN,FRAGILE,CONFLICTED,PRELIMINARY,SUPPORTED,ROBUST,REPLICATED). - Added non-breaking backward compatibility aliases
ConclusionStatusandEvidenceGrade.
- Decoupled into
- Parameter Robustness Auditing:
- Added
audit_parameter_stability()inbionexus.integritycomputing Adjusted Rand Index (ARI) and Jaccard similarity across parameter perturbation sweeps.
- Added
🛡️ Fixed
- Fixed single-cell condition differential expression to require biological replicates (
$n \ge 2$ ) and refuse single-sample pseudoreplication. - Fixed negative binomial GLM inputs to refuse continuous normalized floats and enforce discrete integer counts.
- Fixed cell-type naming invariants to prohibit unverified hallucinated biological labels.