Skip to content

v1.5.0 — retrained contamination model (simulated MAGs) + reorganised resources

Choose a tag to compare

@fmschulz fmschulz released this 21 Apr 05:36
· 31 commits to main since this release

GVClass v1.5.0

Direct release on top of the live v1.4.x line. Ships the combined
v1.4.3 remediation round plus a retrained contamination model trained
on simulated giant-virus MAGs that eliminates the novel-virus
false-positive class observed in earlier releases.

Headline metrics

On a 250-bin simulated-MAG + intact-isolate benchmark:

v1.4.x runtime v1.5.0 retrained
Mean predicted contamination on clean shredded NCLDV 28.8% 4.3%
Clean bins predicting > 10% (50 bins) 100% (50/50) 8% (4/50)
Contaminated-bin MAE 9.4% 2.55%
Pearson r on contaminated bins ~0 0.97
External holdout on 50 unseen isolates — max prediction n/a 0.00%

Example fixture GVMAG-S-1096109-37: 27.15% → 0.08% (false-positive resolved).

What's new

Retrained contamination model on simulated MAGs

  • Training set: 200 simulated giant-virus MAG bins (isolate genomes fragmented
    into MAG-like contigs with bacterial, eukaryotic, NCLDV-host-eukaryotic, and
    viral-mixture contamination donors) + 50 clean intact isolate genomes so the
    model generalises to both shredded and whole-contig inputs.
  • 50 base NCLDV isolates across 13 families; 10 reserved as a held-out
    novel-virus generalisation panel.
  • ExtraTrees regressor (selected over RandomForest / HistGradientBoosting by
    GroupKFold CV); 21 features including five per-contig taxonomic-purity columns.

Per-contig taxonomic-purity classifier

Novel viruses with scattered cellular markers from HGT no longer downgrade to
mixed_viral. Bins with no cellular-coherent contigs but ≥3 viral-bearing
contigs downgrade to uncertain so curators can triage. New output columns:
cellular_coherent_contig_count, cellular_coherent_protein_fraction,
cellular_coherent_bp_fraction, cellular_lineage_purity_median,
cellular_hit_identity_median, viral_bearing_contig_count,
contig_attribution_mode.

Reorganised runtime resources

resources/
├── labels.tsv
├── hmm/                 combined HMM profiles
├── markers/             annotations, stats, order_completeness.tab
├── completeness/        novelty-aware completeness model + configs
├── contamination/       retrained ExtraTrees bundle + model card
└── database/            protein reference FASTAs (dmnd dropped)
  • Dropped the novelty_strategy{2,3}_ prefix — names now describe the file.
  • Contamination model moved out of src/bundled_models/ into
    resources/contamination/; model rotation is now a resources-tarball concern.
  • database/dmnd/ directory dropped (pipeline uses pyswrd directly on
    .faa references). Tarball size 3.2 GB → 1.7 GB; uncompressed 7.3 GB → 3.7 GB.

Correctness (v1.4.3 track rolled in)

  • Multi-HMM dedup in models.out.filtered
  • Contamination model calibrated for sensitive-mode features
  • Atomic resume via JSON SUCCESS sentinel + verified tar
  • TOCTOU-safe joblib load with SHA-256 gate
  • Per-marker taxonomy majority with deterministic ties; new
    taxonomy_confidence column (high / low_support / reduced_fastmode)
  • Transactional contig splitter (--contigs)
  • Hard-fail on short inputs; opt-in --allow-short
  • R² gating on novelty-aware completeness; new estimated_completeness_quality
    and estimated_completeness_advisory columns
  • Redundant per-query BLAST pass removed (halves per-query BLAST cost)

Infrastructure

  • src/__version__.py is the single source of truth; pyproject.toml with
    gvclass console script; committed pixi.lock for reproducibility
  • CI: lint + mypy + pytest on push, golden-file regression on PR, Apptainer
    SIF build on tag
  • Module renames: prefect_flow.py → parallel_runner.py,
    gvclass_prefect.py → gvclass_runner.py

Breaking changes

  • Bundled contamination model replaced; SHA-256 constant rotated. The
    pipeline refuses to load the pickle unless the on-disk digest matches
    the code constant (by design).
  • Resources layout reorganised — REQUIRED_FILES in DatabaseManager now
    expects the new nested paths. v1.4.x tarballs will NOT pass validation;
    use pixi run setup-db to install the v1.5.0 bundle.
  • Module renames (internal consumers): prefect_flow.py → parallel_runner.py,
    gvclass_prefect.py → gvclass_runner.py.

Reference resources

Hosted on Zenodo (concept DOI 10.5281/zenodo.18662445 resolves to latest):

  • v1.5.0 DOI: 10.5281/zenodo.19674504
  • Tarball: resources_v1_5_0.tar.gz (1.75 GB)
  • SHA-256: 5357d96d99aa1eaf4b396ef701ed4c3b22d9015f79b7ae6c6be354c897704c80
  • NERSC mirror: https://portal.nersc.gov/cfs/nelli/gvclassDB/resources_v1_5_0.tar.gz

Install

# Pixi (local / dev)
git clone https://github.com/NeLLi-team/gvclass.git
cd gvclass
pixi install
pixi run setup-db           # downloads resources_v1_5_0.tar.gz from Zenodo
pixi run gvclass example

# Apptainer (HPC) — image pulls transparently
wget https://raw.githubusercontent.com/NeLLi-team/gvclass/main/gvclass-a
chmod +x gvclass-a
./gvclass-a my_genomes my_results -t 32

Known limitations (tracked for v1.5.1)

  • Holdout evaluation on rare HGT-rich families (Mesomimiviridae,
    Schizomimiviridae) can predict 10–20% on clean novel isolates — a
    data-sparsity effect (≤ 3 training examples per family). v1.5.1 will
    augment these families.