Repository navigation
v1.5.0 — retrained contamination model (simulated MAGs) + reorganised resources
GVClass v1.5.0
Direct release on top of the live v1.4.x line. Ships the combined
v1.4.3 remediation round plus a retrained contamination model trained
on simulated giant-virus MAGs that eliminates the novel-virus
false-positive class observed in earlier releases.
Headline metrics
On a 250-bin simulated-MAG + intact-isolate benchmark:
| v1.4.x runtime | v1.5.0 retrained | |
|---|---|---|
| Mean predicted contamination on clean shredded NCLDV | 28.8% | 4.3% |
| Clean bins predicting > 10% (50 bins) | 100% (50/50) | 8% (4/50) |
| Contaminated-bin MAE | 9.4% | 2.55% |
| Pearson r on contaminated bins | ~0 | 0.97 |
| External holdout on 50 unseen isolates — max prediction | n/a | 0.00% |
Example fixture GVMAG-S-1096109-37: 27.15% → 0.08% (false-positive resolved).
What's new
Retrained contamination model on simulated MAGs
- Training set: 200 simulated giant-virus MAG bins (isolate genomes fragmented
into MAG-like contigs with bacterial, eukaryotic, NCLDV-host-eukaryotic, and
viral-mixture contamination donors) + 50 clean intact isolate genomes so the
model generalises to both shredded and whole-contig inputs. - 50 base NCLDV isolates across 13 families; 10 reserved as a held-out
novel-virus generalisation panel. - ExtraTrees regressor (selected over RandomForest / HistGradientBoosting by
GroupKFold CV); 21 features including five per-contig taxonomic-purity columns.
Per-contig taxonomic-purity classifier
Novel viruses with scattered cellular markers from HGT no longer downgrade to
mixed_viral. Bins with no cellular-coherent contigs but ≥3 viral-bearing
contigs downgrade to uncertain so curators can triage. New output columns:
cellular_coherent_contig_count, cellular_coherent_protein_fraction,
cellular_coherent_bp_fraction, cellular_lineage_purity_median,
cellular_hit_identity_median, viral_bearing_contig_count,
contig_attribution_mode.
Reorganised runtime resources
resources/
├── labels.tsv
├── hmm/ combined HMM profiles
├── markers/ annotations, stats, order_completeness.tab
├── completeness/ novelty-aware completeness model + configs
├── contamination/ retrained ExtraTrees bundle + model card
└── database/ protein reference FASTAs (dmnd dropped)
- Dropped the
novelty_strategy{2,3}_prefix — names now describe the file. - Contamination model moved out of
src/bundled_models/into
resources/contamination/; model rotation is now a resources-tarball concern. database/dmnd/directory dropped (pipeline usespyswrddirectly on
.faareferences). Tarball size 3.2 GB → 1.7 GB; uncompressed 7.3 GB → 3.7 GB.
Correctness (v1.4.3 track rolled in)
- Multi-HMM dedup in
models.out.filtered - Contamination model calibrated for sensitive-mode features
- Atomic resume via JSON SUCCESS sentinel + verified tar
- TOCTOU-safe joblib load with SHA-256 gate
- Per-marker taxonomy majority with deterministic ties; new
taxonomy_confidencecolumn (high/low_support/reduced_fastmode) - Transactional contig splitter (
--contigs) - Hard-fail on short inputs; opt-in
--allow-short - R² gating on novelty-aware completeness; new
estimated_completeness_quality
andestimated_completeness_advisorycolumns - Redundant per-query BLAST pass removed (halves per-query BLAST cost)
Infrastructure
src/__version__.pyis the single source of truth;pyproject.tomlwith
gvclassconsole script; committedpixi.lockfor reproducibility- CI: lint + mypy + pytest on push, golden-file regression on PR, Apptainer
SIF build on tag - Module renames:
prefect_flow.py→parallel_runner.py,
gvclass_prefect.py→gvclass_runner.py
Breaking changes
- Bundled contamination model replaced; SHA-256 constant rotated. The
pipeline refuses to load the pickle unless the on-disk digest matches
the code constant (by design). - Resources layout reorganised —
REQUIRED_FILESinDatabaseManagernow
expects the new nested paths. v1.4.x tarballs will NOT pass validation;
usepixi run setup-dbto install the v1.5.0 bundle. - Module renames (internal consumers):
prefect_flow.py→parallel_runner.py,
gvclass_prefect.py→gvclass_runner.py.
Reference resources
Hosted on Zenodo (concept DOI 10.5281/zenodo.18662445 resolves to latest):
- v1.5.0 DOI: 10.5281/zenodo.19674504
- Tarball: resources_v1_5_0.tar.gz (1.75 GB)
- SHA-256:
5357d96d99aa1eaf4b396ef701ed4c3b22d9015f79b7ae6c6be354c897704c80 - NERSC mirror:
https://portal.nersc.gov/cfs/nelli/gvclassDB/resources_v1_5_0.tar.gz
Install
# Pixi (local / dev)
git clone https://github.com/NeLLi-team/gvclass.git
cd gvclass
pixi install
pixi run setup-db # downloads resources_v1_5_0.tar.gz from Zenodo
pixi run gvclass example
# Apptainer (HPC) — image pulls transparently
wget https://raw.githubusercontent.com/NeLLi-team/gvclass/main/gvclass-a
chmod +x gvclass-a
./gvclass-a my_genomes my_results -t 32Known limitations (tracked for v1.5.1)
- Holdout evaluation on rare HGT-rich families (Mesomimiviridae,
Schizomimiviridae) can predict 10–20% on clean novel isolates — a
data-sparsity effect (≤ 3 training examples per family). v1.5.1 will
augment these families.