v0.2.2
A patch release focused on real-data robustness and PyPI-readiness, prompted by running the pipeline against a real published dataset (Okrasińska et al. 2022, Environmental Microbiology — 53 real ITS2/UNITE soil-fungal samples, SRA BioProject PRJNA767765) end to end for the first time at production scale.
Fixed
- Casava filename validation rejected sample identifiers containing dots (common in real SRA-derived names, e.g.
K.BeL.1.1) — a valid file was rejected outright before the pipeline did anything else. - Sample-frequency parsing choked on QIIME2's comma thousands-separator formatting once a sample's frequency reached four digits (e.g.
"110,406.0") — broke the PDF report on any real dataset with more than ~1000 reads/sample, which every prior test dataset was too small to trigger.
Both were found and fixed against the real failure case, not just in isolation — the run that hit each bug went on to complete successfully afterward.
Changed (PyPI packaging prep)
- Fixed README links (logo, license) that were relative to the GitHub repo and would 404 on PyPI's project page.
- Modernized license metadata to the SPDX form (
license = "MIT"+license-files), replacing the table form setuptools was warning would stop being supported. - Added
[project.urls](Homepage/Repository/Documentation/Issues/Changelog) for PyPI's project-page sidebar. - Added
CITATION.cff(enables GitHub's "Cite this repository").
Documentation
- Documented Eukaryome as an additional, directly-fetchable ITS reference source alongside UNITE (
qiime rescript get-eukaryome-data). - Documented why SILVA/GTDB/PR2 aren't usable for ITS classification (none cover the ITS region), including SILVA's pretrained classifier downloads for anyone using this pipeline on a non-ITS marker instead.
- README quick start now leads with installing/activating QIIME2 itself instead of assuming it's already set up.