Skip to content

3.23.0

Latest

Choose a tag to compare

@nextstrain-bot nextstrain-bot released this 19 Aug 12:07

3.23.0

Multi-reference datasets

Datasets can now include multiple reference sequences for auto-detection by nextclade sort and Nextclade Web. This enables reliable dataset suggestion for genetically diverse pathogens (e.g. enteroviruses) where no single reference sequence is close enough to all circulating lineages.

Dataset authors list additional reference FASTA files in the new minimizerIndex field of pathogen.json. The detection fingerprint is built from the union of k-mers across all listed references. The suggestion score divides by the expected hits from a single reference, so adding references broadens the detection range without diluting scores.

Multi-reference detection is transparent: Nextclade Web suggestions, nextclade sort output format, alignment, mutation calling, and all analysis outputs are unchanged. Existing single-reference datasets are unaffected. The minimizer index is validated on load, rejecting malformed entries with clear errors. See datasets documentation, #1771.

Installation and usage

📚 Documentation: docs.nextstrain.org/projects/nextclade

🌍 Nextclade Web: clades.nextstrain.org

🖥️ Nextclade CLI:

  • 📥 Standalone native executables can be downloaded using the links in the "Assets" section just below. Mind the different operating systems and computer architectures.
  • 🐋 Docker images are available on DockerHub
  • 🐍 Conda package is available in bioconda channel.

⌨️ Source code is on GitHub: github.com/nextstrain/nextclade

🚩 Report bugs, suggest features by submitting a GitHub issue

💾 Nextclade datasets are here