Skip to content

3c. RepGenR virus workflow

jaclew edited this page Apr 15, 2024 · 10 revisions

Download and dereplication of Human Mastadenovirus A, B, and E

This example shows how to get a dereplicated for Human Mastadenovirus A.

Download the family-fasta for Adenoviridae:

repgenr vmetadata --target_family adenoviridae --workdir mastadenovirus_test

Useful information exist in the metadata_ncbi_summary.tsv, such as the number of genomes at different taxa and names/identifiers.

Unpack all sequences for subspecies A, B, and E:

repgenr vgenome --workdir mastadenovirus_test --target_species "human mastadenovirus a, human mastadenovirus b, human mastadenovirus e"

Selection can be done at different defined levels: --target_genus, --target_species, --target_serotype. Selection in custom columns is done by, for example, --target_custom "no rank:human adenovirus 14" or --target_custom "undefined_strain:human adenovirus 7d". Matching in selection is case insensitive.

An outgroup dataset is determined by using MASH distance. This feature is disabled by invoking --no_outgroup.

When working with non-curated databases there may be unwanted sequences. Unwanted sequences can be identified through viewing the fasta headers of selected sequences using flags --print_fasta_headers (print selected sequence headers) and --glance (perform a dry-run, do not write output):

`repgenr vgenome --workdir mastadenovirus_test --target_species "human mastadenovirus a, human mastadenovirus b, human mastadenovirus e" --glance --print_fasta_headers`

image
In this image, an unwanted sequence is named "UNVERIFIED", marked with red vertical bars, and can be discarded.

Unwanted sequences can be discarded by invoking --discard flag (case sensitive):

repgenr vgenome --workdir mastadenovirus_test --target_species "human mastadenovirus a, human mastadenovirus b, human mastadenovirus e" --glance --print_fasta_headers --discard UNVERIFIED

Re-do unpack: remove UNVERIFIED sequences:

repgenr vgenome --workdir mastadenovirus_test --target_species "human mastadenovirus a, human mastadenovirus b, human mastadenovirus e" --discard UNVERIFIED

Run derep using virus mode and multiple processes:

repgenr derep --workdir mastadenovirus_test --virus --process_size 100 --num_processes 4

The virus dereplication workflow use the following settings "optimized" for viral genomes: https://drep.readthedocs.io/en/latest/choosing_parameters.html#comparing-and-dereplicating-non-bacterial-genomes

This strategy perform all-vs-all pairwise alignments and yield multiple outputs, for every pairwise alignment. This results in a very large number of files created as the number of genomes increase. It is recommended to use multiple processes if the total number of genomes exceed 200. If a large number of genomes are executed as a single process then be warned that the computer system will spend a lot of time to handle I/O and directory structures.

Run phylo, no virus-specific settings applied:

repgenr phylo --workdir mastadenovirus_test --mode accurate

A note on segmented viruses: Lassa example

Download of segmented viruses may require more effort from the user to select the desired sequences and is due to the algorithm not being designed to handle multiple genome segments in different files. Example using Lassa below:

Download fasta with Mammarenavirus Lassaense (full segments have "complete sequence" as identifiers instead of "complete genome"):

repgenr vmetadata --workdir bunyavirales --target bunyavirales --filter "complete sequence"

Parse Mammarenavirus lassaense sequences:

# "Large" segment:
repgenr vgenome --workdir bunyavirales --target_species "Mammarenavirus lassaense" --length_range 7000-7400

# "Small" segment:
repgenr vgenome --workdir bunyavirales --target_species "Mammarenavirus lassaense" --length_range 3200-3600

# Partial sequences, any segment, not determining an outgroup, and listing parsed fasta-headers in terminal:
repgenr vgenome --workdir bunyavirales --target_species "Mammarenavirus lassaense" --length_range 0-9999999 --no_outgroup --print_fasta_headers

Clone this wiki locally