-
Notifications
You must be signed in to change notification settings - Fork 0
Example queries
Please be aware that the data repository and interface are undergoing continuous improvements and new data ingestion. Therefore the examples hereby provided may at times provide outdated information, counts or figures. We apologize in advance for such misalignments, which in any case should not compromise the clarity of the explanation.
Select sequences from GenBank database that are complete. These are currently 14,260.

Select sequences that have been produced using the high quality Illumina technology with a coverage higher then 1000X. These are currently 1003.

ViruSurf also contains COG-UK data. The last update made publicly available is of May 8th. The interface can be used to filter only sequences that were collected during May 2020, which are now 6,351.

(Tang et al., 2020) claim that there are two clearly definable major types (S and L) of SARS-CoV2 in this outbreak, that can be differentiated by transmission rates. The S and L types can be clearly distinguished by just two tightly linked variants at positions 8782 (within the ORF1ab gene from T to C) and 28144 (within ORF8 gene from C to T, resulting into a change from Serine to Leucine at the 84 position of the ORF8 protein). (MacLean et al., 2020) later discredited the claims made in (Tang et al., 2020) stating the difficulty in demonstrating the existence or nature of a functional effect of a viral mutation. ViruSurf can be used to perform a query to isolate the "S" type, based on the amino acid change in position 84 of gene ORF8.
gene_name: ["orf8"], sequence_aa_original: ["l"], sequence_aa_alternative: ["s"], aa_position: {"min_val":84,"max_val":84}

(Tang et al., 2020) also suggests that patients that have the genotype Y (C or T) at both positions 8782 and 28144 (differing from the general trend of having respectively C and T) could have been infected with viruses from multiple strains of different types.
"sequence_original": ["c"], "sequence_alternative": ["y"], "var_position": {"min_val": 8782, "max_val": 8782} AND "sequence_original": ["t"], "sequence_alternative": ["y"], "var_position": {"min_val": 8782, "max_val": 28144}

To support SARS-CoV2 vaccine design efforts, it is useful to track antigenic diversity. Typically, pathogen genetic diversity is categorized into distinct clades (i.e., a monophyletic group on a phylogenetic tree). In (Gudbjartsson et al., 2020), specific sequence variants are used to define clades/haplogroups (e.g., the “A3 group” is characterized by the 11083 and 29742 nucleotides, originally G mutated to T, by the 1397 nucleotide G mutated to A, and by the 28688, from T to C). Specific definition of all groups is given in Table S3 in the paper Appendix.
The proposed study was performed on residents of Iceland. However, ViruSurf supports all the information required to replicate the definition of SARS-CoV2 clades proposed in the study on public sequences from GenBank.

Researchers at Scripps Research, Florida, found that the mutation, known as D614G (meaning that Aspartic acid mutated to Glycine in position 614 of spyke protein), stabilized the virus’s spike proteins, which protrude from the viral surface and give the coronavirus its name. The number of functional and intact spikes on each viral particle was about five times higher because of this mutation, they found. These spike proteins must attach to a cell for a virus to infect it. As a result, the viruses with D614G were far more likely to infect a cell than viruses without that mutation. The new study has already been spread by mass communication means, but the scientific manuscript (Zhang et al., 2020) has not been peer reviewed yet.
This mutation was previously detected to increase with an alarming speed. At the time of writing their manuscript, they observed that the G614 genotype was not detected in February and observed at low frequency in March (26%), but increased rapidly by April (65%) and May (70%), indicating a transmission advantage over viruses with D614. We use Virusurf to show a similar trend.
When we select this mutation in GenBank sequences collected within the end of March 2020, we the following results:

In total sequences with the mutations are 58.6% of the total.
Instead, when we select this mutation in GenBank sequences collected from the start of April 2020, we obtain the following results:

In total sequences with the mutations are 87.7% of the total.
In addition to L84S, a G-T transversion at 26144 that causes an amino acid change in orf3 protein (G251V) is also being investigated (see (Chaw et al., 2020)) The paper claims that 251 V was first seen on 1/22/2020 and rapidly increased its frequency within 2 weeks. We replicated the search with Virusurf and found that before 1/22/2020 already 3 complete sequences (now available from GenBank) where collected with such mutation.

(Pachetti et al., 2020) report about a mutation located in SARS-CoV2 gene N at position 28881, which is related to a double codon mutation, inducing the substitution of two amino acids, namely 28881 (R to K) and (G to R).
We reproduce this on ViruSurf by first looking for sequences that have two alternative amino acid changes in gene N (16 sequences found) and then filtering only the sequences that have a nucleotide variation specifically at position 28881.
