Skip to content

Results visualization

Anna Bernasconi edited this page Sep 23, 2020 · 13 revisions

Every choice in the Metadata search and the Variant search sections impacts the visualization of the results at the bottom of the page.

In the Metadata search section, to trigger changes in the submitted query, it is sufficient to change values in the drop-down menus. In other words, the results table is updated whenever the user either adds or removes a value by clicking on a drop-down menu element or types/deletes text directly in the text field).

In the Variant search section, to trigger changes in the submitted query, the user needs to press , which consequently registers the choice.

At the bottom of the page we provide a results section, i.e., a table that describes the sequences resulting from the selections of the user.

As an example, using the query is_complete: ['false'], gc_percentage: {min_val: 35, max_val: null, is_null: false}, region: [' connecticut', ' new york'], the obtained result is:

result_table

The count of selected sequences (in the example, 68) is shown in the bottom right corner of the page. In the bottom part of the table the user can select how many rows should be visible in the page. The default is 10, other options are 100 and 1000. Following/previous pages can be visualized using the left/right arrows.

The button allows to re-organize the table visualization by selecting/deselecting fields to be visualized or hidden and by drag-and-dropping the attributes in different table columns. Some fields are not available by default, but are selectable through the button. For each sequence, the complete set of provided fields (when available) follows:

  • Source Page: link (or more links) to view the page related to the sequence on the original source
  • Accession ID: Sequence unique identifier, from original source database
  • Strain: Virus strain name (sometimes hard-coding relevant information such as the species, collection location and date)
  • Reference: True when the sequence is the reference one (from RefSeq) for the virus species, False when the sequence is not the reference one
  • Complete: True when the sequence is complete, False when the sequence is partial
  • Strand: Strand to which the sequence belongs to (either positive or negative)
  • Length: Number of nucleotides of the sequence
  • GC%: Percentage of read G and C bases
  • N%: Percentage of unknown bases
  • Lineage(Clade): Percentage of unknown bases
  • Seq. Technology: Platform used for the sequencing experiment
  • Assembly Method: Algorithms applied to obtain the final sequence (e.g., for reads assembly, reads alignment, variant calling)
  • Coverage: Number of unique reads that include a specific nucleotide in the reconstructed sequence
  • Seq. Lab: Laboratory that sequenced and submitted the sequence to the databank (encoded by 'Database source')
  • Submission date: Date of submission of the sequence to the databank (encoded by 'Database source')
  • BioProject: External reference to the NCBI BioProject database https://www.ncbi.nlm.nih.gov/bioproject/
  • Database: Original database from which information is collected
  • Host Species: Host organism species name from NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Host Species Taxon ID: Host organism species ID from NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Host Gender: Host organism gender (when applicable)
  • Host Age: Host organism age (in years, when applicable)
  • Collection date: Date in which the infected biological sample was collected
  • Isolation source: Tissue from which the infected biological sample was collected
  • Origin Lab: Laboratory that sampled the biological material from the isolation source
  • Country: Country where the biological sample was collected
  • Region: Region (i.e., part of country) where the biological sample was collected
  • GeoGroup: Continent where the biological sample was collected
  • Virus Name: Virus name as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Virus TaxonID: Virus numerical id as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Family: Virus family as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Sub Family: Virus sub-family as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Genus: Virus genus as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Species: Virus species name as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • Equivalent names: Virus alternative names as to the NCBI Taxonomy https://www.ncbi.nlm.nih.gov/taxonomy
  • MoleculeType: Virus molecule type (DNA/RNA)
  • Single stranded: True if virus molecule is single stranded, false if it is double stranded
  • Positive stranded: In case the virus molecule is single stranded, true if virus molecule is positive stranded, false if it is negative stranded
  • Sequence: Full nucleotide sequence of virus

Note that the fields can be used to order the rows of the table. This can be enforced by clicking on the desired column name where an arrow, indicating the ascending (up) or descending (down) order appears by hovering. For example the ascending order by Source ID corresponds to the following

The button can be used to download in .csv format the entire results table corresponding to the performed query. The button can be used to download in .csv format or FASTA file only the sequences corresponding to the performed query.

It is possible to use a drop-down choice to calculate the specific nucleotide and amino acid sequences corresponding to a protein of the virus.

The default choice is "FULL" corresponding to a full nucleotide sequence, while the amino acid sequence column is left to N/D (not defined) When a protein is chosen, the button can be used to download the specific sequences (of nucleotides or amino acids) corresponding to the selected protein.

The "Show control" switch button allows to visualize the sequences of the control group, defined by those sequences selected by the Metadata search filters for which there exist some variant and the variant filters are not satisfied. This option, suggested to us by virologists, is the most sensible for describing the effects of variant analysis.

We show cases positioning the switch in this way and controls, positioning in this other way .

Clone this wiki locally