Skip to content

2. Methods

Aleksandar Anžel edited this page Sep 4, 2021 · 15 revisions

Depending on the data set type, different methods are used to explore the underlying information.

1. Preprocessing

Archived data

Archived data is uploaded as a BytesIO stream. That stream is then written to a file that represents a 1 to 1 copy of the originally uploaded data set. That archive is then extracted into the Data/uploaded/omic_name folder for further use. The next step that is shared among every archived data set is the addition of temporality. If filenames are starting with D or W, then a user is confronted with an input box where a start date should be entered. After inputting the start date, every file name is converted to the ISO 8601 format (yyyy-mm-dd). If filenames are already according to this format, the previous step is skipped. If there is an error during this process, an error message is shown to the user.

FASTA files

If an archived data set contains FASTA files (extensions .fa or .faa), the user is presented with an option for creating an additional data set with physico-chemical properties. These properties are different for genomics and proteomics data, as can be seen below:

  • Genomics: Hydrogen bond, Stacking energy, Solvation (see paper Physico-chemical fingerprinting of RNA genes)
  • Proteomics: Molecular weight, Gravy, Aromaticity, Instability index, Isoelectric point, Secondary structure fraction of helix, Secondary structure fraction of turn, Secondary structure fraction of sheet, Electricity, Fraction aliphatic, Fraction uncharged polar, Fraction polar, Fraction hydrophobic, Fraction positive, Fraction sulfur, Fraction negative, Fraction amide, Fraction alcohol

Since each FASTA file can contain multiple sequences, these physico-chemical values are averaged so that each file has a fixed set of physico-chemical values. When these values are calculated for all files, temporality is added to the newly-created data set using file names.

The next step of preprocessing FASTA files is embedding those files (and sequences within) into 100-dimensional vectors. The length of those vectors was chosen empirically. Word2Vec model is used to embed varying-length sequences into fixed-size vectors. Embedding FASTA files for the first time trains the model on those sequences and then embeds them, which can take a long time depending on the number of sequences and FASTA file. After the first run, model weights are cached as well as the embeddings, so there are no unnecessary re-runs. A vector corresponding to one FASTA file represents an average vector of all sequences contained in that FASTA file. That means MOVIS is robust to the number of sequences in each FASTA file. When the whole archived data set (holding N FASTA files) is processed, the end result is a Nx101-dimensional matrix containing one 100-dimensional vector per FASTA file plus an additional dimension holding the temporal information of that file.

KEGG annotation files

As stated earlier in this Wiki, KEGG files must have a specific header in order to be processed correctly. If that is not the case, an error message is shown, and all computations are stopped. After importing a valid KEGG annotation file, the count matrix is created for each KEGG file name. If there are N KEGG files, then the resulting matrix is Nx(M+1), where M is the number of distinct KO annotations present in the whole archive. One additional dimension is the temporal dimension computed from the file names. Values inside that matrix represent the number of occurrences of the specific KO annotation inside one KEGG file.

Beware: The resulting matrix created from KEGG annotation files can be big in size.

BIN annotation files

MOVIS handles these files by saving only product information from them. That means that every annotation file must have a part where the products (i.e., the functions) of the corresponding FASTA file sequences are stated. We intend to extend the functionality of MOVIS to use even more meaningful data from these files. When different products are imported for each BIN file, they are sorted in descending order with respect to the number of occurrences of each product inside one BIN file. After that, only the TOP 10 products are saved for each BIN file. This is done for easier visualization of those products through time. The result is a Nx(10+1) matrix, where N is the number of BIN files, 10 is the number of columns representing the TOP 10 products, and 1 is an additional temporal column. Each cell of the matrix holds the number of occurrences of the specific product inside one BIN file.

Tabular data

The unknown values are removed immediately after importing a tabular data set. More information about the reasoning behind this decision is in section 1. of this Wiki. Since certain tabular data sets can contain special characters within the column names, MOVIS fixes those columns by replacing each invalid character with a corresponding valid character.

If a multi-tabular data set is uploaded, the previous procedure is applied to each data set individually. Then, since each data set must have the same set of columns (features), a check is run to validate this. If that check is successful, all data sets are merged into one, with one additional column named Type that contains the file names of each corresponding data set. The resulting data set is then treated as a new data set, so all of the following procedures are shared to all tabular data sets.

The next step of MOVIS for curating the uploaded data set is to fix any value that contains , instead of . as a decimal point. After that, MOVIS locates the temporal column in the data set and saves the list of all other columns for further processing. If the temporal column is not located, an error is shown to the user.

The last step in the data set curation is the additional modification of the uploaded data set. MOVIS currently supports the following types of modifications:

  1. Saving only a certain time interval and removing all rows outside of the chosen interval.
  2. Removing one or more rows from a tabular data set.
  3. Removing one or more columns from a tabular data set.

After this step, a data set is considered curated and ready for visualization tasks.

2. Data analysis

Archived data

FASTA files

KEGG annotation files

BIN annotation files

Tabular data

Multi-tabular data

Clone this wiki locally