Skip to content

analysis_multivariate

SlimaniF edited this page Sep 1, 2026 · 8 revisions

This section presents multivariate analysis pipelines that can be performed after quantification. Exploratory routine is designed to unveil structure in rna expression (single molecule number) and in co-localization. This analysis is based on Multivariate Data Analysis by Hair and al.

1. Exploratory analysis
1.1 About independent techniques
1.2 Rna expression analysis
1.3 Co-localization analysis
[2. Analysis parameters]

Exploratory analysis

What I call exploratory analysis is usually refered to as indepedent techniques in multivariate analysis. They are designed to find underlying structure in data.
This pipeline focuses on Principal component analysis (PCA) and on Multiple correspondence analysis (MCA). They are analogue techniques that identify structure in data using dimension reduction.
When looking at indepedent multivariate analysis results keep in mind that the analysis will always yield some sort of structure and should be use as exploratory results to create or investigate hypothesis not as tools to establish causality relations. Some indicators are computed to help you assess the relevance of the structures found.

About indepedent techniques

Multivariate techniques are designed to unveil structural relationship between more than two variates as oposed to direct correlation analysis. To that end we different metrics can be studied (i.e variance or mean) through dimension reduction techniques. Dimensions (RNA) are grouped to explain metric (variance) statistic by a composition of RNA groups. The resulting groups (dimensions) contribute in a different proportion to the global metric, in the case of PCA this is materialised by their eigenvalue. For a sample with $N$ rna distributions we should always be able to find $N-1$ dimensions. However it does not mean that they all contribute equally to the variate, most of the time lower eigenvalues dimensions are discarded are they explain noise in the data. There are a few methods to retain an appropriate number of dimension that are detailed below in the PCA section.

Different techniques were designed to explore different aspect of the data. PCA technique aims at transforming correlated variables into, new, uncorrelated variables. This variables are named principal components and define dimensions that explain best data variability. PCA relies on normality asumption and on the use of metrics following a scale which makes it improper for co-localization analysis where data is categorical (even binary). MCA comes in as an analogue technique compatible with categorical, it is build as an extension of Correspondance Analysis but the interpretation logic of its results is similar to one of PCA results. PCA focuses on explaining data through variance the analogue, in other words it finds the structure that explain the maximum of variance; the analogue value for MCA is called inertia. Intertia is a measure of how much co-occurence pattern in the data deviates from independence which makes it an intuitive indicator for co-localization patterns.

The described methods below are meant to help you find or check patterns in your data using Sequential Fish built-in pipelines.

RNA expression analysis

In this pipeline we investigate reliationship between single molecule number in cells.

Correlation matrix

The correlation matrix is formed by computing Pearson's r coefficient to study direct correlation between RNAs. This is not multivariate analysis but it is a preliminary step to dertime if multivariate analysis is relevant. Indeed a dataset explain by direct correlation will be overly dominated by direct correlations. In such case correlated variates should be grouped beforehand or one should be eliminated to assess such as case one can refer to partial correlations.

Sanity metrics

  • Correlation matrix determinant : A correlation matrix determinant close to 0 (i.e < 10^-5) will result in mathematical instabilities when running PCA in such a case user should proceed with caution.

  • Bartlett sphericity test : This indicator is a pvalue to test that correlation matrix is the identity matrix in such a case that the null hypothesis is not rejected RNAs expression are perfectly orthogonal and no underlying behavior can be found.

  • Kaiser–Meyer–Olkin test (KMO) : This test sample adequacy to be interpreted throught PCA a value below 0.7 reveals critical multicolinearity within sample making dimension reduction inadequate. Sample with KMO > 0.7 are acceptable and >0.8 are interesting to investigate.

About the use of 'foci rnas' option
Using this option user can split a population in its free and clustered part. This two new population will be added to the existing population meaning that population_A can coexist with population_A_clustered and population_A_free. Needless to say this will create artificial correlation and colinearity in the data when running PCA make sure to remvoe your 'focis rnas' option or to filter out the main population using the filter rna population.

Number of dimension

This plot aims at help you consider which dimensions are relevant for analysis and which contribute mostly to noise. There are three ways of selecting an appropriate number of dimensions for the PCA analysis. First is called Kaiser criterion and states that one should retain all dimensions whose eigenvalue is superior to 1. Second is to draw the eigenvalues in an ordinated manner (scree plot) and consider the elbow point of the curve as the threshold for selecting dimensions. Third is to retain as many dimensions as to explain 80% of the data variance this can be visualized in the pie plot.

During this automated analysis we use the Kaiser criterion.

New dimensions and loadings

PCA yields new orthogonal dimensions explaining a part of total variance. On each of these dimensions rnas contribute differently. The contribution of an rna to a dimension is called loadings and is the heart of the analysis. When a group of rnas loads heavily on a dimension it means that part of the sample variance is explained better by the group than by the individuals. The loadings plot allow us to identify the principal components of the new dimensions and thus define and underlying structure of the data.

The results must be interpreted with caution and be used as guides to create or strengthen hypothesis as they are not proving cause-effects relationship, only structural behavior. What is more every data set will yield an underlying structure even if in poorly correlated data this structure will only explain noise.

Noise can random because your data is poorly correlated and there is no interesting underlying structure to study, or coming from biological processes that are not of interest to you (cell cycle, cell stress...) and inducing a correlated behavior to your sample. To account for biological noise we propose a few extra tools : partial correlation and normalisation using control genes.

Partial correlation

Partial correlation isolates the direct linear relationship between two population. It can be used as an indicator on validity to use PCA analysis and on controlling for irrelevant biological effects. The validity of PCA analysis is already controlled using the KMO test, see above, which is based on partial correlation.

In the non-normalised figure, when partial correlation is ran, each pair uses all the other genes as covariates. Meaning we look at the direct linear relation between the pair after removing the effect of all other population. Obtaining very high values for a pair at this point suggest that their expression is explained only by each other and there are no structural effect to uncover. In such a case you can either group them (mean or median rule...) or remove one from analysis so that their correlation effect doesn't dominate the structural analyis.

In the normalised figure partial correlations is computed using control genes as covariates. In this case finding low partial correlations (i.e < 0.3) suggest either that your data set is poorly correlated and there is no structural effect to uncover or that the structure is dominated by noise that is captured is the control genes. In those cases normalisation using control genes is relevant. If partial correlations are strong (i.e > 0.5) PCA is already valid and not driven by the noise captured by the control genes, normalisation is not relevant.

Co-localization analysis

Although this approach was initialy implemented for rna expression analysis I try to adapt it to co-localization. The pipeline is based on the colocalization truth table also computed in the colocalization section submodule. This table indicate for each spot (rows) True in the columns Population_A if the spot co-localizes with population A else False. The column corresponding to the spot population is always True. Principal component analysis requiring normal metrics, the usage of boolean data constrains to switch to Multiple correspondence analysis (MCA) an extension of correspondance analysis with an interpretation analogue to PCA.

Multiple correspondence analysis

Multiple correspondence analysis (MCA) was chosen as an analogue technique of PCA for usable for boolean data. Refer to above section for a thorough insight. Instead of maximizing variance when reducing dimensions MCA aims at maximizing inertia. In this context inertia is defined as how much data deviates from complete independance.

Scree plot and loadings

The scree plot helps you retain the best number of dimensions in your analysis. Similarly to PCA, coordinates x-axis is showing dimension but instead of eigenvalues, on y-axis is plotted contribution to intertia as a fraction. In MCA we usually keep dimensions contribution to more than 10% to global inertia, one can also use the elbow point rule.

Loadings are replaced with categories coordinates (1 category = 1 reduced dimension) in reduced space. While we use them to understand resulting structure unveil by MCA their interpretation is quite different. Firstly, note that vectors point from origin to categories are not unitary vectors their squared sum equals to the inertia contribution of this dimension. What is more the scaling is done in such a way that the distance between categories reflects their dissimilarity in co-occurence patterns. In other words, proximity between categories indicates similar co-occurence patterns and distance from origin indicates strengthness in contributing to intertia. However given the number of population this can be difficult to represent on a scatter plot. To envision this information we propose the use of a heatmap indicating structure of reduced dimensions combined with a bipartite network graph.

The network graph is composed of two types of nodes :

  • Rnas nodes round : Size is proportionate to rna abundancy, their are linked to a category node if they contribute to more than 2% (threshold can be modfied in parameters) to global inertia through a reduced dimension. The edge sizes are proportionate to coordinate value.

  • Categories nodes diamond : Size is proportionate to contribution to global inertia. They are fully connected to each other through a weighted edge representing their proximity, more precisely, chi-squared distance is used (natural metric of MCA).

The network is display in a spring-force layout allowing similar categories to be close and dissimilar categories to be far away.

Hierachical Clustering

Then final indicator of this multivariate pipeline is based on hierarchical clustering which is a clustering technique. Controray to MCA, that looks for structure or patterns of co-occurence this will focus on actual co-localization frequencies ignoring underlying structure. This sequentialy group pairs of rna that have a high number of co-localization events this will favorise highly expressed populations if they show co-localization. Is more meant to identify groups of rnas that have the potential of making a significant biological interaction rather than looking for group of rnas that are driven by a similar effect to co-localize. Hierarchical clustering can be interpreted as a sort of ranking of colocalizations event and is represented trought a dendogram. The sequentialy mechanic of this clustering will merge groups in an order based on their similarity (co-localization frequencies) and ends up by forming a group containing all rnas. The similarity between groups (x-axis) is computed using the Jaccard distance.

Analysis parameters

Control genes (list) : List of names that were imaged as control genes in your experiment. This allow the creation of the normalised figure for the rna expression analysis.

Clone this wiki locally