Skip to content

Analysis

ajdevincentis edited this page Apr 1, 2020 · 12 revisions

The final database of each paper and its corresponding topic probabilities and country of origin was analyzed in several ways to identify temporal and spatial patterns of bright and blind spots of water resources research in Latin America and the Caribbean.

Timeline

The number of new research articles each year was plotted for the three socio-hydrologic clusters to visualize the growth in research output over time. The articles were sourced from the English corpus of water resources research, because these articles were labeled with the country of research and could therefore be associated with a socio-hydrologic cluster. The cumulative research growth showed an exponential trend, although the exponential trend was more evident in Cluster 1 than in Clusters 2 or 3. To better understand trends observed in each socio-hydrologic cluster, a residual analysis was performed. First, the data were transformed with a logarithmic transformation to obtain a roughly linear relationship between time and research output, and then a linear regression was performed. The residuals for each year were plotted and displayed starting with the year 2000 until the end of period of record in 2017, because 2000 was the year in which research output had increased enough to reach at least 30 new articles in each socio-hydrologic cluster per year. The standard deviation of the residuals was plotted (for both positive and negative values) along with the residuals to provide a reference of significance. These results were used to interpret periods of time in which research growth was increasing or decreasing on a region-wide scale, in reference to the growth trajectories of each socio-hydrologic cluster.

Normality

The normality of research topics was plotted for a subset of topics that represent components of the water budget or methodologies. The documents in each subset were sourced from the English corpus. Each subset was analyzed identically, as described here. The subset was filtered for documents that were labeled with a country of research, and then for countries where the sum of documents per country was greater than 30. The distance from standard, normal distribution was calculated to describe the normality of topics from two perspectives: across documents and across documents. The subset was then transformed in 2 ways. To calculate normality across documents, the probabilities were grouped by country and topic and normalized. To calculate normality across countries, the probabilities of each topic per country were aggregated and summed. Total probabilities were grouped by country to calculate the proportion of research each country devotes to each topic, then grouped by topic and normalized.

For each topic, two graphs were produced corresponding to each transformation. Each graph displayed the density distribution of the normalized probabilities and a standard, normal density distribution of 512 observations with a standard deviation of 1. The graph was rendered to obtain a numerical description of the density distributions. The y values of the rendered plots were normalized and used to calculate a distance matrix based on the Jensen-Shannon Divergence with equal weights. The distance matrix was used as the final input to the calculation that described normality: 1 minus the square root of the distance matrix.

This analysis resulted in two values for each topic, which were graphed with normality across documents on the x-axis and the normality across countries on the y-axis.

Chord diagram

Hexagonal map

Citation network

LDA performance

Clone this wiki locally