Skip to content

Analysis

Noelle Patterson edited this page Apr 29, 2020 · 12 revisions

The final database of each paper and its corresponding topic probabilities and country of origin was analyzed in several ways to identify temporal and spatial patterns of bright and blind spots of water resources research in Latin America and the Caribbean.

Timeline

A timeline was created with the number of new research articles per year representing countries in the three socio-hydrologic clusters, to visualize the growth in research output over time. The articles were sourced from the English corpus of water resources research, because these articles were labeled with the country of research and could therefore be associated with a socio-hydrologic cluster. The cumulative research growth showed an exponential trend, although the exponential trend was more evident in Cluster 1 than in Clusters 2 or 3. To better understand trends observed in each socio-hydrologic cluster, a residual analysis was performed. First, the data were transformed with a logarithmic transformation to obtain a roughly linear relationship between time and research output, and then a linear regression was calculated. The residuals for each year were plotted and displayed starting with the year 2000 until the end of the period of record in 2017. Year 2000 was chosen as the starting point because it marks the time by which research output had increased enough to reach at least 30 new articles in each socio-hydrologic cluster per year. The residuals were then plotted along with brackets of the standard deviation (for both positive and negative values) to provide a reference of significance. These results were used to interpret periods of time in which research growth was increasing or decreasing on a region-wide scale, in reference to the growth trajectories of each socio-hydrologic cluster.

Normality

The normality of research topics was plotted for a subset of topics that represent components of the water budget or methodologies. The documents in each subset were sourced from the English corpus. Each subset was analyzed identically, as described here. The subset was filtered for documents that were labeled with a country of research, and then for countries where the sum of documents per country was greater than 30. The distance from standard, normal distribution was calculated to describe the normality of topics from two perspectives: across documents and across documents. The subset was then transformed in 2 ways. To calculate normality across documents, the probabilities were grouped by country and topic and normalized. To calculate normality across countries, the probabilities of each topic per country were aggregated and summed. Total probabilities were grouped by country to calculate the proportion of research each country devotes to each topic, then grouped by topic and normalized.

For each topic, two graphs were produced corresponding to each transformation. Each graph displayed the density distribution of the normalized probabilities and a standard, normal density distribution of 512 observations with a standard deviation of 1. The graph was rendered to obtain a numerical description of the density distributions. The y values of the rendered plots were normalized and used to calculate a distance matrix based on the Jensen-Shannon Divergence with equal weights. The distance matrix was used as the final input to the calculation that described normality: 1 minus the square root of the distance matrix.

This analysis resulted in two values for each topic, which were graphed with normality across documents on the x-axis and the normality across countries on the y-axis.

Chord diagram

Hexagonal entropy map

The entropy across research topics was mapped by country according to three levels of research topics: general, specific, and those corresponding to a component of the water budget. Probabilities were grouped by topic and country and summed. Total probabilities were subset by country and used to calculate entropy of research output for each country.

Citation network

LDA performance

The performance of the LDA was assessed by comparing the topic model output from the individual models corresponding the English, Spanish and Portuguese corpora. The topics from each corpus were tagged with 5 human-derived labels to describe the word distribution corresponding to each topic. The specific level tags were compare the corpora. The number of topics with each specific label were grouped by language and tallied.

Clone this wiki locally