-
Notifications
You must be signed in to change notification settings - Fork 0
Analysis
The final database of each paper and its corresponding topic probabilities and country of origin was analyzed in several ways to identify temporal and spatial patterns of bright and blind spots of water resources research in Latin America and the Caribbean.
A timeline was created with the number of new research articles per year representing countries in the three socio-hydrologic clusters, to visualize the growth in research output over time. The articles were sourced from the English corpus of water resources research, because these articles were labeled with the country of research and could therefore be associated with a socio-hydrologic cluster. The cumulative research growth showed an exponential trend, although the exponential trend was more evident in Cluster 1 than in Clusters 2 or 3. To better understand trends observed in each socio-hydrologic cluster, a residual analysis was performed. First, the data were transformed with a logarithmic transformation to obtain a roughly linear relationship between time and research output, and then a linear regression was calculated. The residuals for each year were plotted and displayed starting with the year 2000 until the end of the period of record in 2017. Year 2000 was chosen as the starting point because it marks the time by which research output had increased enough to reach at least 30 new articles in each socio-hydrologic cluster per year. The residuals were then plotted along with brackets of the standard deviation (for both positive and negative values) to provide a reference of significance. These results were used to interpret periods of time in which research growth was increasing or decreasing on a region-wide scale, in reference to the growth trajectories of each socio-hydrologic cluster.
The normality of research topics was plotted for a subset of topics that represent components of the water budget or methodologies. The documents in each subset were sourced from the English corpus. Each subset was analyzed identically, as described here. The subset was filtered for documents that were labeled with a country of research, and then for countries where the sum of documents per country was greater than 30. The distance from standard, normal distribution was calculated to describe the normality of topics from two perspectives: across documents and across documents. The subset was then transformed in 2 ways. To calculate normality across documents, the probabilities were grouped by country and topic and normalized. To calculate normality across countries, the probabilities of each topic per country were aggregated and summed. Total probabilities were grouped by country to calculate the proportion of research each country devotes to each topic, then grouped by topic and normalized.
For each topic, two graphs were produced corresponding to each transformation. Each graph displayed the density distribution of the normalized probabilities and a standard, normal density distribution of 512 observations with a standard deviation of 1. The graph was rendered to obtain a numerical description of the density distributions. The y values of the rendered plots were normalized and used to calculate a distance matrix based on the Jensen-Shannon Divergence with equal weights. The distance matrix was used as the final input to the calculation that described normality: 1 minus the square root of the distance matrix.
This analysis resulted in two values for each topic, which were graphed with normality across documents on the x-axis and the normality across countries on the y-axis.
The entropy across research topics was mapped by country according to three levels of research topics: general, specific, and those corresponding to a component of the water budget. Probabilities were grouped by topic and country and summed. Total probabilities were subset by country and used to calculate entropy of research output for each country.
We conducted a network-analysis using Gephi 0.9.2 (Bastian et al 2009, https:// gephi.org/). Descriptive parameters and geometry metrics were calculated. We conducted this network analysis at the country and three levels of research topics: general, specific, and those corresponding to a component of the water budget. Probability adjacency matrices were extracted from predictive model results and it was normalized to highlight the relationship between the topics or countries regardless of how much of it we sampled in our corpus. The geometric description was calculated: the number of nodes (countries or research topics), number of edges (citation between countries or research topics), and thickness of the edges (connectivity proportion). We calculated the followed network metrics: Degree range, it is the number of connections of each node (country or topic) has with another node (country or topic). The degree has generally been extended to the sum of weights when analyzing weighted networks and labeled node strength, so the weighted degree and the weighted in- and out-degree was calculated (Barrat et al. 2004, Newman 2001, Opsahl et al. 2010). Network density is a measure of the connectedness of a graph, defined as the number of connections, divided by the number of possible connections. How close the network is to complete, with all possible edges and density =1 (Tonin et al. 2019). Network node size, it represents the research volume in the corpus defined as the sum of probabilities in that topic. A force-directed graph algorithm was selected, Fruchterman-Reingold to produce the network graphs visualization. It simulates the graph as a system of mass particles. The nodes are the mass particles and the edges are springs between the particles. The algorithms try to minimize the energy of this physical system. FR is most suitable for small networks and better performance (Fruchterman, & Reingold 1991, Jacomy et al 2014). A directed network was produced indicating a directionality between the nodes and edges are parallel to each other and automatic merge. Two nodes (countries or topics) are considered connected if they have a probability of citation among them. This network analyses captured the variation in the strength of connectivity-based on the percent of citation by countries or topics and it-self-citation. Force-directed citation network, each node corresponds to each country or water research topic in LAC, and country node colors are groping by Socio-hydrological cluster and NSF specific topics nodes colors are groping by NSF general colors. Showing the degree of connectivity by the nodes size and edge thickness
Four network graphs were produced. The network graphs showed a density value of 1, every topic and countries were connected. Country citation network identified 23 nodes and 529 edges. Country network degree of connectivity has a maximum degree (<50%), medium degree(<30%), and minimum degree less than (<10%). While NSF general topics identified 5 nodes and 25 edges, degree of connectivity from 45% maximum to < 27% minimum. NSF specific topics have 43 nodes and 1849 edges, maximum degree (18%), medium degree (10%), and a minimum degree of less than 5%. The water budget-topic network has 27 nodes and 729 edges. This network showed the smallest degree of connectivity with a maximum degree(less than 5%).
The performance of the LDA was assessed by comparing the topic model output from the individual models corresponding the English, Spanish and Portuguese corpora. The topics from each corpus were tagged with 5 human-derived labels to describe the word distribution corresponding to each topic. The specific level tags were compare the corpora. The number of topics with each specific label were grouped by language and tallied.