One potential issue when computing distances over structural features of language is their non-independence. There's different ways of tackling this, from using a dimensionality reduction technique (e.g. PCA, MCA, FAMD) and computing distances of the x number of dimensions that explain Y% of the data to Chu-Liu-pruning (Chu & Liu 1965, Hammarström & O'Connor 2013) etc. (You can see an example of the Chu-Liu approach on Grambank data in my thesis if you're curious :) (Skirgård 2021).)
For URIEL+, I would suggest considering importing GBI Logical from our recent paper and project on feature dependency in Grambank, World Atlas of Language Structures (WALS), AUTOTYP, PHOIBLE and Lexibank (Graff et al 2025). We have an R function in rgrambank for computing GBI datasets from Grambank - make_GBI I can help you with this, if you'd like?
Refs:
Chu, Yoeng-Jin & Tseng-Hong Liu. 1965. On the Shortest Arborescence of a Directed Graph. Scientia Sinica 14. 1396–1400.
Graff, A., Chousou-Polydouri, N., Inman, D., Skirgård, H., Lischka, M., Zakharko, T., Barbieri, C., & Bickel, B. (2025). Curating global datasets of structural linguistic features for independence. Scientific data, 12(1), 106.
Hammarström, H., & O’Connor, L. (2013). Dependency-sensitive typological distance. Approaches to measuring linguistic differences, 329-352.
Skirgård, H. (2021). Multilevel dynamics of language diversity in Oceania. PhD DissertationCanberra: Australian National University. https://openresearch-repository.anu.edu.au/handle/1885/218982
One potential issue when computing distances over structural features of language is their non-independence. There's different ways of tackling this, from using a dimensionality reduction technique (e.g. PCA, MCA, FAMD) and computing distances of the x number of dimensions that explain Y% of the data to Chu-Liu-pruning (Chu & Liu 1965, Hammarström & O'Connor 2013) etc. (You can see an example of the Chu-Liu approach on Grambank data in my thesis if you're curious :) (Skirgård 2021).)
For URIEL+, I would suggest considering importing GBI Logical from our recent paper and project on feature dependency in Grambank, World Atlas of Language Structures (WALS), AUTOTYP, PHOIBLE and Lexibank (Graff et al 2025). We have an R function in rgrambank for computing GBI datasets from Grambank - make_GBI I can help you with this, if you'd like?
Refs:
Chu, Yoeng-Jin & Tseng-Hong Liu. 1965. On the Shortest Arborescence of a Directed Graph. Scientia Sinica 14. 1396–1400.
Graff, A., Chousou-Polydouri, N., Inman, D., Skirgård, H., Lischka, M., Zakharko, T., Barbieri, C., & Bickel, B. (2025). Curating global datasets of structural linguistic features for independence. Scientific data, 12(1), 106.
Hammarström, H., & O’Connor, L. (2013). Dependency-sensitive typological distance. Approaches to measuring linguistic differences, 329-352.
Skirgård, H. (2021). Multilevel dynamics of language diversity in Oceania. PhD DissertationCanberra: Australian National University. https://openresearch-repository.anu.edu.au/handle/1885/218982