Repository navigation
Ch 15 Clustering - additional method suggestion #862
Replies: 1 comment
|
Thanks, yes we should have a look at this. In general I would recommend clustering data in the original feature space because projecting to few principal dimensions will potentially discard very relevant information, and lead to clustering that models artefacts in the projection (depending on clustering method), rather than real clusters in the actual data. Projections are simplifactions and modeling projections is not the same than modeling the real data. However, ordination can be useful for denoising and focusing on the main features. Then one would often preserve something like 90% of the information (and relatively many principal axes, instead of leading 2-3). Denoising can sometimes lead to better analyses. If we add something on this it should be accompanied by referenced discussion about this. |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I was looking over Ch 15 - Community typing, and thought it might be worth adding a clustering example based on a paper I looked at recently where dimensionality reduction, like MDS, is conducted prior to a distance-based clustering method. This method and R package incorporated the use of compositionally aware distance metrics like UniFrac which made a lot of sense given the data I normally work with.
Being a relatively new mia user and trying to get my head around how to work with data in a TSE structure, I'm loving these workflow examples, a massive help. Great work guys.
R package is MDSMClust: https://github.com/wxy929/MDS
Paper: Multidimensional scaling improves distance-based clustering for microbiome data (https://doi.org/10.1093/bioinformatics/btaf042)
Cheers
Rachele
All reactions