Skip to content

Package for clustering using UMAP and HDBSCAN (with numerical and categorical variables)

License

Notifications You must be signed in to change notification settings

aitorporcel/QuickClus

Repository files navigation

QuickClus

QuickClus is a Python module for clustering categorical and numerical data using UMAP and HDBSCAN. QuickClus allows incorporating numerical and categorical values (even with null values) into the clustering, in a simple and fast way. The imputation of null values, the scaling and transformation of numerical variables, and the combination of categorical variables are performed automatically.

Installation

python3 -m pip install QuickClus

Usage

QuickClus requires a Pandas dataframe as input, which may contain numeric, categorical, or both types of variables. In the case of null values, QuickClus takes care of the imputation and subsequent scaling of all the features. All this process is done automatically under the hood. It is also possible to automatically optimize the algorithm using optuna, calling tune_model(). Finally, QuickClus provides a summary of the characteristics of each cluster.

from quickclus import QuickClus
clf = QuickClus(
    umap_combine_method = "intersection_union_mapper",
)

clf.fit(df)

print(clf.hdbscan_.labels_)

clf.tune_model()

results = clf.assing_results(df)

clf.cluster_summary(results)

Documentation

https://quickclus.readthedocs.io/

Examples

Notebooks with examples of use

References

@article{mcinnes2018umap-software,
  title={UMAP: Uniform Manifold Approximation and Projection},
  author={McInnes, Leland and Healy, John and Saul, Nathaniel and Grossberger, Lukas},
  journal={The Journal of Open Source Software},
  volume={3},
  number={29},
  pages={861},
  year={2018}
}
@article{mcinnes2017hdbscan,
  title={hdbscan: Hierarchical density based clustering},
  author={McInnes, Leland and Healy, John and Astels, Steve},
  journal={The Journal of Open Source Software},
  volume={2},
  number={11},
  pages={205},
  year={2017}
}

About

Package for clustering using UMAP and HDBSCAN (with numerical and categorical variables)

Topics

Resources

License

Stars

Watchers

Forks

Releases

No releases published

Packages

No packages published