Skip to content
Discussion options

You must be logged in to vote

The probabilities that are generated within BERTopic are a result of the soft clustering capabilities of HDBSCAN. More specifically, these probabilities are generated after finding the clusters and are exactly an approximation of the probabilities, not the process by which the clusters are created. In other words, the probabilities that are calculated by HDBSCAN are an approximation of its fitting process and will not exactly match the actual clusters that are created.

Lastly, the output of .reduce_outliers is done with a different process than how the probabilities are calculated, so updating the probabilities with the output of .reduce_outliers is not possible since they are created fro…

Replies: 1 comment 2 replies

Comment options

You must be logged in to vote
2 replies
@jermainkaminski
Comment options

@MaartenGr
Comment options

Answer selected by jermainkaminski
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants