Skip to content
Discussion options

You must be logged in to vote

During training, there are typically two topic representations created, a c-TF-IDF vector per topic and topic embeddings. The former will always be created and the latter only when the embedding model is passed to the BERTopic model.

When topics are reduced, they are reduced based on one of those representations by searching which topic representations, c-TF-IDF or embeddings, can form a cluster and therefore be merged together.

In the case you mentioned above, the c-TF-IDF vectors are used instead of the topic embeddings to perform topic reduction.

Replies: 1 comment 1 reply

Comment options

You must be logged in to vote
1 reply
@johnsonice
Comment options

Answer selected by johnsonice
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants