Repository navigation
How topic reduction is done when an embedding model is not initialized (only given pre-calculated embeddings) ? #1033
|
I am trying to understand exactly how topic reduction was done. I tried to train a topic model with pre-calculated embeddings (from a sentence bert model ), without given a specific embedding model when setting up the topic model. I saved the trained model and latter reloaded it to evaluate topic quality. Then i realize that topic_model.topic_embeddings_ is None. As a result, i can't run topic_model.reduce_outliers with embeddings strategy. But somehow, reduce_topic() function still works. my understanding is _"topic embeddings are the weighted average of the top n words based on their c-TF-IDF values" and "these embeddings are then used for automatically reducing the topics" Am i understand it right that topic embedding is an aggregation of word embeddings (through the embedding model). and if so, if i did not give a embedding model, how is the topic embeddings are generated for topic reduction ? Appreciate your help ! |
Replies: 1 comment 1 reply
|
During training, there are typically two topic representations created, a c-TF-IDF vector per topic and topic embeddings. The former will always be created and the latter only when the embedding model is passed to the BERTopic model. When topics are reduced, they are reduced based on one of those representations by searching which topic representations, c-TF-IDF or embeddings, can form a cluster and therefore be merged together. In the case you mentioned above, the c-TF-IDF vectors are used instead of the topic embeddings to perform topic reduction. |
During training, there are typically two topic representations created, a c-TF-IDF vector per topic and topic embeddings. The former will always be created and the latter only when the embedding model is passed to the BERTopic model.
When topics are reduced, they are reduced based on one of those representations by searching which topic representations, c-TF-IDF or embeddings, can form a cluster and therefore be merged together.
In the case you mentioned above, the c-TF-IDF vectors are used instead of the topic embeddings to perform topic reduction.