Repository navigation
|
Hi, I'm using BERTopic for my master's thesis for the analysis of politicians' tweets during election periods. So far, I left the stopwords in the text when using BERTopic and utilized the count vectorizer model when tokenizing topics to handle stopwords appearing in topics (I followed this guide which recommended this approach: [https://python.plainenglish.io/topic-modeling-for-beginners-using-bertopic-and-python-aaf1b421afeb]). As a consequence, wouldn't that mean that Topic -1 no longer shows outliers (most often "stopwords")? Also, how can I adjust the y-axis ticks on the barchart? I've only just recently joined the Python-community, so bare with me! |
Replies: 1 comment 1 reply
No problem! Let's first start at the beginning, which is how BERTopic works. There is a page in the documentation dedicated to how it works which I highly recommend reading through as, hopefully, that will make BERTopic a bit more intuitive in its usage.
Yes and no, and the reason for that is that the topics do not directly stem from the words themselves but from clusters of documents. What you are doing with removing stopwords that way is simply making sure that a topic will not contain the words but it will have no influence whatsoever on the generation of the clusters themselves. In other words, if you have a topic with keywords ["sports", "hockey", "the", "I"], removing stopwords will simply remove the keywords "the" and "I" but it will not change the meaning of topic but merely its representation.
Typically, you do not interpret the result of Topic -1. Topic -1 consists of outliers that cannot be contained in a single topic but there is often not relationship between all documents in the outlier topic. For example, take the following image: What we can see here is that many documents in Topic -1 (the grey points) are spread throughout and have very little structure to them. In other words, you do not want to analyze Topic -1 as a whole since the whole really does not make sense in most cases. Instead, you would want to analyze sub-clusters in Topic -1 or perhaps reduce them with something like
That is currently not possible since Topic -1 is actually not a topic at all. The name is rather misleading but since it is essentially a bunch of random documents thrown together, there is little meaning to analyzing that topic from a global perspective. |



No problem! Let's first start at the beginning, which is how BERTopic works. There is a page in the documentation dedicated to how it works which I highly recommend reading through as, hopefully, that will make BERTopic a bit more intuitive in its usage.