Repository navigation
Has anyone noticed that sometimes you get another outlier cluster? #908
|
Hey so this was something I just observed in my data that when I set nr_clusters='auto', cluster 0 appears to almost be like a second "outlier" topic. Basically the data I have is a bunch of short text so I am thinking that is part of the problem since I bet the ctfidf calculation is only locking onto a few key terms. But I am more just curious if anyone else has seen this sort of behavior with other data? I probably need to spend more time tuning the params / messing with my data more. |
Replies: 1 comment 1 reply
|
So the term "outlier" is doing a lot of work here. Since you are relying on HDBSCAN to determine the number of topics (clusters) by using 'auto' it is using the BERTopic defaults to determine HDBSCAN's min_cluster_size which will effect the number of clusters formed for your embeddings. (See TopicTuner to easily see how different values will change the HDBSCAN clusters). So for HDBSCAN there is only going to be ONE outlier classification - these are vectors that it can't fit into a cluster given the parameters for that instance of HDBSCAN (altering these can dramatically change this behavior). What you are describing is a cluster that HDBSCAN identified as a 'real' cluster - it just doesn't make sense to you. I've seen this behavior with badly formatted documents - lots of garbage characters and/or "junk" like email auto responds. I've even seen this with text in languages other than English. The ctidf vocabularies are created downstream of the HDBSCAN clustering - so it doesn't effect the formation of clusters. Your intuition about the size of the texts may have something to do with the behavior - but only because it is playing a role in how HDBSCAN is interpreting the embeddings when it forms the clusters. If I were in your shoes I would run TopicTuner and explore how different HDBSCAN parameters effects HDBSCAN cluster formation. It may be that topic 0 in your case actually represents a bunch of texts that are logically related to one another. I've played around with using HDBSCAN to identify junk data that I can then remove. The other possibility is that there are really multiple logical clusters which are getting aggregated into a single cluster. |
So the term "outlier" is doing a lot of work here. Since you are relying on HDBSCAN to determine the number of topics (clusters) by using 'auto' it is using the BERTopic defaults to determine HDBSCAN's min_cluster_size which will effect the number of clusters formed for your embeddings. (See TopicTuner to easily see how different values will change the HDBSCAN clusters). So for HDBSCAN there is only going to be ONE outlier classification - these are vectors that it can't fit into a cluster given the parameters for that instance of HDBSCAN (altering these can dramatically change this behavior).
What you are describing is a cluster that HDBSCAN identified as a 'real' cluster - it just doesn…