Repository navigation
Replies: 1 comment 6 replies
|
If you are using this supervised approach in BERTopic, then it might be important to choose an algorithm that works generally well with class-imbalances, like XGBoost. Also, the evaluation metric here becomes more important since metrics such as accuracy tend to favor the majority class.
Indeed, the easier the topics are to separate, the easier they are to learn. That does not necessarily mean that you will have many documents per topic but it definitely does help. |
6 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment


Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hey Maarten
I have a labelled dataset with the pictured characteristics, I am wondering how many documents I would need per topic, and also how balanced the dataset needs to be (have you observed some sort of threshold for both factors in your experience?): would this also be reliant on how similar the embeddings are for documents within each topic and how dissimilar they are to documents in different topics?
I wonder also how one would go about ensuring the minimum no. docs and the minimum requirements for balancing of classes is met systematically?
All reactions