Repository navigation
|
Hi, Thanks! |
Answered by
MaartenGr
Jan 10, 2023
Replies: 1 comment 1 reply
|
Sure, you can use the from sklearn.feature_extraction.text import CountVectorizer
from bertopic import BERTopic
vectorizer_model = CountVectorizer(stop_words=a_list_of_keywords_i_want_to_exclude)
# Train a model
topic_model = BERTopic(vectorizer_model=vectorizer_model)
topics, probs = topic_model.fit_transform(docs)
# If you want to update an already trained model
topic_model.update_topics(docs, vectorizer_model=vectorizer_model) |
1 reply
Answer selected by
salderma
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Sure, you can use the
CountVectorizerto decide how the words will be tokenized before ending up in the topic representation. Here, you can decide which words you want to include and exclude in the resulting topic representation. More specifically, we can view this exclusion as stopwords that should not be put in the topic labels. In other words, we can approach it like this: