Skip to content
Discussion options

You must be logged in to vote

Generally, if they are Western languages, then it would be okay to just use a multi-lingual embedding model, as is implemented in BERTopic. However, since you also have a number of languages that follow different tokenization schemes, it is important that you make sure tokenization is also multi-lingual. This can be a tricky process, so translating the documents to English might be the easiest way forward, especially if the translation is done very well.

Replies: 1 comment 1 reply

Comment options

You must be logged in to vote
1 reply
@rcote
Comment options

Answer selected by rcote
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants