Repository navigation
|
Hi Maarten and others. I have a set of documents in multiple different languages (English, Arabic, Chinese, Japenese, French) on which I need to do Topic Modeling. What is the best strategy to deal with that? Should I translate all documents to English first? Should I create several sets of embeddings, one per language, using different language models, and then merge them together to complete the topic analysis? Any insights are welcome! Rémi |
Replies: 1 comment 1 reply
|
Generally, if they are Western languages, then it would be okay to just use a multi-lingual embedding model, as is implemented in BERTopic. However, since you also have a number of languages that follow different tokenization schemes, it is important that you make sure tokenization is also multi-lingual. This can be a tricky process, so translating the documents to English might be the easiest way forward, especially if the translation is done very well. |
Generally, if they are Western languages, then it would be okay to just use a multi-lingual embedding model, as is implemented in BERTopic. However, since you also have a number of languages that follow different tokenization schemes, it is important that you make sure tokenization is also multi-lingual. This can be a tricky process, so translating the documents to English might be the easiest way forward, especially if the translation is done very well.