Repository navigation
Consuming too much time for Dynamic Topic Modeling on Google Colab #1056
Replies: 2 comments 8 replies
|
Generally, it should not run for that long of a time. For me, it typically runs within a couple of minutes. Could you share your full code? Also, did you make sure to set |
|
Hi Maarten, ##main Code Create your representation modelrepresentation_model = MaximalMarginalRelevance(diversity=1) Reduce outliersnew_topics = topic_model.reduce_outliers(doc, topics) |

Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hello Marteen,
I really love BERTopic but I have a question about time consuming for DTM.
I have 70,000 scientific articles (**Just ** Title+Abstract) from 2015 to 2021 (10,000 per year). Around 160 words per paper (mean) after stopwords.
I force to just 15 topics (nr_topics=16) and 20 words per topic. The time/date data is just the year on each article.
I use Google Pro+, but for a DTM run consumes around 15-20 hours.
The last one was for almost 20 hours:
7it [19:55:42, 10248.99s/it]
Is that a correct time consuming? Is there a way to reduce time consuming to save compute units =save $$ in Google Colab Pro+? My big issue is because I want to increase articles to 140,000.
Thank you very much in advance,
Roberto Carlos
All reactions