Repository navigation
Replies: 1 comment 2 replies
|
Thank you for sharing this, sounds like an interesting use case!
A custom vectorizer is indeed what is typically used for these kinds of solutions. A nice example of this would be the POS tagger KeyphraseVectorizers which has implemented a tagger component of SpaCy as a vectorizer model. You could do the same with NER.
The one problem with such a heavy-compute vectorizer is that it can slow things down quite a bit when used in the CountVectorizer as it processes all individual documents. Instead, you could focus on NER as a |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I've used the BERTopic approach in several research papers now embedding mainly research abstracts, but I decided to extract keywords using NER before creating topic representations with cTF-IDF.
(I published the fine-tuned NER model for extracting scientific methods/issues and other terminology https://huggingface.co/RJuro/SciNERTopic along with a Colab notebook that reproduces my pipeline).
So far, I'm reconstructing the BERTopic pipeline in ad-hoc notebooks. Such domain-specific keywords have lifted topic representation to a new level, IMO. In other areas they may enable new applications, e.g. clustering job-postings and representing by extracted skill-requirement mentions.
Is there a smarter way to integrate NER-based keyword-extraction into the pipeline?
Would that require coding up a custom vectorizer?
Thanks
Roman
All reactions