A self-supervised seed-driven approach to topic modelling and clustering
Informazioni aggiuntive
Autori
Ravenda F.,
Bahrainian S. A.,
Raballo A.,
Mira A.,
Crestani F.
Tipo
Articolo pubblicato in rivista scientifica
Anno
2024
Lingua
Inglese
Sommario
Topic models are useful tools for extracting the most salient themes within a collection of documents, grouping them to construct clusters representative of each specific topic. These clusters summarize and represent the semantic contents of the documents for better document interpretation. In this work, we present a light approach able to learn topic representations in a Self-Supervised fashion. More specifically, we propose a lightweight and scalable architecture using a seed-word driven approach to simultaneously co-learn a representation from a document and its corresponding word embeddings. The results obtained on a variety of datasets of different sizes and natures show that our model is capable of extracting meaningful topics. Furthermore, our experiments on five benchmark datasets illustrate that our model outperforms both traditional and neural topic modelling baseline models in terms of different coherence and clustering accuracy measures.
Periodico
Journal of Intelligent Information Systems
Volume
63
Numero ( Mese )
1
Pagine (o numero dell’articolo)
333-353
ISSN
0925-9902, 1573-7675
Diffusione
Licenza
Licenza non definita
Visibilità
Privato