Improving LDA topic modeling in Twitter with graph community detection

Texts can be characterized from their content using machine learning and natural language processing techniques. In particular, understanding their topic is useful for different tasks such as personalized message recommendation, fake news detection or public opinion monitoring. Latent Dirichlet Allo...

Descripción completa

Guardado en:
Detalles Bibliográficos
Autores principales: Albanese, Federico, Feuerstein, Esteban
Formato: Objeto de conferencia Resumen
Lenguaje:Inglés
Publicado: 2021
Materias:
Acceso en línea:http://sedici.unlp.edu.ar/handle/10915/140126
http://50jaiio.sadio.org.ar/pdfs/agranda/AGRANDA-04.pdf
Aporte de:
id I19-R120-10915-140126
record_format dspace
institution Universidad Nacional de La Plata
institution_str I-19
repository_str R-120
collection SEDICI (UNLP)
language Inglés
topic Ciencias Informáticas
Topic modeling
Community detection
Twitter
Text mining
Text clustering
spellingShingle Ciencias Informáticas
Topic modeling
Community detection
Twitter
Text mining
Text clustering
Albanese, Federico
Feuerstein, Esteban
Improving LDA topic modeling in Twitter with graph community detection
topic_facet Ciencias Informáticas
Topic modeling
Community detection
Twitter
Text mining
Text clustering
description Texts can be characterized from their content using machine learning and natural language processing techniques. In particular, understanding their topic is useful for different tasks such as personalized message recommendation, fake news detection or public opinion monitoring. Latent Dirichlet Allocation (LDA) is an unsupervised generative model for the decomposition of topics, which seeks to represent texts as random mixtures over topics with a Dirichlet distribution, and each topic is characterized by a distribution over words. However, this method is challenging to apply when the text is short and sometimes incoherent, as is often the case with posts on social networks such as twitter. Therefore, different works have shown that tweet pooling (aggregating tweets into longer documents) improves LDA results, but its performance depends on which method was used to aggregating the texts. We propose the new method to detect topics on twitter: “Community pooling”. In this novel scheme, first we define the retweet graph where users are the nodes and retweets between them are the edges. Then, we use the Louvain method for community detection in order to uncover the communities (a group of users who mainly interact with each other but not with other groups). Finally we aggregate into a single document all the tweets authored by all users in a community. Therefore, this method drastically reduces the number of total documents and makes denser word co-occurrence matrix, which is beneficial to LDA algorithm. With the intention of evaluating our model, we created two datasets of tweets with different characteristics. A first generic dataset involving various topics such as music, health and movies and a second dataset corresponding to an event: Biden’s presidential inauguration day in the United States. We compare the performance of our model with state of the art schemes and previous pooling models in terms of document retrieval performance, cluster quality and supervised machine learning classification score. Results showed that Community pooling had a better performance on all datasets and tasks, with the only exception of the retrieval task on the event dataset. Moreover, Community polling was faster than all other aggregation techniques (less than half the running time), which is particularly useful in big data scenarios.
format Objeto de conferencia
Resumen
author Albanese, Federico
Feuerstein, Esteban
author_facet Albanese, Federico
Feuerstein, Esteban
author_sort Albanese, Federico
title Improving LDA topic modeling in Twitter with graph community detection
title_short Improving LDA topic modeling in Twitter with graph community detection
title_full Improving LDA topic modeling in Twitter with graph community detection
title_fullStr Improving LDA topic modeling in Twitter with graph community detection
title_full_unstemmed Improving LDA topic modeling in Twitter with graph community detection
title_sort improving lda topic modeling in twitter with graph community detection
publishDate 2021
url http://sedici.unlp.edu.ar/handle/10915/140126
http://50jaiio.sadio.org.ar/pdfs/agranda/AGRANDA-04.pdf
work_keys_str_mv AT albanesefederico improvingldatopicmodelingintwitterwithgraphcommunitydetection
AT feuersteinesteban improvingldatopicmodelingintwitterwithgraphcommunitydetection
bdutipo_str Repositorios
_version_ 1764820458284253184