Topic detection is a fundamental task in text mining, aimed at uncovering latent thematic structures within large textual corpora. In this study, we propose the use of co-occurrence networks to identify topics through community detection algorithms. To compare and evaluate which algorithm best fits a given network, we introduce a new methodology named robin, which implements a procedure for comparing communities through robustness analysis. We applied our approach to a corpus of 71 articles by Italian newspapers between 1914 and 1919, with the aim of exploring the evolution of economic and financial discourse in Italy during World War I and the early postwar period. By constructing word co-occurrence networks across multiple years, we traced the evolution of thematic areas over time. Our results demonstrate that the choice of community detection algorithm significantly influences the identified topics and their temporal dynamics. By combining network analysis and algorithm evaluation, we offer a replicable and data-driven procedure for improving topic detection in complex text corpora.

Robust community detection for topic identification in narratives of financial crisis / V. Policastro, G.D.L. - In: JADT 2026 Proceedings. 1 / [a cura di] A. Plaia, M. Sciandra, A. Albano. - Prima edizione. - Palermo : University of Palermo, 2026. - ISBN 978-88-5509-882-3. - pp. 21-27 (( 18. International Conference on Statistical Analysis of Textual Data Palermo 2026.

Robust community detection for topic identification in narratives of financial crisis

G. De Luca;
2026

Abstract

Topic detection is a fundamental task in text mining, aimed at uncovering latent thematic structures within large textual corpora. In this study, we propose the use of co-occurrence networks to identify topics through community detection algorithms. To compare and evaluate which algorithm best fits a given network, we introduce a new methodology named robin, which implements a procedure for comparing communities through robustness analysis. We applied our approach to a corpus of 71 articles by Italian newspapers between 1914 and 1919, with the aim of exploring the evolution of economic and financial discourse in Italy during World War I and the early postwar period. By constructing word co-occurrence networks across multiple years, we traced the evolution of thematic areas over time. Our results demonstrate that the choice of community detection algorithm significantly influences the identified topics and their temporal dynamics. By combining network analysis and algorithm evaluation, we offer a replicable and data-driven procedure for improving topic detection in complex text corpora.
text mining; topic detection; network; co-occurence; robustness
Settore STEC-01/B - Storia economica
2026
Book Part (author)
File in questo prodotto:
File Dimensione Formato  
Proceedings_JADT_2026 Volume I (1).pdf

accesso riservato

Descrizione: Pdf della pubblicazione
Tipologia: Publisher's version/PDF
Licenza: Nessuna licenza
Dimensione 3.65 MB
Formato Adobe PDF
3.65 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1260876
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact