Topic detection is a fundamental task in text mining, aimed at uncovering latent thematic structures within large textual corpora. In this study, we propose the use of co-occurrence networks to identify topics through community detection algorithms. To compare and evaluate which algorithm best fits a given network, we introduce a new methodology named robin, which implements a procedure for comparing communities through robustness analysis. We applied our approach to a corpus of 71 articles by Italian newspapers between 1914 and 1919, with the aim of exploring the evolution of economic and financial discourse in Italy during World War I and the early postwar period. By constructing word co-occurrence networks across multiple years, we traced the evolution of thematic areas over time. Our results demonstrate that the choice of community detection algorithm significantly influences the identified topics and their temporal dynamics. By combining network analysis and algorithm evaluation, we offer a replicable and data-driven procedure for improving topic detection in complex text corpora.
Robust community detection for topic identification in narratives of financial crisis / V. Policastro, G.D.L. - In: JADT 2026 Proceedings. 1 / [a cura di] A. Plaia, M. Sciandra, A. Albano. - Prima edizione. - Palermo : University of Palermo, 2026. - ISBN 978-88-5509-882-3. - pp. 21-27 (( 18. International Conference on Statistical Analysis of Textual Data Palermo 2026.
Robust community detection for topic identification in narratives of financial crisis
G. De Luca;
2026
Abstract
Topic detection is a fundamental task in text mining, aimed at uncovering latent thematic structures within large textual corpora. In this study, we propose the use of co-occurrence networks to identify topics through community detection algorithms. To compare and evaluate which algorithm best fits a given network, we introduce a new methodology named robin, which implements a procedure for comparing communities through robustness analysis. We applied our approach to a corpus of 71 articles by Italian newspapers between 1914 and 1919, with the aim of exploring the evolution of economic and financial discourse in Italy during World War I and the early postwar period. By constructing word co-occurrence networks across multiple years, we traced the evolution of thematic areas over time. Our results demonstrate that the choice of community detection algorithm significantly influences the identified topics and their temporal dynamics. By combining network analysis and algorithm evaluation, we offer a replicable and data-driven procedure for improving topic detection in complex text corpora.| File | Dimensione | Formato | |
|---|---|---|---|
|
Proceedings_JADT_2026 Volume I (1).pdf
accesso riservato
Descrizione: Pdf della pubblicazione
Tipologia:
Publisher's version/PDF
Licenza:
Nessuna licenza
Dimensione
3.65 MB
Formato
Adobe PDF
|
3.65 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




