Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application

Crema, C.; Buonocore, T.M.; Fostinelli, S.; Parimbelli, E.; Verde, F.; Fundarò, C.; Manera, M.; Ramusino, M.C.; Capelli, M.; Costa, A.; Binetti, G.; Bellazzi, R.; Redolfi, A.

doi:10.1016/j.jbi.2023.104557

The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1score 84.77 %, Precision 83.16 %, Recall 86.44 %. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "low-resource" approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.

Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application / C. Crema, T.M. Buonocore, S. Fostinelli, E. Parimbelli, F. Verde, C. Fundarò, M. Manera, M.C. Ramusino, M. Capelli, A. Costa, G. Binetti, R. Bellazzi, A. Redolfi. - In: JOURNAL OF BIOMEDICAL INFORMATICS. - ISSN 1532-0464. - 148:(2023), pp. 104557.1-104557.10. [10.1016/j.jbi.2023.104557]

Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application

Crema, Claudio;Buonocore, Tommaso Mario;Fostinelli, Silvia;Parimbelli, Enea;F. Verde;Fundarò, Cira;Manera, Marina;Ramusino, Matteo Cotta;Capelli, Marco;Costa, Alfredo;Binetti, Giuliano;Bellazzi, Riccardo;Redolfi, Alberto

2023

Abstract

The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1score 84.77 %, Precision 83.16 %, Recall 86.44 %. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "low-resource" approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.

Scheda breve

Scheda completa

Scheda completa (DC)

	Parole chiave
	
				Biomedical text mining; Deep learning; Language model; Natural language processing; Transformer
			
	Settori scientifico-disciplinari dell'articolo (sola visualizzazione)
	
				Settore MED/26 - Neurologia
			
	Data di pubblicazione
	
				2023
			
	Rivista in ANCE
	
				JOURNAL OF BIOMEDICAL INFORMATICS
			
	DOI
	
				https://dx.doi.org/10.1016/j.jbi.2023.104557
			
	Tipologia
	
				Article (author)
			
	Appare nelle tipologie:
	
				01 - Articolo su periodico

File in questo prodotto:

File	Dimensione	Formato
1-s2.0-S1532046423002782-main.pdf accesso aperto Descrizione: Original Research Tipologia: Publisher's version/PDF Dimensione 1.22 MB Formato Adobe PDF Visualizza/Apri	1.22 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1032034

Citazioni

0

7

5

ND

IRIS Institutional Research Information System - AIR Archivio Istituzionale della Ricerca