Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation

Saibene, A.M.; Allevi, F.; Calvo-Henriquez, C.; Maniaci, A.; Mayo-Yáñez, M.; Paderno, A.; Vaira, L.A.; Felisati, G.; Craig, J.R.

doi:10.1007/s00405-023-08372-4

Purpose This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifi- cally for the multidisciplinary management of odontogenic sinusitis (ODS). Methods A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS- related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators. Results While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs’ responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators’ specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs. Conclusions LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs’ performance as they evolve over time.

Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation / A.M. Saibene, F. Allevi, C. Calvo-Henriquez, A. Maniaci, M. Mayo-Yáñez, A. Paderno, L.A. Vaira, G. Felisati, J.R. Craig. - In: EUROPEAN ARCHIVES OF OTO-RHINO-LARYNGOLOGY. - ISSN 0937-4477. - (2024), pp. 1-7. [Epub ahead of print] [10.1007/s00405-023-08372-4]

Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation

A.M. Saibene^Co-primo;F. Allevi^Co-primo;Calvo-Henriquez, Christian;Maniaci, Antonino;Mayo-Yáñez, Miguel;Paderno, Alberto;Vaira, Luigi Angelo;G. Felisati^Co-ultimo;Craig, John R.

2024

Abstract

Purpose This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifi- cally for the multidisciplinary management of odontogenic sinusitis (ODS). Methods A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS- related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators. Results While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs’ responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators’ specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs. Conclusions LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs’ performance as they evolve over time.

Scheda breve

Scheda completa

Scheda completa (DC)

	Parole chiave
	
				Artificial intelligence; Chronic rhinosinusitis; Computer-assisted diagnosis; Dental implant; Maxillary sinusitis; Oroantral fistula
			
	Settori scientifico-disciplinari dell'articolo (sola visualizzazione)
	
				Settore MED/31 - Otorinolaringoiatria
Settore MED/29 - Chirurgia Maxillofacciale
			
	Data di pubblicazione
	
				2024
			
	Rivista in ANCE
	
				EUROPEAN ARCHIVES OF OTO-RHINO-LARYNGOLOGY
			
	DOI
	
				https://dx.doi.org/10.1007/s00405-023-08372-4
			
	Tipologia
	
				Article (author)
			
	Appare nelle tipologie:
	
				01 - Articolo su periodico

File in questo prodotto:

File	Dimensione	Formato
chat GPT vs ODS (2024).pdf accesso aperto Descrizione: Online First Tipologia: Publisher's version/PDF Dimensione 565.54 kB Formato Adobe PDF Visualizza/Apri	565.54 kB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1023199

Citazioni

0

22

21

26

IRIS Institutional Research Information System - AIR Archivio Istituzionale della Ricerca