Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Lorenzi, A.; Pugliese, G.; Maniaci, A.; Lechien, J.R.; Allevi, F.; Boscolo-Rizzo, P.; Luigi Angelo Vaira,; Saibene, A.M.

doi:10.1007/s00405-024-08746-2

Purpose: This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI’s ChatGPT 4 and Google’s Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes. Methods: This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test. Results: Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2–4]) compared to Gemini Advanced (median 2 [2–3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection. Conclusions: While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced / A. Lorenzi, G. Pugliese, A. Maniaci, J.R. Lechien, F. Allevi, P. Boscolo-Rizzo, L. Angelo Vaira, A.M. Saibene. - In: EUROPEAN ARCHIVES OF OTO-RHINO-LARYNGOLOGY. - ISSN 0937-4477. - (2024), pp. 1-6. [Epub ahead of print] [10.1007/s00405-024-08746-2]

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Andrea Lorenzi^Co-primo;G. Pugliese^Co-primo;Antonino Maniaci;Jerome R. Lechien;F. Allevi;Paolo Boscolo-Rizzo;Luigi Angelo Vaira;A.M. Saibene^Ultimo

2024

Abstract

Purpose: This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI’s ChatGPT 4 and Google’s Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes. Methods: This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test. Results: Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2–4]) compared to Gemini Advanced (median 2 [2–3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection. Conclusions: While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.

Scheda breve

Scheda completa

Scheda completa (DC)

	Parole chiave
	
				Artificial intelligence; Computer-assisted diagnosis; Head and neck cancer; Head and neck oncology; Large language models; Laryngeal carcinoma; Nasopharyngeal carcinoma; Oncological diagnosis; Oropharyngeal carcinoma; Parotid carcinoma; Tongue carcinoma;
			
	Settori scientifico-disciplinari dell'articolo (sola visualizzazione)
	
				Settore MED/31 - Otorinolaringoiatria
Settore MED/29 - Chirurgia Maxillofacciale
			
	Data di pubblicazione
	
				2024
			
	Data ahead of print o data di stampa
	
				25-mag-2024
			
	Rivista in ANCE
	
				EUROPEAN ARCHIVES OF OTO-RHINO-LARYNGOLOGY
			
	DOI
	
				https://dx.doi.org/10.1007/s00405-024-08746-2
			
	Tipologia
	
				Article (author)
			
	Appare nelle tipologie:
	
				01 - Articolo su periodico

File in questo prodotto:

File	Dimensione	Formato
chat gpt vs gemini vs onco h&n (2024).pdf accesso aperto Tipologia: Publisher's version/PDF Dimensione 605.22 kB Formato Adobe PDF Visualizza/Apri	605.22 kB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1053609

Citazioni

0

25

22

28

IRIS Institutional Research Information System - AIR Archivio Istituzionale della Ricerca