Training of Latin language models is rarely done with consideration of important historical watersheds. Here we demonstrate how this leads to a poor performance when specific socio-temporal contextualisation is sought, something common to humanities research. We perform an eval- uation that compares the historical adequacy of Latin language models, i.e. their ability to generate tokens, representative for a historical period. We adopt a previously established method and refine it to overcome limitations due to Latin being an under-resourced language and one with intense tradition of intertextuality. To do this we extract word lists and concordances from the LatinISE corpus and use them to compare seven masked language models trained for Latin. We further perform statistical analysis of the results in order to identify the best and worst performing models in each of the historical contexts of interest. We show that BERT medieval multilingual best captures the Classical linguistic context. Four models are indistinguishably good in our evaluation of the the Neo-Latin linguistic con- text. These findings have broad implications for wider historical language research and beyond. Among these, we emphasise the need to train historical language models with due attention on consistent historical periods and we discuss the possible usefulness of noisy predictions. Historical research of language models provides a neat demonstration of how model biases could impact their performance in specific domains.
The Latin Language Evolved Over Time, Masked Models Disregard That / M. Cuscito, A.F. - In: Computational Humanities Research 2025 / [a cura di] T. Arnold, M. Fantoli, R. Ros. - [s.l] : Association for Computers and the Humanities, 2025. - pp. 1336-1347 [10.63744/slahynqda8fu]
The Latin Language Evolved Over Time, Masked Models Disregard That
A. Ferrara;M. Ruskov
2025
Abstract
Training of Latin language models is rarely done with consideration of important historical watersheds. Here we demonstrate how this leads to a poor performance when specific socio-temporal contextualisation is sought, something common to humanities research. We perform an eval- uation that compares the historical adequacy of Latin language models, i.e. their ability to generate tokens, representative for a historical period. We adopt a previously established method and refine it to overcome limitations due to Latin being an under-resourced language and one with intense tradition of intertextuality. To do this we extract word lists and concordances from the LatinISE corpus and use them to compare seven masked language models trained for Latin. We further perform statistical analysis of the results in order to identify the best and worst performing models in each of the historical contexts of interest. We show that BERT medieval multilingual best captures the Classical linguistic context. Four models are indistinguishably good in our evaluation of the the Neo-Latin linguistic con- text. These findings have broad implications for wider historical language research and beyond. Among these, we emphasise the need to train historical language models with due attention on consistent historical periods and we discuss the possible usefulness of noisy predictions. Historical research of language models provides a neat demonstration of how model biases could impact their performance in specific domains.| File | Dimensione | Formato | |
|---|---|---|---|
|
unpaywall-bitstream--430395345.pdf
accesso aperto
Tipologia:
Publisher's version/PDF
Licenza:
Creative commons
Dimensione
430.18 kB
Formato
Adobe PDF
|
430.18 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




