Recipient organisationKing's College LondonSource-published name: King's College London
Funding£1.7M
PeriodFeb 2025 — Feb 2030
In plain English
AI plain-English summary
Latin words shift their meanings depending on who wrote them, when, and in what context—but until now, no one has tracked those shifts systematically across two thousand years of texts. The problem is practical: to study how word meanings change over time and across genres, researchers need vast datasets where each word usage is tagged with its precise sense. For historical languages like Latin, this annotation is so labour-intensive that large-scale studies have been impossible. COALA aims to automate the process using computational methods for historical word sense disambiguation, building a corpus annotation system that can handle Latin’s full recorded history. If successful, the project will produce the first large-scale quantitative account of semantic variation and change in a historical language. This is fundamental science—it will not directly change infrastructure or daily life. But it could transform how linguists study meaning in any language with a written record, and it may eventually inform how digital tools handle ambiguous language in search engines, translation software, or historical archives. For now, the immediate payoff is a deeper, data-driven understanding of how Latin—a language long considered fossilised—actually evolved across centuries of use.
View original technical description
Understanding language crucially requires capturing words’ meanings, but these are not directly observable. Words’ meanings change over time and vary by register, genre, style, social and geographic factors. Our knowledge of semantic variation and change in historical languages is largely based on qualitative evidence from dictionaries and small-scale studies. Large quantitative studies are not possible yet because they require high-quality data with rich semantic annotation indicating the meaning of each word’s usage. Since this is time-consuming and complex, we lack large-scale quantitative accounts of semantic variation and change over long time spans. Recent computational methods in historical word sense disambiguation allow us, in principle, to automate semantic annotation. Hence, large-scale quantitative semantic analyses are now within reach. Latin has one of the longest recorded histories, an unprecedented set of tools and digital corpora covering over two thousand years and is a key part of Europe’s cultural heritage. This context places Latin in an excellent position to lead the way in quantitative historical semantics. Uniquely integrating computational methods in a novel corpus annotation system to analyse Latin words’ meaning quantitatively at scale, COALA can transform the way historical lexical semantics research is done. The impact spans multiple fields: in corpus linguistics, addressing open challenges for consistent sense annotation at scale for a historical language; in computational semantics, advancing state-ofthe-art methods as a reliable basis for lexical semantics research; in Latin and historical semantics, answering open questions on how polysemy varies by text genre, how words in the same lexical field change their meaning, and how the timing of semantic innovations relates to lasting changes. Our analysis will also be the first extensive empirical semantic investigation of Latin’s status as a fossilised language throughout its history.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know