Recipient organisationCardiff UniversitySource-published name: Cardiff University
Funding£1.4M
PeriodMar 2016 — Nov 2020
In plain English
AI plain-English summary
Welsh speakers will soon be able to donate their everyday conversations, text messages, and social media posts to build the first comprehensive digital record of their living language. The problem is straightforward: Welsh has never had a large, searchable database of real language use—what linguists call a corpus. Without one, dictionary makers, teachers, translators, and tech developers must rely on intuition or outdated rules about how Welsh "should" be spoken, rather than evidence of how it actually is. This gap has held back everything from Welsh-language predictive text and voice recognition software to teaching materials that reflect how people genuinely speak. If CorCenCC succeeds, it will change how Welsh is taught, used in government and business, and integrated into everyday technology. The corpus will be open-source and publicly accessible, allowing a learner to find real examples of a dialect, a publisher to check readability, or a developer to train machine translation tools. Because the project involves crowdsourcing—speakers recording and uploading their own language via a mobile app—it captures regional and social variation that traditional methods miss. The design is also shaped from the start by the people who will use it, including the Welsh Government, the BBC, and the Welsh Joint Education Committee, ensuring the resource meets actual needs rather than academic assumptions.
View original technical description
This project will create a major corpus of Welsh language: CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes: National Corpus of Contemporary Welsh). A corpus is a principled collection of language data sampled from real-life contexts, presented as a searchable database. This will be the first corpus to represent spoken, written and electronically-mediated Welsh, and the first in any language with a functional design informed, from the outset, by representatives of all anticipated academic and community user groups. CorCenCC will provide societal, economic and academic benefits by: - Facilitating uses of Welsh in public, commercial, educational and governmental settings. - Redefining the scope, relevance and design infrastructure of corpus development methodology. A corpus allows users to identify and explore language as it is actually used, rather than relying on intuition or prescriptive accounts of how it 'should' be used. This evidence-based approach is used by academic researchers, lexicographers, teachers, language learners, assessors, resource developers, policy makers, publishers, translators and others, and is essential to the development of technologies such as predictive text production, word processing tools, machine translation, voice recognition and web search tools. Welsh has had no comprehensive corpus facility able to meet these requirements. CorCenCC will capitalise on extensive community interest in sustaining and 'growing' Welsh, using the novel integration of crowdsourcing, a powerful data collection method which has the potential to revolutionize corpus construction. Recruited through social and broadcast media, roadshows and existing networks, Welsh speakers will record and upload their own data via a mobile app, and even contribute to data coding. This approach promises representative language across genres, language varieties (regional and social) and contexts. Traditional, data collection will supplement the crowdsourcing, ensuring a representative balance of data as specified in the project targets. Preliminary engagement with stakeholders (including a briefing event at the Senedd) generated collaboration from the Welsh Government, Welsh Language Commissioner, Welsh Joint Education Committee, Welsh for Adults, BBC, Gwasg y Lolfa press, and University of Wales Dictionary; all have identified current needs which CorCenCC can meet, and all will be represented in the project advisory group, so the corpus design is user-informed throughout. A language corpus able to inform delivery of Welsh has been called for by e.g. National Foundation for Educational Research (2008:48) and Welsh Government (2013:27,71). CorCenCC, with its integrated pedagogical toolkit, will impact significantly on Welsh language teaching practice, enabling data-driven, inductive learning and assessment. CorCenCC will be open-source and publicly accessible, with user interfaces for specific groups. It will enable, for example, community users to investigate dialect variation or idiosyncrasies of their own language use; professional users to profile texts for readability or develop digital language tools; language learners learn from real life models of Welsh; and researchers to investigate patterns of language use and change. In order to ensure that CorCenCC remains a sustainable, permanent and user-oriented record of language, an in-built facility will allow data to be added and moderated beyond the life of the project. The project team comprises experts in corpus linguistics, Welsh, and language pedagogy and assessment, who specialise in the application of linguistic tools to real world issues. Working with an advisory body of stakeholder representatives, they are optimally placed to meet the project aims: creating a permanent, sustainable and fit-for-purpose record of the living language, and pioneering an approach to content generation and user-driven applications that will provide a model for future corpus creation.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know