Scotland’s historical birth, marriage, and death records—24 million handwritten image files—are being turned into machine-readable text so researchers can search and analyse them at scale. Currently, these records exist only as indexed images. A researcher wanting to study, say, causes of death across generations must look up each person by name and copy the information by hand. That makes large-scale population research effectively impossible. The project will transcribe every record into structured data, covering the vast majority of people who have lived in Scotland since 1855. This will allow researchers to link individuals across generations—parents, grandparents, and descendants—and connect that information to existing health and longitudinal studies. If successful, the UK will gain a population data system comparable to those in Scandinavia and the Netherlands, which have already transformed understanding of demography, economic history, and social mobility. The immediate impact is on fundamental science: researchers will be able to ask questions about how family background, occupation, and longevity interact across multiple generations. In the longer term, this richer historical context could improve health informatics and genetic studies by revealing how past populations differ from the people who survived to have children—a bias that current datasets cannot correct.
View original technical description
This project aims to digitise the 24 million vital events record images (births, marriages and deaths) for all people in Scotland since 1855 (ie transcribe them into machine encoded text). This will allow research access to individual level information on some 18 million individuals, a large proportion of those who have ever lived in Scotland between 1855 to the present day. At the moment these records are kept as indexed images. This means that to extract any data, a researcher must search for an individual record by name and then manually transcribe the information they need themselves (eg cause of death, occupation etc.); this of course makes any large scale research impossible. A one off investment would mean that these records could made available for major population research. This dataset will be prepared for linkage to existing longitudinal studies, primarily the Scottish Longitudinal Study (SLS), and more generally to the already highly developed Scottish health informatics systems. This will allow the characteristics (place of birth, age at marriage, occupation, longevity, cause of death etc) of parents, grandparents and other relatives of those followed by the SLS to be analysed and will therefore enhance contemporary Scottish and UK health datasets, health informatics systems, longitudinal datasets and genetic studies. People surviving to have children, grandchildren and other descendents are unlikely to be fully representative of the historic population of Scotland, and the complete transcription will importantly enable a complete linkage exercise between the different to create full or partial life histories for all those experiencing vital events in Scotland since 1855.This will mean that for the first time the UK will have a data system of a similar potential depth and breadth as the Scandinavian and Low countries, whose population registers provide such life histories, with countries, where work such as those using "The Demographic Database" at Umeá University, Sweden; the "Historical Population Registers" Project at the Norwegian Historical Data Centre, University of Tromso, Norway and the Historical Sample of the Netherlands, based at the International Institute of Social History, Amsterdam, Netherlands is currently extending knowledge of demography as well as economic and social history over the nineteenth and early decades of the twentieth century. Such a dataset will bring the possibility of exploring the condition of the present Scottish population within the context of their families through multiple generations of micro-data.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know