Active Computing & AI Public Health & Healthcare

Stewardship of Privacy-Enhancing Technologies for Scientific Research and Policymaking

In plain English

AI plain-English summary

A new framework will let researchers and regulators test whether privacy-protecting techniques actually work before releasing sensitive data about people’s health, movements, and habits. The problem is a growing tension. Governments and researchers want to use digital traces—electronic health records, mobile location data—to tackle cancer, disasters, and other large-scale challenges. But techniques meant to protect privacy, such as adding noise to data or generating synthetic datasets, can distort the information in ways that are poorly understood. Rare diseases may vanish from synthetic data. Vulnerable communities may become invisible. A group of 50 US academics recently warned that the secrecy around these anonymisation methods can produce “biases that have never been publicly quantified.” Meanwhile, public distrust has already derailed major NHS data-sharing schemes. If this Fellowship succeeds, it will give regulators, researchers, and civil society the tools to make evidence-based decisions about which privacy technique to use for a given dataset. It will produce statistical methods to extrapolate lab results to real-world settings, technical standards to quantify privacy impacts, and online tools that let researchers independently audit anonymised data. The work will also inform regulation of generative AI used to create synthetic data. The goal is to make research using digital traces both safe and reliable—without sacrificing the integrity of the science or the privacy of the people behind the data.

View original technical description
We generate vast amounts of data concerning our health, movements, and habits when interacting with technology and digital services. These digital traces are a vital key to solving society's biggest problems-for example, electronic health records can support cancer surveillance efforts, and mobile location data can support humanitarian action for disaster relief. While privacy researchers have proposed numerous techniques to safely collect, analyse, and share personal data, these systems are not without their limits. Indeed, a number of supposedly anonymous datasets have been re-identified, and a lack of public confidence derailed the NHS's care.data and GP data collection scheme that tried to share de-identified health data for research. To address the privacy threats involved in releasing sensitive human data, regulators have advocated for use of modern privacy-enhancing technologies (PETs) that have stronger privacy guarantees. However, some PET techniques-such as injecting noise into the data, or creating 'synthetic' datasets-can fundamentally distort data in unknown but potentially harmful ways, for example if rare diseases are suppressed from synthetic data, or vulnerable communities are further marginalised. A group of 50 US academics led by Prof. Gary King recently warned the US Census Bureau that the secrecy of anonymisation techniques can lead to "biases that have never been publicly quantified". This lack of understanding of how PETs will impact research and data analysis-and the policy interventions that rely on it-complicates recent calls to "unlock the power of data" for the public good. Over the course of this Fellowship, I will provide a pathway to guarantee both the privacy of data subjects *and* the utility and integrity of research data. My proposal pioneers a statistical learning and computational approach to guide the development of fair and usable PETs, allowing regulators and civil society-for the first time-to make evidence-based determinations for which privacy mechanisms to use when collecting and releasing sensitive datasets, and researchers to independently audit the validity and integrity of any anonymised data they receive. It will pioneer computationally-heavy replication studies to understand how PETs can cause harm (WP1); statistical methods to help PET developers 'extrapolate' guarantees from lab studies to the real world (WP2); technical standards and certification to quantify the impact of PETs (WP3); will produce online tools to allow researchers to audit deployed PETs (WP4); and undertake a broad programme of outreach and engagement to inform practice in policy, industry and academia (WP5). The Fellowship will thereby provide a framework to make research using digital traces safe and reliable; support data-driven policy interventions that rely on anonymised administrative data; and inform the regulation of underlying AI technologies-such as generative AI for synthetic data.

View the original record at the funder ↗

Researchers

Luc Rocher (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Towards Practical Federated Analytics and Multi-Target Privacy Enhancing Technologies (PETs)
Missing Data as Useful Data
Removing Legal Hurdles in Copyright and Data Privacy for AI-driven Research: Unleashing the Potential of AI for Science
Methods for the privacy preserving analysis of sensitive health data: text analysis and data visualisation
Patients, the public and the uses of big data; practical engagement, education, scrutiny and leadership

Original classification

Fellowship

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.