The 100,000 Genomes Project is building a shared, secure computing platform to let academic researchers analyse whole genome sequences and clinical data from NHS patients with cancer and rare inherited disorders. This matters because the project’s industrial arm already lets companies work with anonymised data in a secure environment, but academic researchers have lacked equivalent access to the richer, read-level genomic data needed for robust discovery. Without this infrastructure, valuable clinical data collected through the NHS sits underused for fundamental science. If the infrastructure succeeds, it will create a standardised, reusable environment where researchers can produce “research-ready” datasets, manage patient and genomic data securely, and collaborate on large-scale clinical studies. The platform reuses designs and software already tested by UK Biobank and the European Bioinformatics Institute, so it builds on proven foundations rather than starting from scratch. Public, charitable, and philanthropic funders will have a formal mechanism to engage, and projects that add value to the Genomics England programme can access the compute infrastructure at no charge. The result is a quieter but critical piece of national research infrastructure—one that keeps the data flowing securely between clinics, labs, and companies.
View original technical description
This proposal to the MRC will establish a shared, secure, high performance data and compute infrastructure as a platform for large-scale clinical genomics research based on the data flows of the UK 100,000 Genomes project. Samples and data from patients with cancer and rare, inherited disorders will be provided by NHS England, working in collaboration with Cancer Research UK and programmes funded by the NIHR and the MRC. Genomics England, a company wholly owned by the Department of Health, will pay for the generation of whole genome sequence data. Genomics England will pay also for the generation of summary reports, based upon clinical annotations of this data, and will return these to the NHS to support patient care. Genomics England will make anonymised, redacted versions of the data available for industrial research strictly within a secure, managed environment. The proposed infrastructure will provide a similar environment for academic research, with a more comprehensive collection of genomic and patient data, including the read-level data used for the generation of variant calls and summary reports. The infrastructure will include software tools to support the production of 'research-ready' data sets, the effective management of patient and genomic data, and the delivery of collaborative clinical research. The project partners have experience in infrastructure development and clinical genomics research, and will be able to reuse designs, procedures, and software developed and tested within existing programmes and organisations, including UK Biobank and the European Bioinformatics Institute. A formal mechanism will be established for engagement with public, charitable, and philanthropic funders, and with the clinical research projects that they fund. Subject to capacity constraints, projects that add appropriate value to the Genomics England programme will be provided with access to the compute infrastructure at no charge.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know