Completed Genetics & Molecular Biology Cells, Biochemistry & Physiology

Trace Archive and 1KG DCC.

In plain English

AI plain-English summary

The raw DNA sequence data from thousands of genomes—the foundational layer beneath every genetic discovery—is at risk of being lost unless funding is secured to keep its archive running. This proposal covers a critical funding gap for the Trace Archive, a repository that stores unprocessed sequencing data before it is analysed. Unlike polished genome assemblies, these raw traces can be reanalysed years later with new tools, combining datasets from different labs and time periods. The rise of next-generation sequencing has vastly increased data output, especially for clinical applications, but the archive itself lacks sustained support. Without it, the 1,000 Genomes Project—a major international effort to catalogue human genetic variation—would lose its raw data backbone. If this succeeds, the archive will remain accessible for future reanalysis, enabling researchers to spot variants missed by older algorithms and to integrate data across studies. This is infrastructure work: invisible to the public, but essential for every genetic test, drug target, and disease association study that relies on comparing patient genomes to a reference. The project is fundamentally about preserving a shared scientific resource, not generating new discoveries itself.

View original technical description
The trace archive stores raw sequencing information before analysis, and thus represents the foundational dataset for all DNA sequence resources. This data can then be reanalysed in the future, in particular combining datasets over time and geographically separated institutes. The presence of next generation sequencing machines has provided a step increase in the capacity of sequencing centres and has broadened the use of sequencing to many new applications, in particular clinical cases. This pr oposal aims to bridge a funding gap for the trace archive until the European infrastructure proposals are completed, and focuses on maximising the utility to one major next generation sequencing project, the 1,000 genomes project.

View the original record at the funder ↗

Researchers

Ewan Birney (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

MicrobesNG: A scalable replicable biological sample repository incorporating whole-genome sequence data and analysis of thousands of microbial strains
Improved quality of high-throughput sequencing by computational methods.
16 ERA-CAPS: 1001 Genomes Plus
High throughput Sequencing Hub for the North of England
Development of Novel Computational Strategies to Store and Interpret Next Generation Sequencing Data and Their Application to Multi-Genomic Analyses

Original classification

Strategic Award - Science

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.