Active Genetics & Molecular Biology Computing & AI

Genealogical and statistical methods for large-scale genomic analysis

Summary

Original abstract (not yet simplified)

Biobank datasets, containing genomic, environmental, and health data for millions of individuals, have provided insights into disease susceptibility, biological mechanisms, and human evolution, facilitating applications such as drug development and genetic risk prediction. However, these datasets present significant challenges: processing their large volumes is computationally very demanding, and modeling the heterogeneity they contain poses substantial statistical obstacles. These issues risk...

View original technical description
Biobank datasets, containing genomic, environmental, and health data for millions of individuals, have provided insights into disease susceptibility, biological mechanisms, and human evolution, facilitating applications such as drug development and genetic risk prediction. However, these datasets present significant challenges: processing their large volumes is computationally very demanding, and modeling the heterogeneity they contain poses substantial statistical obstacles. These issues risk leaving genomic datasets underutilized or amplifying existing biases, such as those linked to genetic ancestry.To address these challenges, this proposal will develop scalable statistical methods to reduce computational costs, improve the modeling of genetic ancestry, and increase accuracy and statistical power across several genomic analyses. We will focus on three specific aims. First, we will develop scalable methods to reconstruct large-scale genome-wide genealogical graphs, capturing evolutionary relation-ships and enabling applications such as simulation, phasing, imputation, and data sharing. Second, we will extend this framework to analyze both ancient and modern genomes, using genealogical graphs to define new ancestry descriptors and study human evolutionary history at fine resolution. Finally, we will create Bayesian machine learning approaches to improve the detection of trait- and disease-associated variants, model multiple traits and ancestries, learn biological function from raw genomic data, and enable distributed cross-biobank analyses. We will implement these models as high-quality, freely available open-source software.

Related Research

Grants with similar aims, by meaning.

Using hidden genealogical structure to study the architecture of human disease
Haplotype-based inference of hidden structure in the human genome
Statistical methodology for population genetics inference from massive datasets with applications in epidemiology.
Inference and analysis of gene genealogies from large genomic data sets
Leveraging genealogies for powerful inference of evolutionary processes driving human genetic diversity

Original classification

HORIZON

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.