Active Genetics & Molecular Biology Cancer

AI-driven Biomedical Discovery from Spatial and Single-cell Cancer Data

In plain English

AI plain-English summary

A single tumour can contain millions of cells, each with different genes switched on or off, and researchers are drowning in the data that describes them. The problem is that existing scientific literature holds vast knowledge about how genes work, but no one can read all of it and connect it to the complex patterns seen in real tumours. This project builds a computer system that does both at once: it analyses spatial and single-cell data to see exactly where different cells sit inside a tumour and which genes they are expressing, while simultaneously using natural language processing to automatically read millions of research papers and extract causal relationships between genes and disease mechanisms. If it works, the system will generate testable hypotheses about which gene changes actually drive cancer progression, rather than just correlating with it. This could lead to more accurate diagnostic signatures that transfer reliably between different cancer types, and ultimately help identify new targets for treatment. The research is primarily a data science and fundamental biology project—it will not directly change patient care tomorrow, but it aims to solve a bottleneck that currently prevents existing knowledge from being translated into clinical tools.

View original technical description
Genes and their functions can change dramatically in cancerous tissues, often leading to increased disease progression and treatment resistance. Understanding these changes is complex, as gene activity varies between cells and tumor environments. While there is a breadth of scientific literature describing gene functions, relating them to disease mechanisms, and distinguishing causal from correlative relationships is time-consuming and overwhelming for researchers. To address this, we will integrate omics data from spatial assays and single-cell data. This allows visualization of cell locations within tissue and the genes they express. Using data science tools, we will map, compare, and predict roles of genes and cells in tissue contexts across different cancers, e.g., identifying gene signatures indicative of disease-associated tissue properties. We will use Natural Language Processing (NLP) to automatically read and extract knowledge from scientific literature to better understand and inform our omics data analysis, generating hypotheses to explain the differences we see across tissues and cancers. We hypothesise that using the literature will establish causal hypotheses grounded in mechanistic knowledge, leading to more robust and transferable disease signatures. Our work is expected to facilitate interpreting complex gene activity, providing clearer insights, more accurate omics-informed diagnostics, and new avenues for targeted treatment.

View the original record at the funder ↗

Researchers

Maureen Ng'etich (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Decoding the lung cancer microenvironment through spatial approaches
Exploring disentangled generative factors of cancer transcriptomes
Topological Analysis for Predictive Spatial Signatures of Cancer progression and Treatment Response
Single cell profiling of breast tumors and the residing immune cells leveraged by integration with multi-dimensional molecular data from thousands of tumors
Investigating the Prognostic and Predictive Values of Spatial Cellular Interactions in Colorectal Cancer through Spatial Data Analysis and Machine Lea

Original classification

PhD Studentship (Basic)

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.