Completed Cells, Biochemistry & Physiology Computing & AI

Building a Next Generation Image Repository: Molecular Annotation and Cloud-based Data Processing and Analysis

In plain English

AI plain-English summary

Biologists rarely share their original microscope images, even though those images contain data as rich as gene sequences or protein structures. The problem is not reluctance—it is that imaging datasets are complex, heterogeneous and often enormous, so no standard repository exists to host them. This project will build that repository at EMBL-EBI, the European hub for molecular data, using open-source technologies already proven to handle terabytes of image data. It will link image data with existing genomic, structural and phenotypic databases, so a scientist can browse and compute across all scales—from molecules to whole organisms—without switching platforms. If it succeeds, the infrastructure will quietly change how biological research is validated and accelerated. Just as shared gene sequence databases transformed molecular biology, a shared image repository could make microscopy data as routinely reusable as DNA sequences, enabling new discoveries from old experiments and reducing the need to repeat costly imaging work.

View original technical description
Access to primary research data is vital for the advancement of the scientific enterprise. It facilitates the validation of existing observations and provides the raw materials to build on those observations. In the life sciences, there are numerous examples where members of a research community determined that a particular type of data would be useful and necessary to share. These include gene sequences, protein structural data, and gene and protein expression profiles. In these cases the community united to standardize the structure of the data and its associated metadata, and to create centralized repositories to facilitate deposition, promote discoverability, and ensure the longevity of the data. Imaging in the life sciences has undergone a revolution in recent years and is now used as a quantitative assay technology throughout the life and biomedical sciences. Imaging is used to understand the behavior of organisms, the formation of embryos, the structure and dynamics of cells, and the function and interactions of molecules that are the building blocks of life. Imaging datasets are complex, heterogeneous, and often extremely large, so they are rarely shared or published. Based on the recent development of several image data management technologies and the rapidly decreasing cost of large data storage facilities, we propose to create a resource to host, serve, and make available original scientific image data that underpins life sciences research. Our proposal is based on open source technologies with proven utility and performance that already run on-line resources serving several terabytes (TBs) of image data. We propose to place this resource at EMBL-EBI, which is the established home of molecular and structural life sciences data and interface the resource with ELIXIR, Europe's research infrastructure for life science informatics. In particular we will build links with established molecular and structural resources and work towards a seamless integration of these data, so that any scientist can easily browse, query and compute on genomic, structural and phenotypic data across several scales.

View the original record at the funder ↗

Researchers

Alvis Brazma (Co-Investigator)Jason Swedlow (Principal Investigator)Rafael Edgardo Carazo Salas (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Public archiving and data integration in the era of multi-modal imaging
Intuitive Large-scale Image Processing for Biologists
World-wide E-infrastructure for structural biology
Integrating 3D biological data on scales from molecules to cells
BioStudies and the Image Data Resource: Expanding Imaging Datasets, Linkage, Metadata, and Value

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.