A computer model is learning to spot which animal viruses could jump to humans before they cause the next pandemic. The problem is that we cannot reliably predict which of the thousands of known animal viruses will spill over. The COVID-19 pandemic showed that SARS-like coronaviruses could have been flagged as high risk in hindsight, but we lack tools to do this prospectively. This research tackles that gap by training machine learning models on large datasets—genome sequences, published research, and data on which tissues viruses infect. If successful, the models would produce a ranked list of priority viruses for surveillance, allowing public health agencies to monitor the most dangerous candidates before they emerge. The work also identifies specific viral protein "hotspots" and host receptor proteins that could become targets for pre-emptive drugs or vaccines. This is fundamental computational biology with a direct public health payoff: better preparedness, fewer surprises.
View original technical description
Despite substantial research, we have so far failed to successfully predict which viruses would emerge to cause outbreaks with large burdens to public health and economies. Research addressing SARS-CoV-2 (the virus causing the COVID-19 pandemic) has shown that, in retrospect, SARS-like coronaviruses could have been predicted as high risk. To prepare for future pandemics, we need more reliable and specific predictions of which viruses have potential to be 'zoonotic', i.e., capable of transmitting from animals to humans. This research will investigate new ways of making predictions by taking advantage of large contemporary datasets, e.g., genome sequence repositories and text-mined published research. Machine learning will be used as a state-of-the-art computational toolkit that can build models to find patterns in complex information (e.g., images, text, genetic sequences) and apply them to specific tasks (e.g., predicting whether a virus is zoonotic or not). I will model mammal and bird RNA viruses as the most likely sources of emerging infections. By incorporating traditionally neglected data that better captures how viruses interact with host proteins and tissues, models will predict potential of viruses to zoonotically infect and cause disease in humans with improved quality and precision. Although viral sequencing has improved in coverage, different viruses have been sampled unequally. Resulting biases can lead to poor performance or misidentified relationships if data used to train machine learning models is not selected cautiously. Alongside three analytical objectives, I will also innovate new methods to improve model representation of differently sampled viruses based on evolutionary relatedness. Firstly, I will build models using protein sequences to predict which viruses are likely to be zoonotic and from which hosts they will originate. To better represent how viruses interact with host cells, I will build models to use information about their physical and chemical protein properties. Further models will use newer methods that can automatically find important properties straight from raw sequences. These properties can be used to find protein 'hotspots' where important signals for predicting hosts are concentrated. Models will be tested by searching for predicted zoonotic viruses in surveillance data from ongoing hospital sampling. Secondly, I will build models using host tissue and organ data. Data describing which tissues/organs are infected by each virus has already been extracted from scientific literature using text mining methods. Based on this new data, I will model the three-way network of viruses, hosts and their tissues and predict which additional tissues viruses are likely to infect. These predictions can then be tested through experimental in-vitro infection of cells from different tissues and hosts, taking advantage of synthetic viral protein toolkits. Once models are validated, further properties can be built into the network, e.g., disease severity or fatality, to predict which animal viruses have potential to cause severe human disease based on tissue patterns. Finally, I will investigate virus-host interactions in more detail by focusing on host proteins underlying patterns of infection. By combining data on infected tissues and how often those tissues express potential viral-interacting proteins, I will predict which proteins may act as barriers to viral infection and which proteins may act as viral receptors (i.e., structures that directly bind viruses and allow cell entry). Experimental in-vitro infection of cells that do/do not express potential receptor proteins will further support viral interactions identified. The proposed research will generate significant public heath impact by identifying priority viruses for targeted surveillance to prevent disease emergence and priority protein interactions for targeted experiments to develop pre-emptive therapeutics.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know