Completed Computing & AI Public Health & Healthcare

PAIR: Building a cloneable Pipeline for utilizing foundation AI on EHRs

In plain English

AI plain-English summary

Clinicians in NHS hypertension clinics have just 10 minutes per patient to sift through complex electronic health records, and this project will build an AI pipeline to automate that summary process. Hypertension affects over 14 million UK adults and contributes to half of all heart attacks and strokes. Manual data extraction from electronic health records is slow and error-prone, increasing clinician cognitive load and decision fatigue. This six-month project, based at NHS Greater Glasgow and Clyde, will develop a proof-of-concept pipeline that uses pretrained multimodal large language models to produce automated patient summaries. The pipeline will be built on synthetic data in a sandbox environment, then tested on anonymised real-world data within a trusted research environment. If successful, the pipeline will be cloneable—meaning other NHS trusts and research centres across the UK can reproduce it efficiently within their own secure data environments. This would reduce medical errors, speed up clinical decision-making, and enable personalised, risk-stratified care for hypertension patients. The project also explores patient views on AI use in healthcare, addressing concerns about data privacy and depersonalisation. The resulting tools—including reusable prompts, labelled data, and fine-tuned models—could boost NHS productivity by streamlining how clinicians interact with electronic health records.

View original technical description
Context and Challenge Hypertension is a leading global health concern, affecting over 14 million adults in the UK and contributing to 50% of heart attacks and strokes. Hypertension clinics face severe time constraints, with clinicians having just 10 minutes per patient to review complex, multimodal electronic healthcare record (EHR) data. Manual data extraction is time-consuming and error-prone, increasing cognitive load and decision fatigue. Artificial intelligence (AI)-driven patient summaries streamline this process by integrating structured and unstructured data, highlighting key clinical insights, and ensuring personalized, risk-stratified care. This reduces medical errors, enhances decision-making, and improves patient outcomes—a critical necessity amid rising hypertension cases and NHS capacity challenges. Aims and Objectives This 6-month project aims to build automated patient summaries for hypertension clinics at the NHS Greater Glasgow and Clyde Trust. This exemplar will use real-world data to develop a proof-of-concept pipeline for streamlining the application of pretrained multimodal large language models or LLMs (Llava, Qwen-VL, NVLM) for health data research and clinical care within a secure data environment. Technically, the goal is to develop a cloneable pipeline within a sandbox environment on synthetic data which can be reproduced efficiently in a trusted research environment (TRE) on anonymized routinely collected data. The planned activities are separated into three main domains: 1) the Safe Haven hosted by NHS Greater Glasgow and Clyde who are responsible for data extraction, linkage and anonymization of routinely collected health data; 2) the Trusted Research Environment (TRE) within University of Glasgow, which has necessary governance approvals to house anonymized health data for model testing; and 3) the Sandbox environment that holds synthetic and open data for pipeline development by the research team, building tools that can be exported and adopted by others. Crucial to success of this project is understanding its impact on the use of technology in healthcare delivery for patients with hypertension. Initial discussions with patients about this project have raised both positive sentiments about AI’s role in improving care through faster and more precise treatments, but also concerns about data privacy, accuracy, and de-personalization of the patient-doctor relationship. We will continue to explore patient views about AI use on EHRs for care and related information governance processes needed for building and scaling reproducible pipelines for AI models in TREs. Applications and Benefits The project will generate lessons on the application of a generic and cloneable pipeline for foundation AI on EHRs, including: an efficient approach to incorporate domain knowledge (e.g., clinical guidelines, disease codes) as well as local practitioner knowledge (disease prevalence/treatments/population characteristics) for adapting generative AI the process and interactions of human-AI collaborative working the provenance and reproducibility of model development the reuse of engineered prompts, labelled data and fine-tuned models the data schema and vocabulary for enabling LLM-based research, where we will build upon and extend existing standard common data models. This first of its kind pipeline should be easily reproducible at scale in safe havens and TREs across the UK. Its implementation would greatly facilitate both clinical and scientific research using real-world data and, more importantly, boost the productivity of the NHS through AI-enabled care.

View the original record at the funder ↗

Researchers

Alexis Webb (Principal Investigator)Honghan Wu (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Predictive Machine Learning and Digital Health for Improving Patient Outcomes
Efficient AI tools for equitable handling of missing values in population-wide e-health records to advance prevention of chronic diseases
Advanced AI-based Digital Twins For Emergency Respiratory Care (SLAIDER-QA)
PharosAI: navigating the path to AI-assisted healthcare
Integrating hospital outpatient letters into the healthcare data space

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.