The AI-SAFE-GP study: Evaluating Artificial Intelligence (AI) to SAFEly improve General Practitioner (GP) documentation and safety-netting advice in complex consultations.
GPs spend 14% of their time writing consultation notes, yet one in three appointments lacks any verbal safety-netting advice—the crucial guidance that tells patients what symptoms to watch for and when to seek help. This project tests whether AI can safely take over two tasks that human GPs currently do poorly or not at all: writing accurate consultation notes and generating written safety-netting advice. The researchers will simulate 60 complex GP consultations with actor-patients, then compare the time GPs spend writing notes against the time they spend reading and correcting AI-generated notes. They will also assess whether AI language models can reliably detect vague or missing safety-netting and produce clear written advice for patients. If the AI performs well, the impact could be substantial. NHS England has already designated AI-scribe evaluation a priority. With 160 million GP appointments in 2023 and rising, even modest improvements in documentation speed and safety-netting completeness could reduce administrative burden and improve patient safety at scale. If the AI performs poorly, the study will still produce a reusable evaluation framework and open-source code that other researchers can use to test future systems.
View original technical description
Research question Can Artificial Intelligence (AI) safely improve GP documentation and safety-netting advice in complex consultations? Background GPs report having to manage increasingly complex clinical presentations and an overwhelming administrative burden. My research shows GPs spend 14% of their time writing consultation notes, but these frequently contain errors and omissions. 'AI-scribes' can automatically generate notes from audio-recorded consultations. Companies claim AI-scribes produce time-efficiency savings and improve documentation quality. However, these claims lack independent evidence. My research also shows that verbal safety-netting advice is absent in one-third of GP consultations, more than half of GPs' safety-netting is vague, and (contrary to patients' wishes) it is rarely written. AI-Large-Language-Models (LLMs) have the potential to detect and evaluate verbal safety-netting advice, provide feedback to GPs on their safety-netting, and automatically generate written advice for patients. However, AI can omit important information and 'hallucinate' false details. Therefore, research into AI's quality and safety is required. Work package (WP) aims Assess how AI-scribes may impact on complex GP consultations. Assess the quality of AI-scribes. Assess the capability of AI-LLMs to classify safety-netting behaviours and provide written safety-netting advice for patients. Methods WP1: AI-scribe impact on consultations (months 1-21) Co-develop (with patients and GPs) 60 diverse complex clinical scenarios, and record the resulting consultations (10 GPs, 6 actor-patients each) in a high-fidelity simulated clinical environment. Compare time taken for GPs to write consultation notes vs. time to read/amend AI-scribe notes. Conduct semi-structured interviews with the participating GPs about their experiences and perceived risks/benefits of AI-scribes. WP2: AI-scribe quality (months 16-35) Update the 'consultation checklist' methodology, where everything verbalised during a consultation is itemised and evaluated as 'critical', 'non-critical' or 'irrelevant' for documentation, to include GP consensus on list items and their criticality. Recruit 5 GPs to judge the criticality of consultation items for documentation. Compare WP1 GP and AI-scribe notes in terms of: Completeness: quantity of critical and non-critical items. Correctness: quantity of false items. Conciseness: quantity of irrelevant items / overall length. WP3: AI to improve safety-netting (months 26-48) Develop AI-LLM instructions to classify GP safety-netting behaviours using my published Safety-Netting Coding Tool (SaNCoT). Compare AI-LLM SaNCoT coding vs. human expert (me) to evaluate if AI-LLMs could automatically give feedback to GPs. Co-produce a template for written safety-netting advice with patients and GPs. Assess the accuracy of AI-completed templates and evaluate the quality/readability with 20 patients. Impact and dissemination 160 million GP appointments occurred in 2023 and this number is predicted to rise annually. AI utilisation in healthcare is likely to become widespread over the next 5-10 years. Therefore, the following outputs will be highly impactful: Understanding strengths/limitations of AI-scribes - this has been designated a priority by NHS England. Reusable process for evaluating AI-scribes. Open-source code for classifying and generating written safety-netting advice with an evaluation of its quality. Reusable archive of complex GP consultations. Findings will be presented in traditional (peer-reviewed publications, conferences) and non-traditional (blog posts, podcasts, animated videos, webinars) formats. A targeted stakeholder map will be generated to maximise impact.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know