A patient’s prescription record might reveal a hidden cancer months before other symptoms appear. This study will systematically mine GP prescribing data for eight hard-to-diagnose cancers—including pancreatic, ovarian, and stomach cancer—to find which medications people start taking in the year before their diagnosis, compared with matched healthy controls. Currently, many cancers are caught late because early signs are vague, like indigestion or back pain. If a sudden new prescription for, say, antacids or painkillers reliably precedes a cancer diagnosis, GPs could use that signal to trigger earlier investigation. The team will analyse the 100 most common medication classes from linked Welsh health records, using both standard statistics and machine learning to spot patterns. Success would turn routine prescribing data into a low-cost, passive early-warning system—no extra tests or appointments needed. This is applied health data science, not fundamental biology; it aims to create a practical tool for primary care, not to explain how cancers develop.
View original technical description
Background Many patients with undiagnosed cancer have non-specific symptoms making early diagnosis difficult. Changes in prescription medications could offer an opportunity to identify patients for earlier cancer investigation and detection. There has not been an attempt to systematically investigate new use of prescription medications before cancer onset in a range of cancers to determine whether they could offer opportunities to detect cancer earlier. Aims To systematically assess prescription medications to identify previously unrecognised medications which are newly used in the period before cancer diagnosis. Methods The study will be conducted within the SAIL Databank utilising linked cancer registry data and primary care data. The following eight cancers will be investigated separately: multiple myeloma, pancreatic cancer, stomach cancer, ovarian cancer, lung cancer, non-Hodgkin’s lymphoma, renal cancer, and colorectal cancer. A series of nested case-control studies will be conducted comparing patients with cancer to five matched randomly selected cancer-free population-based controls. New medication use in the year before diagnosis\index date will be determined for the 100 most common medication classes from general practice prescription records. The proportion starting medications will be determined in cases and controls. Conditional logistic regression models will be used to calculate odds ratios and confidence intervals for each medication class accounting for age, gender, year and comorbidities. Medication classes associated with cancer, after correction for multiple testing, will be considered signals and further investigated. Additional medication classes will be identified after applying random forest, machine learning, algorithms. For each identified medication-cancer signal, detailed analyses will be conducted on the timing of new use of the medication to identify the month before cancer diagnosis when changes in prescribing occurred. How the results of this research will be used Our study has the potential to identify previously unrecognised medications which are newly used in the period up to two years before cancer diagnosis. New use of these medications could themselves act as an alert for considering earlier cancer investigation or point to unrecognized symptom patterns. Further research will be required to determine the utility of new use of these medications to act as signals both alone and alongside other predictors of undiagnosed cancer in an independent dataset.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know