The Clinical Language Informatics group is led by Arlene Casey. Their research includes developing language technologies, health intelligence and privacy science to make information in healthcare narratives useful for rigorous, trustworthy research. Summary Some of the richest information about health and care is written rather than coded. Clinical notes, reports and letters capture symptoms, treatment decisions, response to care, changing health and wellbeing, and the context surrounding a person's healthcare journey. These narratives can transform research, but they are also sensitive, complex and created within a relationship of trust.Our work asks two questions together: what can healthcare narratives tell us, and how can that information be used safely and responsibly? We combine computational methods with clinical validation, privacy assessment, governance, and patient and public involvement.Our research covers three connected themes:Clinical text analytics and language technologies: developing and evaluating natural language processing, foundation models and information extraction methods for clinically meaningful information.Health intelligence and discovery science: turning information from narratives into research-ready data for epidemiology, population health, healthcare analytics and translational research.Privacy science, governance and responsible AI: developing evidence-based approaches to privacy risk, Trusted Research Environments and the responsible use of clinical language models.We work closely with DataLoch – a partnership between the university and NHS Lothian - to connect methodological research with the secure data, governance and Trusted Research Environment expertise needed to translate new approaches into sustainable research capability. This collaboration supports safe access to clinical free text, privacy-risk assessment, NLP-derived research variables and responsible model development. Primary Contact Arlene Casey Group Leader | Vivensa Senior Research Fellow | Strategic and Operational NLP Lead, DataLoch Contact details Email: arlene.casey@ed.ac.uk People NameRoleArlene CaseyGroup Leader | Vivensa Senior Research Fellow | Strategic and Operational NLP Lead, DataLochFranz GruberNLP Research FellowJudit KutiNLP Research FellowFahrurrozi RahmanNLP Research FellowMatúš FalisNLP Research FellowSam McInerneyClinical Fellow, PhD Projects AMBER — Antidepressant Measures and Biological Exposure and ResponseWellcome, PI Cathryn Lewis Kings College LondonAMBER investigates the biological, genetic and clinical determinants of antidepressant response by integrating electronic health records, genomics and participant-focused research. Our group leads the development of clinical text informatics methods for deriving robust measures of antidepressant exposure and response. We compare large and smaller language models to identify depression-related consultations, treatment response, medication switching, side effects and other indicators of drug response in GP and hospital records.Can AI Tell the Story of Cancer?Sam McInerney, Clinical Fellow PhD, evaluates whether large language models can reliably extract information about disease, treatment, biomarkers, response and progression from cancer records. The longer-term ambition is to develop patient knowledge graphs representing an individual's cancer journey over time and concise summaries of complex treatment and outcome histories.Senior Proleptic FellowshipArlene Casey's five-year Vivensa Senior Research Fellowship investigates how healthcare narratives can be used safely to improve research into later-life health. Understanding Alcohol, Drug and Self-Harm Presentations in Emergency CareCollaboration with Dr Chris HumphriesRoutine emergency department coding does not always capture the full circumstances surrounding a person's attendance. This work evaluates whether clinical narratives can provide a more complete picture of alcohol-, drug- and self-harm-related presentations by comparing routine coding with clinician-adjudicated annotations and validated language-model approaches.STAR-TREMRC, PI Arlene CaseySTAR-TRE investigates how privacy risk within healthcare narratives can be understood well enough to enable safe research access. The project studies variation in risk across patient groups and data types, investigates language models for identifying contextual privacy risks, and develops practical governance methods for free-text access within Trusted Research Environments. A UK-wide patient and public involvement programme explores views on sensitive information and AI-assisted privacy-risk assessment.TransPECT — Beyond the AirlockMRC, PI Arlene CaseyTransPECT examines whether language models trained on sensitive healthcare data can memorise or reveal private information. It identifies the characteristics of models, data and training approaches that influence disclosure risk and develops methods for evaluating safe model release from Trusted Research Environments. The project also contributes to national guidance and training through the UK TREvolution programme. Patient and Public Involvement and Engagement Research involving sensitive healthcare narratives cannot be guided by technical performance alone. Patients and members of the public help us understand what information feels sensitive, which safeguards are expected and where human judgement should remain central.Public involvement was central to the SARA DARE UK Driver Project, where workshops and wider consultation informed approaches to privacy-risk assessment in clinical free text and data provenance. STAR-TRE extends this work through a UK-wide programme exploring public expectations of AI-assisted de-identification and secure researcher access to free-text data. Publications Publications from this research group can be found on Arlene's Edinburgh Research Explorer pages. Arlene Casey | Edinburgh Research Explorer Humphries C, et al. Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model. medRxiv. 2026.Humphries C, et al. Natural language processing-driven knowledge graphs for public health intelligence in urgent and emergency care. 2026.Alex B et al. GS-BrainText: A multi-site brain imaging report dataset from Generation Scotland for clinical natural language processing development and validation. 2026.Ford E, et al. What is the patient re-identification risk from using de-identified clinical free text data for health research? AI and Ethics. 2025.Casey A et al. A systematic review of natural language processing applied to radiology reports. BMC Medical Informatics and Decision Making. 2021;21:179. Themes and keywords Scientific Themes Clinical language informatics; health intelligence; population health; discovery science; later-life health; cancer informatics; privacy science; responsible AI; patient and public involvement. Methodological Keywords Natural language processing; large language models; foundation models; information extraction; clinical phenotyping; knowledge graphs; clinical validation; privacy-risk assessment; de-identification; Trusted Research Environments; information governance. This article was published on Wednesday 20 May 2026