Clinical Language Informatics

The Clinical Language Informatics group is led by Arlene Casey. Their research includes developing language technologies, health intelligence and privacy science to make information in healthcare narratives useful for rigorous, trustworthy research.

Summary

Some of the richest information about health and care is written rather than coded. Clinical notes, reports and letters capture symptoms, treatment decisions, response to care, changing health and wellbeing, and the context surrounding a person's healthcare journey. These narratives can transform research, but they are also sensitive, complex and created within a relationship of trust.

Our work asks two questions together: what can healthcare narratives tell us, and how can that information be used safely and responsibly? We combine computational methods with clinical validation, privacy assessment, governance, and patient and public involvement.

Our research covers three connected themes:

  1. Clinical text analytics and language technologies: developing and evaluating natural language processing, foundation models and information extraction methods for clinically meaningful information.
  2. Health intelligence and discovery science: turning information from narratives into research-ready data for epidemiology, population health, healthcare analytics and translational research.
  3. Privacy science, governance and responsible AI: developing evidence-based approaches to privacy risk, Trusted Research Environments and the responsible use of clinical language models.

We work closely with DataLoch – a partnership between the university and NHS Lothian -  to connect methodological research with the secure data, governance and Trusted Research Environment expertise needed to translate new approaches into sustainable research capability. This collaboration supports safe access to clinical free text, privacy-risk assessment, NLP-derived research variables and responsible model development.

Primary Contact

Arlene Casey

Group Leader | Vivensa Senior Research Fellow | Strategic and Operational NLP Lead, DataLoch

Contact details

People

NameRole
Arlene CaseyGroup Leader | Vivensa Senior Research Fellow | Strategic and Operational NLP Lead, DataLoch
Franz GruberNLP Research Fellow
Judit KutiNLP Research Fellow
Fahrurrozi RahmanNLP Research Fellow
Matúš FalisNLP Research Fellow
Sam McInerneyClinical Fellow, PhD

Projects

AMBER — Antidepressant Measures and Biological Exposure and Response

Wellcome, PI Cathryn Lewis Kings College London

AMBER investigates the biological, genetic and clinical determinants of antidepressant response by integrating electronic health records, genomics and participant-focused research. Our group leads the development of clinical text informatics methods for deriving robust measures of antidepressant exposure and response. We compare large and smaller language models to identify depression-related consultations, treatment response, medication switching, side effects and other indicators of drug response in GP and hospital records.

Can AI Tell the Story of Cancer?

Sam McInerney, Clinical Fellow PhD, evaluates whether large language models can reliably extract information about disease, treatment, biomarkers, response and progression from cancer records. The longer-term ambition is to develop patient knowledge graphs representing an individual's cancer journey over time and concise summaries of complex treatment and outcome histories.

Senior Proleptic Fellowship

Arlene Casey's five-year Vivensa Senior Research Fellowship investigates how healthcare narratives can be used safely to improve research into later-life health. 

Understanding Alcohol, Drug and Self-Harm Presentations in Emergency Care

Collaboration with Dr Chris Humphries

Routine emergency department coding does not always capture the full circumstances surrounding a person's attendance. This work evaluates whether clinical narratives can provide a more complete picture of alcohol-, drug- and self-harm-related presentations by comparing routine coding with clinician-adjudicated annotations and validated language-model approaches.

STAR-TRE

MRC, PI Arlene Casey

STAR-TRE investigates how privacy risk within healthcare narratives can be understood well enough to enable safe research access. The project studies variation in risk across patient groups and data types, investigates language models for identifying contextual privacy risks, and develops practical governance methods for free-text access within Trusted Research Environments. A UK-wide patient and public involvement programme explores views on sensitive information and AI-assisted privacy-risk assessment.

TransPECT — Beyond the Airlock

MRC, PI Arlene Casey

TransPECT examines whether language models trained on sensitive healthcare data can memorise or reveal private information. It identifies the characteristics of models, data and training approaches that influence disclosure risk and develops methods for evaluating safe model release from Trusted Research Environments. The project also contributes to national guidance and training through the UK TREvolution programme.

Patient and Public Involvement and Engagement

Research involving sensitive healthcare narratives cannot be guided by technical performance alone. Patients and members of the public help us understand what information feels sensitive, which safeguards are expected and where human judgement should remain central.

Public involvement was central to the SARA DARE UK Driver Project, where workshops and wider consultation informed approaches to privacy-risk assessment in clinical free text and data provenance. STAR-TRE extends this work through a UK-wide programme exploring public expectations of AI-assisted de-identification and secure researcher access to free-text data.

Publications

Publications from this research group can be found on Arlene's Edinburgh Research Explorer pages.

Themes and keywords

Scientific Themes

Clinical language informatics; health intelligence; population health; discovery science; later-life health; cancer informatics; privacy science; responsible AI; patient and public involvement.

Methodological Keywords

Natural language processing; large language models; foundation models; information extraction; clinical phenotyping; knowledge graphs; clinical validation; privacy-risk assessment; de-identification; Trusted Research Environments; information governance.