Jennings, Charlotte Nicola
ORCID: https://orcid.org/0000-0003-1584-4057
(2026)
Defining and addressing data needs in pathogenomic cancer research.
PhD thesis, University of Leeds.
Abstract
Multimodal artificial intelligence approaches that integrate histopathology whole slide images with high-throughput genomic data offer significant potential to advance precision oncology. While substantial research effort has focused on developing increasingly complex predictive models, comparatively less attention has been paid to the quality, structure, and availability of the data required to support robust and generalisable pathogenomic AI. This thesis addresses this gap by examining the data ecosystem underpinning pathogenomic cancer research, with a focus on data generation, curation, and reuse.
A systematic review is first conducted to evaluate machine learning–based prognostic models that integrate pathology images and genomic or other high-throughput -omic data for overall survival prediction in cancer. The review assesses model performance, data sources, integration strategies, and methodological quality, and identifies substantial heterogeneity in data handling, limited external validation, and high risk of bias across much of the literature. These findings highlight that reported performance gains from multimodal modelling are difficult to interpret without greater transparency and methodological rigor.
Surveys of key stakeholders involved in the creation and use of pathology whole slide image datasets are then presented, providing empirical insight into practical challenges surrounding data access, standardisation, quality control, and workforce capacity.
The thesis subsequently describes the design, creation, and evaluation of the Genomics Pathology Imaging Collection, a large-scale multimodal dataset linking digitised pathology slides with genomic data from a national cancer cohort, and characterises its coverage, quality, and operational limitations.
Finally, approaches for curating whole-case pathology image datasets are developed and evaluated, including manual and automated methods. The results demonstrate that auto-assisted curation can substantially reduce time and cost while maintaining the quality associated with expert oversight.
Overall, this work argues that data curation and infrastructure are central scientific challenges in pathogenomic AI, and that addressing these foundations is essential for enabling reproducible, scalable, and clinically meaningful multimodal cancer research.
Metadata
| Supervisors: | Treanor, Darren and Westhead, David |
|---|---|
| Keywords: | digital pathology; genomics; pathogenomics; multimodal; data; curation; artificial intelligence |
| Awarding institution: | University of Leeds |
| Academic Units: | The University of Leeds > Faculty of Medicine and Health (Leeds) > School of Medicine (Leeds) |
| Date Deposited: | 16 Jul 2026 09:10 |
| Last Modified: | 16 Jul 2026 09:10 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:38905 |
Download
Final eThesis - complete (pdf)
Embargoed until: 1 July 2028
Please use the button below to request a copy.
Filename: Jennings_CN_Medicine_PhD_2026.pdf
Export
Statistics
Please use the 'Request a copy' link(s) in the 'Downloads' section above to request this thesis. This will be sent directly to someone who may authorise access.
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.